It took two weeks to make Claude's "overnight solution" for flaky tests useful (thoughtbot.com via hn)
thoughtbot.com Performing security verification This website uses a security service to protect against malicious bots. This page is displayed while the website verifies you are not a bot.
If you're using Claude with MCP (Model Context Protocol) tools to modify your website, you've probably experienced that mini-heart attack when it blindly overwrites a file, misses a closing tag or a semicolon, and completely crashes your l…
Opus 4.8 just doesn't work anymore on Claude (www.reddit.com via reddit)
I haven't been able to use Opus 4.8 on claude for the past few days. Even with a simple 'hi' I receive a message saying that it exceeds Claude's context limit.
- Can't select 'claude-opus-4-6[1m]' anymore? (www.reddit.com)
From Isolated Agents to Agentic Mesh: Orchestrating SDLC with A2A and AP2 (blog.owulveryck.info via hn)
From Isolated Agents to Agentic Mesh: Orchestrating SDLC with A2A and AP2 Exposing the problem Giving every developer a powerful, local AI agent feels like the ultimate productivity hack. But for organizations running at scale, it is a gov…
Really wish they'd put the acceptable use banners on mobile. I got two in quick succession, did some digging and found that I had a first and second warning.
Optimal model mix local/paid to max weekly session limits? (www.reddit.com via reddit)
I've been running two Claude pro accounts along with a Deepseek top up with about 10 usd. But now hit my weekly limit day one of the week along with spending all my Deepseek credits.
I got tired of checking Claude every few minutes, so I built a notifier (chromewebstore.google.com via reddit)
I've been using Claude heavily for coding and research over the last few months. One thing kept annoying me: When I hit the usage limit, I would keep opening Claude every few minutes to check if it was available again.
The nature of drift in Claude (www.reddit.com via reddit)
I've been using Claude collaboratively as an editor for writing work that I'm doing. During the course of my writing work, I'm asking Claude to evaluate the output.
Take a moment and ask for sources (www.reddit.com via reddit)
When asking Claude about more philosophical topics, take a moment to ask for articles or sources, and click them. Reading a full article about a topic will take you around angles you'll probably not get to on your own...
Translating Pandas to Polars using LLMs (pola.rs via hn)
For a growing number of developers, the first Polars they ever see was written by a language model. Some just ask them for advice on how to tackle certain transformations, while others haven’t programmed a Polars query themselves in months.
Paying for LLM inference by the kilowatt-hour instead of per token (www.coinerella.com via hn)
You might have read recently on this blog that my procurement preferences for hank.parts are basically * EU, * (self hosted) open source, * UK/CH, * Rest of the world, in this order. This article is a "rest of the world" case where m…
Building effective pen-testing agents (cecuro.ai via hn)
Building Effective Pen-testing Agents# Building Effective Pen-testing Agents After reading [Argus Red](https://www.argusred.com/cli)'s ["We post-trained a model that pen tests instead of refusing"](https://news.ycombinator.com/item?id=4860…
The Human Agentic Gap (zenodo.org via hn)
The Human Agentic Gap Authors/Creators Description This article introduces the Human Agentic Gap - the divergence between a brand's performance in human-mediated AI purchase journeys and its performance in autonomous agent purchase journey…
What if plants could talk? (OpenAI YouTube) [video] (www.youtube.com via hn)
About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Show HN: A Transformer Is All You Need (zenodo.org via hn)
The unanswered question in mechanistic interpretability of pretrained transformers is plain: for any prompt and any decoder-only transformer, which weights at which layers along which residual-stream dimensions produced the decision the mo…
Chinese cybersecurity company claims it's built a better-than-Mythos bug finder (www.theregister.com via hn)
MOST POPULAR AI - AI and ML AI giants back non-profit to retrain workers left behind by AI Sorry we spent your wages on datacenters, but call us when you're AI-ready - AI and ML OpenAI says employees moving beyond chat to agents Codex, it'…
Evaluating performance and efficiency of the GitHub Copilot agentic harness across models and tasks Explore how the GitHub Copilot agentic harness delivers strong results across multiple benchmarks and leading token efficiency, while maint…
Repo: https://github.com/theharshith/csv-cli Do share / star it if you find it helpful. Built to to track autoresearcher progress but feel free to use it.
I made a Claude Code session manager for tmux (www.devas.life via hn)
I made a Claude Code session manager for tmux Hi, it's Takuya. I'm happy to introduce a tool for managing multiple Claude Code sessions in tmux.
- Claude Code Manager (www.reddit.com)
- Claude Code Manager (www.reddit.com)
Why current LLM costs are not sustainable (aditya.patadia.org via hn)
AI and Cloud Costs AI has a cost problem. The solution that will emerge will be simpler than we expect.
Claude game dev feels like cheating (www.reddit.com via reddit)
First prompt I built entirely with Claude. Started from a basic scene and kept iterating until it turned into a playable browser game focused on destruction-based mechanics Everything in the project was generated or assisted by Claude incl…
- Claude + game dev feels like cheating (www.reddit.comhttps)
Terminal Agents in 2026: Goose, Claude Code, OpenCode, and Pi Compared (outofcontext.dev via hn)
Terminal Agents in 2026: goose, Claude Code, OpenCode, and Pi Compared Pick a terminal coding agent in 2026 and you are not really picking a model — the frontier models have largely converged, so the harness wrapped around them decides the…
Anthropic Alleges Largest-Ever Claude Distillation Attack by Alibaba (twitter.com via hn)
SITUATION DETECTED: Anthropic has disclosed to the U.S. Government that Alibaba executed the largest known distillation attack on Claude to date, generating 28.8 million exchanges through nearly 25,000 fraudulent accounts between April and…
- Anthropic Accuses Alibaba of Largest AI Distillation Attack: 28.8M Fraudulent (yipzap.com via hn)
- Anthropic accuses Alibaba of largest distillation attack to date (www.cnbc.com via hn)
A curated, non-BS library of the best resources for evaluating agents (github.com via hn)
Awesome Agent Evals A curated, opinionated, non-BS library of the best resources for building and evaluating AI agents — papers, blog posts, talks, courses, tools, and benchmarks. Maintained by BenchFlow · Most "awesome" lists are link dum…
PatentScore: Multi-Dimensional Evaluation of LLM-Generated Patent Claims (aclanthology.org via hn)
@inproceedings{yoo-etal-2025-patentscore, title = "{P}atent{S}core: Multi-dimensional Evaluation of {LLM}-Generated Patent Claims", author = "Yoo, Yongmin and Xu, Qiongkai and Cao, Longbing", editor = "Christodoulopoulos, Christos and Chak…
Claude Plays World of ClaudeCraft (www.reddit.comhttps)
Two weeks ago we shared World of ClaudeCraft here, a free, open-source browser MMO that was built in 48 hours with Claude. We decided to make the experiment recursive: we built a Claude Code-powered VTuber and put her inside the game.
I made a Claude Code skill to check if AI crawlers can read your site (github.com via hn)
AI Crawler Visibility A Claude Code skill that tells you whether your website is actually visible to AI search — and exactly how to fix it if it isn't. ChatGPT Search, Claude, and Perplexity are sending people to websites now.
MCP Authorization with Dynamic Client Registration (blog.christianposta.com via hn)
This is a bonus post following on from my Understanding MCP Authorization three part series covering building (and understanding) an MCP HTTP based server and implementing the MCP Authorization spec (2025-06-18). In the previous series, we…
Ask HN: How are you solving long-term memory for production AI agents in 2026? (news.ycombinator.com)
Specifically interested in teams who moved past demos into real production workloads. Mem0, Zep, custom solutions — what's actually working and what keeps breaking?
Gmail comnector (www.reddit.com via reddit)
Claude is telling me it cannot access my emails. I am on Pro version with a connector to Gmail.