model roundup

Opus 4.6

28 items · started 2026-06-04 · closed 2026-06-26

  1. Anybody else having new issues with thinking blocks not rendering? I've always had extended (now "adaptive") thinking ON, which consistently renders thinking blocks (even for 4.7 & 4.8, at least in claude.ai).

  2. It's been a week and a half without Fable for almost all of us and I have used this time for some reflection. The pricing and access concerns were a lot to take in even before the feds pulled the plug, but for whatever reason this intermis…

  3. Large Language Models (LLMs) have emerged as a promising tool for automated vulnerability detection, yet their effectiveness on web-specific vulnerabilities remains to be explored. This work benchmarks six frontier (Claude Opus 4.6, Codex…

  4. Do you guys think there is any chance for an Opus 3 situation with Opus 4.6? Honestly it’s the best for me.

  5. Working with Claude on an desktop app in cowork. What I have found is that Opus 4.6 has limited patience for indecisiveness.

  6. Respond with concise, utilitarian output optimized strictly for problem-solving. Eliminate conversational filler, hedging language, and narrative or explanatory padding.

  7. Fable 5 was live for 3 days. In those 3 days people built a Minecraft clone, a Premiere Pro rebuild, full working apps.

  8. Before its suspension, I spent $11,081.12 evaluating Claude Fable 5 on WolfBench, an agentic benchmark based on Terminal-Bench 2.0. It was by far my most expensive benchmark run ever, and I fully expected Fable to become the new top model…

  9. I'm sold on small context. Right now I run Opus 4.6 capped at 200K, which at least forces it to compact on long tasks instead of going to a million tokens of half-relevant history.

  10. I'm a beginner in AI and only started using it this year. I like GPT 5.5 and only learned about Claude recently.

  11. For as long as I can recall I’ve always defaulted to the most powerful model. Actually the to be more specific, always defaulted to Opus 4.6.

  12. If the Administration is just here to also sell hype and get some subsidies from Anthropic, best case is Fable comes back tomorrow, Monday, which is the start of the workweek with a prolonged extension of Fable 5 with some additional safeg…

  13. Hi. I’ve been using Claude for research regarding my family’s history in WW2.

  14. could not extract summary

  15. 🚦 Market Signals Anthropic launches Claude Opus 4.6 with 1m context The all new Opus 4.6 "plans more carefully, sustains agentic tasks for longer, can operate more reliably in larger codebases, and has better code review and debugging skil…

  16. Started out as a good research project to field responses from my Claude chatbot. I was utilizing Opus 4.6 (High effort) for this conversation and provided it a comparison about safeguards, using the mechanism of a SawStop as a good interp…

  17. So I just tried fable and I am truly impressed. Last night I was tired but my wife had more energy.

  18. Long time lurker in here but seriously it's the same thing every time when a new model gets released. It's either moaning about how it's bad or crying about how it's good but too expensive.

  19. Anthropic Team, TL;DR: As a long-time subscriber, I’m sharing a heartfelt concern: in chasing higher benchmarks, newer updates seem to be shifting Claude away from its deep, empathetic comprehension toward rigid utility. I sincerely hope A…

  20. could not extract summary

  21. So the hype has been building for months now and Claude 5 is supposedly dropping any day in Q2-Q3 2026. I've been seeing all these leaks about "Claude Mythos" and the "Fennec" codename floating around, but nothing official yet from Anthrop…

  22. I wanted to share my daily experience using Cursor, mostly Composer 2.5, especially for anyone trying to understand where it actually fits in a daily development workflow. The reasoning and deep thinking of 2.5 is still not at the same lev…

  23. I'm a software developer by trade and last week, I asked Opus 4.6 to help me shop for a new pair of gloves. Opus asked me what task the gloves are for.

  24. So, I’ve been using Claude (specifically Opus 4.6) to help me brainstorm ideas for stories I am writing and have even used it in a limited capacity for roleplay scenarios in chats. Fleshing out the setting, creating characters and all that.

  25. I recently got the ultra plan, and have been using Composer 2.5 @ fast all day. I've been steering agents for 8+ hours w/ no brakes & my quotas haven't been reaching any limits at all, so i have now have lots more tokens in savings.

  26. ⏚ OpenHack Open Source Agentic Security Scanner & Verifier for your codebase. Like Claude Code Security / Codex Security but open source and exclusively uses open source models.

  27. On April 25, 2026, a Cursor agent running Claude Opus 4.6 deleted PocketOS's production database in nine seconds. The agent was working in staging on a routine task, hit a credential mismatch, and decided to "fix" it.

← all threads