model roundup

Opus 4.6

19 items · started 2026-06-24 · closed 2026-07-11

  1. I've been writing toy kernels and working on operating system projects since my childhood, and it's partly how I learned C. That includes this project, MontaukOS, which I started early in 2025, where I wrote a lot of the fundamental kernel…

  2. Ever since the release of Opus 4. 7, I made sure to stick to Opus 4.

  3. Go back to Opus 4.6. You'll thank me!

  4. # RFC: A Distributed Behavioral Policy Mesh for Cross-Model Skill Evolution **Status:** Request for Comments **Author:** J.S. Colson (GitHub: [swordsman](https://github.com/swordsman)) — jscolson+decentralfabcollab@gmail.com **AI Collabora…

  5. I love Fable for law related matters, but I mostly used Claude for biological theorizing and speculative biology. Opus4.6 is so far still the best corroborator I ever had, nerfed and unnerfed.

  6. Since yesterday, whenever I use opus 4.6 for one of my projects, something really weird happens after the conversation gets a bit long. As soon as the chat history hits about 10 messages or more, the model starts claiming in every single r…

  7. We have been focused on AI error distribution for the past year, and in our last research paper, "Architecture of Errors" showed mathematically that an AI solution needs a finite set of interventions to perform well in a bounded patch doma…

  8. With Fable’s return the first thing I tested was just a silly prompt to see how over-reactive the guardrails are. As expected, they are just as bad as initial release.

  9. I’m late to this, but I couldn’t find a post about it here. If this has already been shared, feel free to remove.

  10. When Opus 4.6 first came out and was amazing, I wanted to explore the limits of LLM creativity and play around with ways of injecting novelty/expanding range. Initially I just wrote up a small skill that told Claude essentially you're an a…

  11. Please advise. I'm currently only running one project seriously.

  12. So recently I decided that it would be nice to run my agent against some popular benchmarks. And oh my god, the cost to run a single benchmark, such as terminal-bench or swe-bench will cost you thousands of dollars in tokens just for a sin…

  13. I'm an SDE and I feel like I'm not getting much productivity out of coding agents. Yeah, they can generate code and build features, but most of what I get isn't really deployment-ready or easy to maintain.

  14. I’m just generally curious what conversations you’re having with the models. How much are you using it throughout the day?

  15. Anybody else having new issues with thinking blocks not rendering? I've always had extended (now "adaptive") thinking ON, which consistently renders thinking blocks (even for 4.7 & 4.8, at least in claude.ai).

  16. It's been a week and a half without Fable for almost all of us and I have used this time for some reflection. The pricing and access concerns were a lot to take in even before the feds pulled the plug, but for whatever reason this intermis…

  17. Large Language Models (LLMs) have emerged as a promising tool for automated vulnerability detection, yet their effectiveness on web-specific vulnerabilities remains to be explored. This work benchmarks six frontier (Claude Opus 4.6, Codex…

  18. Do you guys think there is any chance for an Opus 3 situation with Opus 4.6? Honestly it’s the best for me.

  19. Working with Claude on an desktop app in cowork. What I have found is that Opus 4.6 has limited patience for indecisiveness.

← all threads