Public LLM benchmarks are mostly garbage (grandpacad.com via hn)
model roundup
Opus 4.7
-
OpenRouter throughput, Design Arena ELO, and token prices said skip Opus 4.7. My own 3D eval suite said the opposite.
-
How did we make DeepSeek outperform Opus (twitter.com via hn)
how did we make deepseek outperform opus 4.7? i've been thinking about why "open model bad at tool calling" is almost always a harness problem, not a model problem.
-
Best Claude workflow for converting large PDFs/books into one revision handbook? (www.reddit.com via reddit)
Hi everyone, I'm using Claude Pro (Opus 4.7 with Extended Thinking) and I'm trying to build a single consolidated UPSC revision handbook from multiple sources. My inputs include: Multiple PDFs (detailed notes) Another set of concise notes…
-
Oikoumene: Autonomous Agent Civilization Simulator (github.com via hn)
A research project by GeoLambda GmbH This simulation was developed primarily with Claude Code, Anthropic's agentic CLI, using both Claude Opus 4.6 and Opus 4.7. The collaboration served as a real-world stress test of the latest coding LLM…
-
Switch from Claude Opus 4.7 to Claude Fable without losing anything. (www.reddit.com via reddit)
Hello, I had a very long conversation with Claude Opus 4.7. With the reappearance of Fable, I would like to use this model to resolve a complex legal interpretation within a short timeframe (48 hours).