model roundup

Haiku 4.5

3 items · started 2026-07-16 · closed 2026-07-19

  1. I built a small benchmark to answer one question: is Claude Opus 4.6 actually worse than 4.8, or does it just feel older? It scores the half of code review that most evals ignore, which is restraint - not flagging correct code that looks s…

  2. So I just checked my usage and I found this Session Total cost: $17.77 Total duration (API): 24m 43s Total duration (wall): 40m 54s Total code changes: 1513 lines added, 0 lines removed Usage by model: claude-haiku-4-5: 604 input, 18 outpu…

  3. Setup: - 38 tasks - 2 Claude models (Haiku 4.5, Sonnet 5) × 5 reps, + a 1-rep Opus 4.8 probe, - same live workspace. - Deterministic state checks + an arm-blind LLM judge.

← all threads