model roundup

Haiku 4.5

6 items · started 2026-07-17 · closed 2026-07-24

  1. We measured AI gateway performance with an open-source TTFT benchmark: 75 cold + 75 warm interleaved runs of claude-haiku-4.5 against LLM Gateway and OpenRouter, with phase-by-phase timings and raw data published. LLM Gateway reached first…

  2. I benchmarked Claude Haiku 4.5 against 8 free/alt models (NVIDIA NIM plus a second free-tier provider) on a real task: writing outreach proposals from actual job postings, not a synthetic prompt. Same production prompts, same 3 real jobs,…

  3. I think Haiku is pretty much underestimated and I feel it every day when the main agent spawns the Explore tool with Haiku 4.5 which hallucinates the F out of the codebase. Of course I could include in CLAUDE.md that a "better" model shoul…

  4. I built a small benchmark to answer one question: is Claude Opus 4.6 actually worse than 4.8, or does it just feel older? It scores the half of code review that most evals ignore, which is restraint - not flagging correct code that looks s…

  5. So I just checked my usage and I found this Session Total cost: $17.77 Total duration (API): 24m 43s Total duration (wall): 40m 54s Total code changes: 1513 lines added, 0 lines removed Usage by model: claude-haiku-4-5: 604 input, 18 outpu…

  6. Setup: - 38 tasks - 2 Claude models (Haiku 4.5, Sonnet 5) × 5 reps, + a 1-rep Opus 4.8 probe, - same live workspace. - Deterministic state checks + an arm-blind LLM judge.

← all threads