model roundup

Gemini 3.7

4 items · started 2026-08-30 · closed 2026-09-05

  1. Mushroom identification with AI: GPT-5.6-Sol, Gemini 3.7 Flash, GLM-5.3-Flash and Claude Fable 5.1 benchmarked on poisonous and edible species of FungiTastic. A lot of dangerous errors.

  2. I ran Claude Fable 5.1 on the current 98-task MindTrial set with the same Python executor available as in the earlier Fable 5, Opus 5 and Sonnet 5 runs. The result was stronger than I expected: 90/98, which is currently the highest raw pas…

  3. Magnus is built for agents that need to think, use tools, and get things done, without waiting around. All 97 tasks of τ³-bench banking, head-to-head against gpt-5.6-sol, gpt-5.6-luna and gemini-3.7-flash, reasoning effort as published.

  4. could not extract summary

← all threads