model roundup

Gemini 3.1

8 items · started 2026-06-07 · closed 2026-06-13

  1. Pelican on a Bicycle: Claude Fable 5 vs GPT-5.5 Pro vs Gemini 3.1 Pro We asked the top frontier AI models — launch-day Claude Fable 5, GPT-5.5 Pro and Gemini 3.1 Pro — to draw a pelican riding a bicycle as SVG code. Same prompt, one shot,…

  2. Is this benchmark broken, or is Anthropic benchmaxing? LiveBench

  3. Overall an improvement over Opus 4.8, but I'd still say Gemini 3.1 Pro has more of an artistic vision even tho it fails tool calls and writes buggy code sometimes. Ik almost everyone is interested just in the SWE stuff, but this has been a…

  4. Hopefully this isn't too low effort of a post. I just finished the benchmarks and I figured I'd post them online because they certainly were insightful for me.

  5. could not extract summary

  6. Over the last few weeks I've been comparing the latest frontier AI models, including Claude Opus 4.8, GPT-5.5, Gemini 3.1 Pro, Grok 4.3, Perplexity AI and DeepSeek V4-Pro. Instead of focusing only on benchmark scores, I looked at: Real-wor…

  7. Tried Gemma, Qwen and a few others. Need vision and larger context windows for an application I am working on.

  8. Title

← all threads