model roundup

GPT 5

2 items · started 2026-07-24 · closed 2026-07-29

  1. This started from a simple observation: agents are great at judgment and terrible at discipline. Every rule I enforced through prompts ("don't poll", "don't claim completion") eventually broke.

  2. I ran the 62-item politicalcompass.org test 30 times each on sixteen models: OpenAI's GPT-5.x and GPT-4o, Claude, Gemini, Grok, Llama, Mistral, and China's DeepSeek, Qwen, Kimi and GLM. Fifteen land in the libertarian-left quadrant.

← all threads