model roundup

GLM 5.3

2 items · started 2026-09-06 · closed 2026-09-09

  1. 10-task GLM 5.3 harness bench: claude, opencode, pi, zcode, hermes and 3code I'm performing a series of harness benchmarks on the same 10 SWE-bench verified tasks representatively chosen for difficulty. This is far from a perfect measure a…

  2. I am working on a big project with Claude Code only context7 mcp added no others tools. With opus 5 is all ok it seems to remember what we have done days before follow the repo conventions etc.

← all threads