model roundup

Qwen 2.5

3 items · started 2026-09-16 · ongoing (last activity 2026-09-18)

  1. monkeyDcode The coding agent that makes the model you already have actually reliable — not "beat GPT with a 7B," but make qwen2.5-coder:7b on your laptop produce work you can merge, every time, not on lucky rolls. Quick start · Why this ex…

  2. Who the judge is can affect an LLM-as-judge result, but measuring that effect without confusing it with candidate quality is difficult. We study four open-weight families (Llama 3.1, Qwen 2.5, Gemma 2, and Yi 1.5) in a fully crossed pairwi…

  3. Parallel Constrained Decoding for Apple Silicon A high-throughput inference engine for structured information extraction, decision routing, and categorical classification on Apple Silicon using MLX. Parallel Constrained Decoding evaluates…

← all threads