model roundup

Gemma 4

5 items · started 2026-06-29 · closed 2026-07-06

  1. I was auditing my codes using the now back fable 5 and it kept failing to run runtime tests and this the error i got. so is anthropic now using gemma-4-12b-agentic-fable5-composer2.5-v2-3.5x-tau2 to run tests?

  2. Gemma 4 is now significantly faster in Ollama 0.31 on Apple Silicon via multi-token prediction (MTP), powered by MLX. Performance is now up to 90% faster when used with coding agents, as measured using the Aider polyglot benchmark.

  3. HF Realtime Voice Voice chat over WebSocket against a HF speech-to-speech The result is a speech-to-speech experience that feels dramatically more natural. Instead of waiting for an AI to respond, conversations flow with the responsiveness…

  4. Gemma 4 on Cerebras—The Fastest Inference is Now Multimodal Gemma 4 31B is now running at over 1,800 tokens per second on Cerebras Inference. This multimodal model unlocks an entirely new class of applications, from computer use to image-d…

  5. One thing I miss when using local models is the artifact experience from Claude. With Claude, if you ask for a dashboard, chart, diagram, or landing page, you actually get the thing rendered in the chat.

← all threads