model roundup

Qwen 3.5

7 items · started 2026-08-21 · closed 2026-08-29

  1. A sticky popped up "Hey, AMA today". At first I thought I missed something, but I didn't see a single mention of it here so far, aside from having never heard of it.

  2. So I was testing this technique of runtime steering on tiny versions of Qwen 3.5 and Gemma 4 (2B and 4B). Basically, without changing the weights (like with Heretic/ablation, for example), we steer the model in the opposite direction of a…

  3. I am a contributor and part time employee at sktime, a framework for all time series related tasks, but it is collection of large number of estimators which sometimes make it difficult for new commers or people who want to do simple task t…

  4. Qwen 3.5 opus 4.6 distilled said I need to give him some maintenance. Featuring the wolfbox

  5. Qwen 3.5 9B q6?

  6. I was finally able to replicate tensor level allocation outside the Gemma family. https://huggingface.co/ByteOtter/Qwen3.5-4B-CADA-IQ2_XS After the Gemma 4 12b, e4b and gemma 3 4b results, I attempted to expand into qwen and ran into a few…

  7. Around 3 months ago, we were thinking why none of the iPhone apps running an LLM are built as a full harness (as in inference + agentic loop + context management + tools + MCP servers and etc.). It became more interesting when we noticed e…

← all threads