I've built a new method for steering LLMs called Semantic Overlays, small trained adapters on a frozen model which change how it perceives a piece of its context. The most readily applicable usage is to mitigate prompt injection, and it le…
model
Qwen3.5-9B
huggingface.co/Qwen/Qwen3.5-9B ↗
5662081 downloads1256 likesimage-text-to-texttransformers
from the model card
Qwen3.5-9B [!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. Qwen3.5 Highlights Qwen3.5 features the following enhancement: Unified Vision-Language Foundation: Early fusion training on multimodal tokens achieves cross-generational parity with Qwen3 and outperforms Qwen3-VL models across reasoning, coding, agents, and visual understanding benchmarks. Efficient Hybrid Architecture: Gated Delta Networks combined with sparse Mixture-of-Experts deliver high-throughput inference with minimal latency and cost overhead. Scalable RL Generalization: Reinforcement learning scaled across million-agent environments with progressively complex task distributions for robust real-world adaptability. Global Linguistic Coverage: Expanded support to 201 languages and dialects, enabling inclusive, worldwide deployment with nuance…
discussions
recent items
Show HN: Semantic Overlays – an NX bit for LLM prompt injection (live demo) (semantic-overlays.vercel.app via hn) First Make It Playable, Then Make It Good: Staged Interaction Learning for Small Dialogue-Game Agents (arxiv.org) We present Qwen-GuidePlay-2B, a 2B-parameter language model for dialogue-game interaction. We fine-tune Qwen3.5-2B using three steps: a) SFT on only successful game trajectories from Playpen, b) weighted turn-level SFT, and c) teacher-guid…
Show HN: Moe-Direct – MoE Models far larger than your RAM, on a consumer desktop (github.com via hn) I wanted to try using the larger models on my computer (32GB RAM, RTX 5080, Gen5 NVMe), but the best I could do was around 30B. So I started with the idea that it might be possible by taking advantage of the fact that MoE models use only s…
RL'd Qwen-3.5-35B to paint hibiscus (twitter.com via hn) had to try this. RL'd Qwen-3.5-35B to paint hibiscus by writing p5.js code, rewarded by rendering the sketch and judging it against my own hand-rated favorites.
What's this Apodex thing? (AMA prep) (www.reddit.com via reddit) A sticky popped up "Hey, AMA today". At first I thought I missed something, but I didn't see a single mention of it here so far, aside from having never heard of it.
Reducing Sycophancy in Qwen and Gemma Using Runtime Activation Steering (www.reddit.com via reddit) So I was testing this technique of runtime steering on tiny versions of Qwen 3.5 and Gemma 4 (2B and 4B). Basically, without changing the weights (like with Heretic/ablation, for example), we steer the model in the opposite direction of a…
Show HN: Sktime-CLI A command line tool to do all time-series tasks (github.com via hn) I am a contributor and part time employee at sktime, a framework for all time series related tasks, but it is collection of large number of estimators which sometimes make it difficult for new commers or people who want to do simple task t…
clean your dust bunnies (www.reddit.comhttps) Qwen 3.5 opus 4.6 distilled said I need to give him some maintenance. Featuring the wolfbox
Best model you can run on a 16gb phone? (www.reddit.com via reddit) Qwen 3.5 9B q6?
Qwen 3.5 4B IQ2_XS: +16.67% Reasoning Performance From Tensor-Level Allocation (www.reddit.com via reddit) I was finally able to replicate tensor level allocation outside the Gemma family. https://huggingface.co/ByteOtter/Qwen3.5-4B-CADA-IQ2_XS After the Gemma 4 12b, e4b and gemma 3 4b results, I attempted to expand into qwen and ran into a few…
Show HN: Building a full agentic harness around a 4B model is hard (orvena.app via hn) Around 3 months ago, we were thinking why none of the iPhone apps running an LLM are built as a full harness (as in inference + agentic loop + context management + tools + MCP servers and etc.). It became more interesting when we noticed e…
Inferencing at 10.33 t/s on Qwen 3.5 35B on a $300 laptop (www.reddit.com) https://preview.redd.it/u8062juegq3h1.png?width=1919&format=png&auto=webp&s=a213f6929c6cad58e92bc1681dac9f0545b04d13 Overview: As the market for consumer computing parts becomes more scarce due to the AI boom, finding ways to use lower-end…
Is a 128 GB MacBook Pro M5 Max actually too slow for large-context local LLM coding workflows? (www.reddit.com) People are warning me about the prompt-processing speed of a MacBook Pro M5 Max with 128 GB RAM. My main concern is prompt ingestion / prefill latency and large-context handling — not raw token generation speed (which I think is OK).
ran qwen3.5 locally on a flight with no wifi. claude code started straight-up hallucinating (www.reddit.com) heavy travel period last month, lots of offline time, and i could not stop building. airplane wifi was unusable so we switched models inside Claude Code and fired up qwen3.5 locally on an M4 macbook.
Stop QwenLLama! Every other 4th post in this sub is about Qwen models in the past month (www.reddit.com) Disclaimer: I use Qwen models on a day to day basis.. You could take it as a rant or even my concern about innovation in other models.
Long-context performance at lower quants (www.reddit.com) I've been using Qwen3.5 122B A10B (Q3_K_XL) a lot lately for coding, and it's been pretty incredible overall like it feels not far off from frontier-level for most tasks -- but I've been noticing that usually once I hit around 75-80k conte…
Harbor v0.4.19 - vllm/sglang/llama.cpp launch codex/claude/pi/opencode (www.reddit.com) I'm usually not posting about Harbor releases out of the respect for the community here, but I think v0.4.19 might save a lot of people some time. Harbor can now launch your local agentic coding tools with local inference backends.
ReAct tool-calling issue: Orchestration model computes internally instead of using tools (www.reddit.com) Built a local ReAct-style calculator agent with 6 tools: add subtract multiply divide modulo etc. The setup is: orchestrator agent dynamic tool selection ReAct loop tools exposed as functions Problem: Even when the user asks multi-step ari…
Server build for local inference. 128 gb 3200 or 256 gb 2133mhz RAM? (www.reddit.com) Hi, I am building a server so that my dual rtx 3090 setup runs at full speed. - asrock romed8 t2 revision 1.3 - epyc 7642 - ddr4 128 gb 3200 or 256 gb 2133 (256 gb is a bit cheaper) 8 channel - dual rtx 3090 - gigabyte psu 1600 w What do y…
Old Mac Pro still proving its worth (www.reddit.com) The “Trash Can” Mac Pro, once the most expensive machine you could buy from Apple, mine was just shy of £10,000 in 2016 — that’s £14k in today’s money. Until recently mine was just running as a kubernetes single node development platform,…
Want Built a React-style looping agent with small LLMs (Qwen 3.5 9B / Gemma4) + LangGraph? (www.reddit.com) Currently experimenting with building a React-style looping agent system using small LLMs like Qwen 3.5 9B and Gemma 4 (E2B), and I wanted to ask if anyone here has worked on something similar. Current setup: Using LangGraph Around 5 tools…
Hermes Agent issues with directory creation (www.reddit.com) I'm having issues with Hermes Agent actually processing commands through the terminal. I'm doing something simple like asking it to make a dir and it tells me it has, but it hasn't.
Is there something wrong with Local LLM ability to read file? (www.reddit.com) So I've been feeding the sub file of anime episodes into Claude/ChatGPT/Deepseek and ask them to find all full name of Japanese character in it and put it into a python array so I can run a script to flip the name back to the original Japa…
qwen 2B model - thinks for 600 tokens on a simple "Hi" (www.reddit.com) Using llama.cpp Model - Q8 - unsloth/Qwen3.5-2B-GGUF Is this expected with tiny models like this one? I am trying tiny models for a since most of the task I have involves searching local files etc and need less of the models own knowledge.
At wits end for optimizing settings in llama.cpp for 100k context (www.reddit.com) Long story short, I am running Qwen3.5-35B-A3B (GGUF format) and other models on MacOS and getting around 1500 tokens/sec for prompt processing and around 35-50 tokens per second for prompt processing. I'm using the latest version of llama…
Show HN: Prism Coder – Qwen3.5-14B fine-tuned for MCP tool-routing decisions (github.com via hn) 🧠 Prism Coder 🌐 Read in your language: 🇬🇧 English · 🇪🇸 Español · 🇫🇷 Français · 🇵🇹 Português · 🇷🇴 Română · 🇺🇦 Українська · 🇷🇺 Русский · 🇩🇪 Deutsch · 🇯🇵 日本語 · 🇰🇷 한국어 · 🇨🇳 中文 · 🇸🇦 العربية Persistent memory + tool-calling intelligence for AI a…
40+tok/s - optimized recipe for Qwen 3.5 122B Int4 on a single DGX Spark with vLLM (www.reddit.com) Hello guys, two days ago i ran the spark-arena for my Qwen 3.5 122B Recipe on a single DGX Spark and I got the highest score on speed for any context length and concurrency across all 3.5 122B Int4 Recipes. Just wanted to share if somebody…
Floor for local meeting summarization on a 6GB GPU: qwen3.5:0.8b works at 57s, Granite 4 350M hallucinates (www.reddit.com) Disclosure: I made this. Open-source, MIT, Windows + Linux.
Full Hermes Agent tutorial (Spanish with English auto-translation). Computer Use, MCP Blender, Hindsight memory and multi-agent setup (www.reddit.com) Spent weeks running Hermes Agent in production on my Mac Mini M4 before recording this. Wanted to show things nobody else was covering.