model roundup

Qwen 3.6

12 items · started 2026-06-23 · closed 2026-07-08

  1. I am building a custom skill for a specific searching tool that I use. It uses custom search commands similar to SQL.

  2. Hi HN, I was once given the advice: Don't waste expensive frontier model credits (GPT/Claude/etc.) on bulk work. Send the boring, repetitive, high-volume jobs to a smaller model, and save the expensive prompts for when you actually need fr…

  3. To save cost, I had a local LLM (qwen 3.6 35B) execute some basic maintenance on a project (relaunch a pipeline to resume it at a specific point). When Qwen struggled to do it, but eventually figured it out, I asked it to add to the docume…

  4. If you have an Apple Silicon Mac you can run Claude Code completely locally (and free) by pointing it at a local server. Here's how: Setup (2 minutes) brew install mlx-serve mlx-serve run gemma-4-e4b-it # downloads + starts the server Then…

  5. Running Qwen 3.6 Locally on a Mac Mini M4 with 16GB RAM Two days ago Qwen open-sourced Qwen 3.6-35B-A3B — a 35-billion parameter Mixture of Experts model that only activates 3 billion parameters per token. It's Apache 2.0 licensed, ships w…

  6. Upgrading a lab NAS into one box for VMs + loomcycle server + local LLM inference, on a budget that ruled out the DGX Spark and a Mac Studio. Why AM5 + DIMM DDR5 + the 8700G APU was the only shape that hosted all three workloads without so…

  7. I've been running Qwen3.6-35B-A3B locally on llama.cpp and noticed that prompt processing throughput gets too low with MTP. I got nerd-sniped.

  8. Inference Cards Why When someone says “I run Qwen 3.6 at 25 tokens per second”, or makes any similar performance claim about their self-hosted LLM setup, this is only meaningful if we know several other things. - Which model variant?

  9. is there anything with a 1M context window I can spend 100-200usd a day on that actually works? I don't have 5-10m to wait for claude to think about how to respond to a three word prompt.

  10. I've been running Qwen3.6-35B-A3B locally on llama.cpp and noticed that prompt processing throughput gets too low with MTP. I got nerd-sniped.

  11. 1 qwen3-vl-2b-instruct 25.7M/mo 56.7/mo 2.13B 2025-10-19 8mo ago Qwen/Qwen3-VL-2B-Instruct + 2 more variants 2 qwen3-6-35b-a3b 21.6M/mo 2.2K/mo 36.0B 2026-04-15 2mo ago Qwen/Qwen3.6-35B-A3B + 15 more variants 3 qwen3-6-27b 20.3M/mo 2.1K/mo…

  12. Hi all I'm coming here because I'm a bit desperate. I've a Web and mobile app project and I'm trying to use AI Agent to dev it.

← all threads