claude is a token maxxing f*ckboi, what's next? (www.reddit.com via reddit)
model roundup
Qwen 3.6
-
is there anything with a 1M context window I can spend 100-200usd a day on that actually works? I don't have 5-10m to wait for claude to think about how to respond to a three word prompt.
-
I was curious why MTP affects PP TPS in llama.cpp. My PoC recovers it? (news.ycombinator.com)
I've been running Qwen3.6-35B-A3B locally on llama.cpp and noticed that prompt processing throughput gets too low with MTP. I got nerd-sniped.
-
Show HN: A leaderboard of most popular open-weight LLMs (osolmaz-leaderboard.hf.space via hn)
1 qwen3-vl-2b-instruct 25.7M/mo 56.7/mo 2.13B 2025-10-19 8mo ago Qwen/Qwen3-VL-2B-Instruct + 2 more variants 2 qwen3-6-35b-a3b 21.6M/mo 2.2K/mo 36.0B 2026-04-15 2mo ago Qwen/Qwen3.6-35B-A3B + 15 more variants 3 qwen3-6-27b 20.3M/mo 2.1K/mo…
-
I keep restarting vibe dev in CC! (www.reddit.com via reddit)
Hi all I'm coming here because I'm a bit desperate. I've a Web and mobile app project and I'm trying to use AI Agent to dev it.