Ask HN: DeepSeek v4.1 Flash on Local hardware, what tok/s do you see? (news.ycombinator.com)
model roundup
DeepSeek 4.1
-
could not extract summary
-
DeepSeek v4.1 Flash avg 102 tps on 4x RTX6000 pro max-q, 2.1x up from v4-flash (forum.level1techs.com via hn)
Recently saw Supermicro at Super Compute 2025 Overview and thought to share my build. I got a 4x RTX6000 Blackwell Max-Q node in place that I built out from my original rtx6000 ada workstation from 2024.
-
The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV c…
-
Setting Up Pi with DeepSeek v4.1 Flash on OpenRouter (www.vincentschmalbach.com via hn)
Inference Is the Last LLM Moat OpenAI and Anthropic's moat for developers is subsidized inference, meaning reliable, fast, Western-hosted access to running models at an affordable monthly price.… When I run out of Codex and Claude Code usa…
-
DeepSeek v4.1 Flash Is Now Our Best Hacking Model (enclave.ai via hn)
DeepSeek’s 11/11 result showed why advanced agent benchmarks need to check both the outcome and the attack path: our audit confirmed six planned exploits and found five unexpected routes.
-
Ask HN: What's the most economical approach to the most tokens? (news.ycombinator.com)
I'm doing web developement, and game development for a hobby project. I've tried lots of harnesses / IDE's - Best I've found is VSCodium.+ Cline + Openrouter, using discounted models (GLM 5.3 Flash is 50% off atm for example) I used Cursor…
-
SGLang and Miles Add Day-0 Support for DeepSeek-v4.1 (www.lmsys.org via hn)
SGLang and Miles Add Day-0 Support for DeepSeek-V4.1 1. Architecture overview DeepSeek-V4.1 introduces several architecture choices that shape the serving stack.
-
ARGODRIVE Layout, balancer and instruments for running mixture-of-experts models from SSDs. A 518 GB model on a 128 GB laptop: every token waits on disk, so what matters is not how much bandwidth you own but how long the slowest required r…
-
Show HN: Local historical wind and gust visualizer (tamtamhero.github.io via hn)
Quite a niche project that shows in a neat way historical wind force and orientation, anywhere on earth. No particular use case in mind, apart from checking that my next house will not be affected by the Mistral wind (the horrendous wind d…
-
cognition-claude-proxy A local proxy that lets Claude Code (or any Anthropic-API client) use Devin's model catalog — SWE-2, GLM-5.2, DeepSeek V4.1 Flash, and 200+ others — as its backend. It translates the Anthropic Messages API to the Con…
-
I wanted my chats to talk to my chats so I made a chat (macOS) (www.reddit.com via reddit)
For months I've been using a CC/Codex plugin (agent-talk) to use GPT as an adversarial reviewer for Claude. Then I recently switched to herdr -- very cool terminal session server/multiplexer.
-
Dwarf Star Support for DeepSeek 4.1 Flash on MBPro 128GB (news.ycombinator.com)
Antirez just landed a commit[1] to the github repo for DwarfStar that adds support for his heavily quantized variant of DeepSeek V4.1 Flash. He claims it runs with SSD streaming on a macbook pro M5-Max 128 GB laptop.
-
We are late to this but better than never. Have been busy finalizing the second AIE NYC, which is happening in one month.
-
DeepSeek v4.1 flash runs 23 seconds/token on a 2020 16gb M1 Mac Mini (twitter.com via hn)
FP4 Brain on X: "got deepseek V4.1 flash running locally on a 16GB m1 mac mini original FP4/FP8 weights, ssd streaming + custom mlx runner 108s ttft and about 23s/token (not to be confused with tok/s)" got deepseek V4.1 flash running local…
-
DeepSeek v4.1 does the same coding task for $.04 while Fable 5.1 costs $3.64 - LiveBench (www.reddit.comhttps)
Anthropic, you better get your shit together
-
DeepSeek v4.1 Flash vs. ChatGPT Pro Subscription (aicharts.io via hn)
Subscription vs API vs GPUs One fully used ChatGPT Pro 20x seat implies a monthly token volume. This calculator prices that same volume five ways: the subscription sticker, the GPT-5.6 Sol API, the DeepSeek-V4.1-Flash API, GPUs you buy, an…
-
DeepSeek v4.1 Flash – Artificial Analysis (artificialanalysis.ai via hn)
DeepSeek V4.1 Flash (Reasoning, Max Effort) Intelligence, Performance & Price Analysis Model summary DeepSeek V4.1 Flash (Reasoning, Max Effort) is amongst the leading models in intelligence and reasonably priced when comparing to other op…
- DeepSeek v4.1 Flash Uncensored (huggingface.co)
-
DeepSeek-v4.1-Flash: smarter, faster, more efficient (www.deepseek.com via hn)
- Introducing the smallest model in our new architecture family, with native visual understanding. - Designed for greater capability, faster inference, higher throughput, and scaling to larger models.
-
DeepSeek-v4.1-Exp (huggingface.co via hn)
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Technical Report 👁️ Introduction We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to…
-
DeepSeek v4.1 Flash Beta on Vercel (vercel.com via hn)
DeepSeek V4.1 Flash Beta An experimental beta version of DeepSeek V4.1 Flash that expires on September 10th 2026. View API reference- Input and output price - Input $0.22, Output $0.66, Per 1M tokens - 24h uptime - Loading AI Gateway uptim…
- DeepSeek v4.1 Flash is now available for internal beta testing (news.ycombinator.com)