There is an adversarial relationship between developers and big model labs. Model labs charged developers higher API prices to subsidize their own agent harness offerings.
model
DeepSeek-V4-Pro
huggingface.co/deepseek-ai/DeepSeek-V4-Pro ↗
78864 downloads2553 likestext-generationtransformers
from the model card
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence Technical Reportšļø Introduction We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models ā DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) ā both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: Hybrid Attention Architecture: We design a hybrid attention mechanism combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to dramatically improve long-context efficiency. In the 1M-token context setting, DeepSeek-V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache compared with DeepSeek-V3.2. Manifold-Constrained Hyper-Connections (mHC): We incorporate mHC to strengthen conventional residual connections, enhancing stability of signal propagation across layers while preserving model expressivity. Muon Optimizer: We employ the Muon optimizer for faster convergence and greater training stability. We pre-train both models on more than 32T diverse and high-quality tokens, followed by a comprehensive post-training pipeline. The post-training features a two-stage paradigm: independent cultivation of domain-specific experts (through SFT and RL with GRPO), followeā¦
discussions
- DeepSeek 4 9 ongoing since 2026-06-22
- DeepSeek 4 20 2026-06-11 – 2026-06-21
- DeepSeek 4 37 2026-05-29 – 2026-06-13
- DeepSeek 4 181 2026-04-22 – 2026-06-01
recent items
Show HN: DeepSeek Flash inverted the economics of agent products (www.rtrvr.ai via hn) We got DeepSeek-V4-Pro serving in 20 seconds (inferize.ai via hn) Inferize is building highly optimized, elastic inference for AI workloads. Ridiculously fast, efficient LLM serving that scales with demand.
Ask HN: How to avoid LLMs struggling with Lisp parens? (news.ycombinator.com) LLMs seem to love certain languages (Python, Bash, etc.), but they all seem to struggle with Lisp (e.g. Racket or Emacs Lisp).
DeepSeek V4 Flash optimized framework and model variants for DGX Spark (github.com via hn) ds4 - Mixed NVFP4 serving of DeepSeek V4 Flash on the NVIDIA Spark family (GB10) ā ļø This GitHub repository is for archival / mirror purposes only. Active development happens at git.kokoham.com/sleepy/ds4-nvfp4-spark.
Microsoft considers DeepSeek as OpenAI costs mount (www.digitimes.com via hn) Microsoft is reportedly considering introducing a fine-tuned version of the Chinese open-source model DeepSeek V4 into its enterprise artificial intelligence (AI) tool Copilot Cowork, as a lower-cost alternative to models from OpenAI and Aā¦
Cheapest way to run Claude Opus 4.8 on a <$30 monthly budget? (www.reddit.com via reddit) Which option gives the most actual Opus 4.8 usage volume: Kiro Pro, Claude Pro or something else? My monthly budget is $30.
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence (arxiv.org) We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) -- both suā¦
GLM 5.2 via Claude Code is the first non-Claude model that feels close to Opus (www.reddit.com via reddit) Iāve been using GLM 5.2 with Claude Code through its Anthropic-compatible API endpoint. Iāve tested it on various projects, including but not limited to database development, backend payment API work, backend and frontend debugging, Laraveā¦
Using Claude Opus as planner + DeepSeek as worker in Claude Code ā anyone solved the single-session routing problem? (www.reddit.com via reddit) I've been running a hybrid planner/worker setup with Claude Code and hit a tricky constraint I'm hoping the community has thoughts on. The setup Planner ā Claude Opus for architecture, planning, and review Worker ā DeepSeek V4 Pro / DeepSeā¦
Native Coding Agent Optimized for Local LLM and DeepSeek v4 with Vector Memory (code.intellios.ai via hn) cwcode A terminal coding agent built around DeepSeek V4 Pro, Qwen3.6ā27B, Kimi, Azure, and anything else that speaks OpenAIās chat API. Written in Go.
DeepSeek V4 Pro at 5% the cost of Claude ā what it takes to close the gap (howardchen.substack.com via hn) DeepSeek V4 Pro at 5% the cost of Claude ā what it takes to close the gap Hash-anchored edits, a sticky prefix cache, and the autonomous loops we run on production code Weāve been using DeepSeek V4 Pro as our daily-driver coding model forā¦
Kimi 2.7 vs. DeepSeek Coder (simpletechguides.com via hn) Kimi K2.7 Code vs MiMo Code vs DeepSeek V4 Pro: Three Open-Source Coding Tools Compared Three Chinese AI labs shipped major coding tools in the same window this spring: Moonshot AI released Kimi K2.7 Code, Xiaomi shipped MiMo Code, and Deeā¦
DeepSeek-V4 Can't Read Images? I Made It Read (www.dataleadsfuture.com via hn) DeepSeek-V4 Can't Read Images? I Made It Read Don't wait for a multimodal model, you can use it now Introduction Have you ever had that frustrating moment: you are coding with deepseek-v4 in OpenCode, your code throws an error, you want toā¦
International Market Retention Strategy After the Fable 5 Export Ban (www.reddit.com via reddit) Like many of you, I lost access to Fable 5 on June 12. The next day, I co-authored a strategy paper with Claude addressing the core business problem: how does Anthropic retain its international market now that cloud-only deployment has beeā¦
Ask HN: Which cheap Chinese LLM are you using? (news.ycombinator.com) In the last one or two months, starting from DeepSeek V4 Pro, there are quite many low-price Chinese models coming out. Their performance looks more or less similar to me: Mimo V2.5 Pro, MiniMax M3, and the just released GLM 5.2, etc.
Fable 5 Max confidently wrong about PDF encryption status (www.reddit.com via reddit) I just ran into a bizarre hallucination with Fable 5 Max regarding file analysis. i uploaded several PDF to Fable 5 Max, and out of two of it claude completely refused to process it, claiming the files was password-protected.
↯ Hallucination↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4hallucinationdeepseek
How can Deepseek v4 top the coding leaderboards and still sit 8 months behind the frontier? (www.reddit.comhttps) Two numbers on this model that don't sit comfortably with each other. The Pro config posts coding scores near the top of every board, 80.6 on SWE-bench Verified and 93.5 on LiveCodeBench.
↯ Swe Bench↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4swe-benchgpt-5deepseek+1
FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention (www.reddit.com via reddit) Conventional LLMs keep the full KV cache loaded during decoding, causing a severe GPU memory bottleneck for ultra-long context serving. In this report, we propose Lookahead Sparse Attention (LSA), a novel inference paradigm powered by a Neā¦
↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4deepseek
Deepseek v4 pro (www.reddit.com via reddit) Hello, Ive ran out of Pro+, is it possible to use DS4 in cursor ide? thanks
Bit of a lull or Winter is Coming? (www.reddit.com via reddit) It feels as though weāre at an inflection point and I was wondering what othersā take is on the current situation: On the frontier end we have OpenAI and Anthropic gearing up for their IPO, so itās all Mythos and wow and it seems plausibleā¦
↯ Mistral↯ Anthropic Mythos↯ DeepSeek 4↯ DeepSeek 4mistralmythosopenai+1
Can I finetune Deepseek V4-flash with two rtx pro 6000s (www.reddit.com via reddit) Well I knew, it may be very tight on 192GB. However, is there any framework to do finetuning of DS4-flash with 4bit QLoRA?
DOA model by Cohere Labs (www.reddit.com via reddit) So apparently the model gets beaten by qwen 3.6 on every benchmark reported by cohere labs. You are getting lower RAM (considering model offload) usage and slightly better performance for imo significantly less output quality.
Running DeepSeek-V4-Flash on a Raspberry Pi (twitter.com via hn) Article Conversation Running DeepSeek-V4-Flash on a Raspberry Pi I ran DeepSeek-V4-Flash on a Raspberry Pi 5 (8GB edition) by streaming model weights from a PCIe attached NVMe SSD. Codex (GPT-5.5 xhigh) and Claude Code (Opus 4.8 max) droveā¦
Here are some tips on hitting nearly 200 tok/s for DeepSeek v4 Flash on Hopper (dnhkng.github.io via reddit) I needed a smarter model for my local Hermes Agent setup, so I moved to DeepSeek v4 Flash. First things first: Running 4 concurrent threads on vLLM, I can hit ~400 tok/s 400 x 60 x 60 x 24 x 30 is ~1B TOKENS per month!!!
Share your agentic LLMs and average cost ($/MTokens) (www.reddit.com via reddit) DStudio ā local DeepSeek V4 with a design studio, reachable from your phone (github.com via hn) DStudio A native, local-first desktop app for DeepSeek V4 ā chat, a coding agent and a design studio, all running on your Mac. Nothing leaves the device.
Mimo v2.5 is better deal than DeepSeek v4 flash (news.ycombinator.com) So Hear me out. Not only on almost all benchmarks is mimo v2.5 is better than dsv4f flash, but also the pricing.
Show HN: One API Key for 45 AI Models ā Pay per Token, OpenAI Compatible (modelhub-api.com via hn) DeepSeek V4 math score equals GPT-5.5 (91) and trails by just 4-6 points in other categories ā at 97% lower cost. Is the AI quality as good as GPT?
DeepSeek V4 Pro beats GPT-5.5 Pro on precision (runtimewire.com via hn) DeepSeek V4 Pro takes this matchup 38.0 to 33.0, and the margin feels earned. Across the scored tasks, the pattern is simple: Model A was tighter, more literal, and more reliable under constraints, while Model B was good but a little too wā¦