Deepseek V4 Pro hallucinates A LOT The most out of any frontier model, actually. Interestingly, this comes from a very specific behavioural pattern it has - it tends to confabulate responses to questions it doesn’t know.
model
DeepSeek-V4-Flash-0731
huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731 ↗
4561861 downloads3839 likestext-generationtransformers
from the model card
DeepSeek-V4-Flash-0731 Technical Report👁️ Introduction DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. It has the same model structure as DeepSeek-V4-Flash-DSpark, i.e. it comes with a speculative decoding module attached. DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) on benchmarks listed below despite its far smaller activated parameter count, and is broadly competitive with the strongest proprietary models available. | Benchmark | DeepSeek-V4-Flash-0731 | DeepSeek-V4-Flash (Preview) | DeepSeek-V4-Pro (Preview) | GLM-5.2 | Opus-4.8 | | :--- | :---: | :---: | :---: | :---: | :---: | | Terminal Bench 2.1 | 82.7 | 61.8 | 72.1 | 81.0 | 85.0 | | NL2Repo | 54.2 | 39.4 | 38.5 | 48.9 | 69.7 | | Cybergym | 76.7 | 38.7 | 52.7 | - | 83.1 | | DeepSWE | 54.4 | 7.3 | 12.8 | 46.2 | 58.0 | | Toolathlon-Verified | 70.3 | 49.7 | 55.9 | 59.9 | 76.2 | | Agents' Last Exam | 25.2 | 15.8 | 16.5 | 23.8 | 25.7 | | AutomationBench Public | 25.1 | 10.8 | 12.8 | 12.9 | 27.2 | | DSBench-FullStack † | 68.7 | 37.0 | 41.8 | 61.8 | 71.6 | | DSBench-Hard † | 59.6 | 25.8 | 31.1 | 54.5 | 71.7 | Notes: For the Code Agent tasks among the public benchmarks above, DeepSeek-V4-Flash-0731 is evaluated with the minimal mode of DeepSeek Harness (to be released) as the agent framework, using the max reas…
discussions
- DeepSeek 4 11 2026-09-07 – 2026-09-16
- DeepSeek 4 14 2026-08-30 – 2026-09-05
recent items
Prompt steering can reduce DeepSeek v4 hallucinations to below GPT and Claude (propensitylabs.substack.com via hn) DeepSeek Peak Hours (deepseek-peak-hours.sivaram.dev via hn) DeepSeek-V4 API · time-of-day pricing DeepSeek API pricing status: OFF-PEAK Checking the clock… A live tracker for DeepSeek-V4-Pro and DeepSeek-V4-Flash time-of-day API pricing, where every token costs exactly 2× during peak hours. Read th…
Anthropic is no longer a frontier lab (twitter.com via hn) Anthropic's latest PR push for creating a regulatory framework for shutting down access to competing models is likely due to the July release of Kimi K3 and DeepSeek V4. Kimi K3 was launched as a model that on benchmarks looked competitive…
Do output compression tools still matter with newer AI models? (www.reddit.comhttps) We tried to answer this question by validating one of the most popular tool in this area: RTK (Rust Token Killer). We used Claude Code with Fable 5.0, and OpenCode with DeepSeek V4 Pro 0813 through OpenRouter on Terminal-Bench 2.1 Research…
A Claude Code skill pushed DeepSeek V4 Flash from 67.42% to 82.02% (www.reddit.comhttps) Autoprompt closes much of the manual coding loop by planning, building, testing, reviewing, and repairing from one prompt. Autoprompt v2 is now out, with support for Claude Code and 10 other coding tools.
DeepSeek-v4-flash-0731-spark-sparkinfer: DeepSeek V4 Flash on one DGX Spark (github.com via hn) DeepSeek V4 Flash on one DGX Spark A pinned Docker recipe for serving 0xSero/deepseek-v4-flash-0731-spark on one NVIDIA DGX Spark with Local Inference Lab's SparkInfer. The validated configuration exposes a 262,144-token model limit and us…
DeepSeek v4 Pro discontinuation (email) (news.ycombinator.com) Dear DeepSeek API user, DeepSeek has officially released the V4.1 Flash model on September 10, 2026 (Beijing Time). In the meanwhile, we plan to postpone the discontinuation of the V4 Pro service to 12:00 Beijing Time on September 14, 2026.
DeepSeek V4 Flash Provider Cache Benchmark (olafdsouza.com via hn) Your inference provider sucks (at caching) Originally published as an article on X. I used to believe that if two providers serve the same model, token prices would tell you which one is cheaper.
RedKnot-MLA: Multi-Head Offline-Online Reuse for DeepSeek-V4 Long-Context Serving (arxiv.org) Multi-head latent attention (MLA) exposes many logical query heads through one packed latent KV stream. This representation is memory efficient, but it removes the physical per-head cache boundary assumed by conventional head-wise reuse.
DeepSeek V4 Flash across 14 providers: cost, speed and caching (www.inference.academy via hn) DeepSeek V4 Flash across 14 providers: cost, speed and caching Where you send a prompt changes what you pay, how long you wait, and whether the call succeeds. One controlled setup: 1k, 10k and 100k input-token targets, each with a 100 or 1…
Show HN: Aidcrew a team of coding agents, each on its own model, in one terminal (github.com via hn) aidcrew A team of coding agents, each on its own provider and model, every job in a git worktree of its own, in one terminal. architect claude-opus-5 coder deepseek-v4-flash ◆ reviewer free-tier ▸ Plan in PLAN.md: rotate… ⠹ thinking ▸ The…
Antirez Brings Vision to DS4, DeepSeek V4 Flash Runs Locally on an M5 Max (pasqualepillitteri.it via hn) could not extract summary
Ask HN: Why Groq didn't provide DeepSeek-v4-flash (news.ycombinator.com) could not extract summary
You are an AI Assistant, but what am I? Why I always tell my Agent I'm an expert (shimin.io via hn) Do 2026 models still give worse answers to users they judge less educated? 1000 multiple-choice questions and 30 advice scenarios across Sonnet 5, GPT 5.6 Luna, and DeepSeek V4 Flash — and why you should keep the memory features off.
Claude gets surprisingly hostile when I try asking about other models (www.reddit.com via reddit) I am trying to set up an email triage setup for my needs. After a bit of discovery I started exploring using Hermes Agent as an option with Claude.
An update to my memory system that is long overdue. (www.reddit.com via reddit) Hey everyone, it's been a while since I updated anyone on the memory system I built for my AI assistant Friday. Well, I had Deepseek V4 Flash write up a system map for itself, and I figured that probably people here would be interested in…
↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4deepseek
OpenCode Senses: The most advanced Local Vision Plugin for OpenCode That Actually Understands Images (www.reddit.comhttps) OpenCode Senses can inspect screenshots, extract exact OCR, detect and locate objects, zoom into regions, compare two images, measure colors, crop and annotate images, and even reverse-search them. Everything runs locally, so it's private,…
↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4deepseek
Qwen3.8-Flash-Next better then DeepSeek V4 Pro (www.reddit.com via reddit) could not extract summary
↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4deepseek
Questions on optimism speed/intelligence on this rig (www.reddit.com via reddit) Rig: 3945WX (12C, 2 CCDs, no AVX-512) · 8×32GB DDR4-3200 · 4× 5060 Ti 16GB · PCIe 4.0. Agentic workload (Hermes Agent).
↯ Vllm↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4vllmdeepseekagentic