model roundup

DeepSeek 4

10 items · started 2026-09-07 · closed 2026-09-16

  1. Deepseek V4 Pro hallucinates A LOT The most out of any frontier model, actually. Interestingly, this comes from a very specific behavioural pattern it has - it tends to confabulate responses to questions it doesn’t know.

  2. DeepSeek-V4 API · time-of-day pricing DeepSeek API pricing status: OFF-PEAK Checking the clock… A live tracker for DeepSeek-V4-Pro and DeepSeek-V4-Flash time-of-day API pricing, where every token costs exactly 2× during peak hours. Read th…

  3. Anthropic's latest PR push for creating a regulatory framework for shutting down access to competing models is likely due to the July release of Kimi K3 and DeepSeek V4. Kimi K3 was launched as a model that on benchmarks looked competitive…

  4. We tried to answer this question by validating one of the most popular tool in this area: RTK (Rust Token Killer). We used Claude Code with Fable 5.0, and OpenCode with DeepSeek V4 Pro 0813 through OpenRouter on Terminal-Bench 2.1 Research…

  5. Autoprompt closes much of the manual coding loop by planning, building, testing, reviewing, and repairing from one prompt. Autoprompt v2 is now out, with support for Claude Code and 10 other coding tools.

  6. DeepSeek V4 Flash on one DGX Spark A pinned Docker recipe for serving 0xSero/deepseek-v4-flash-0731-spark on one NVIDIA DGX Spark with Local Inference Lab's SparkInfer. The validated configuration exposes a 262,144-token model limit and us…

  7. Dear DeepSeek API user, DeepSeek has officially released the V4.1 Flash model on September 10, 2026 (Beijing Time). In the meanwhile, we plan to postpone the discontinuation of the V4 Pro service to 12:00 Beijing Time on September 14, 2026.

  8. Your inference provider sucks (at caching) Originally published as an article on X. I used to believe that if two providers serve the same model, token prices would tell you which one is cheaper.

  9. Multi-head latent attention (MLA) exposes many logical query heads through one packed latent KV stream. This representation is memory efficient, but it removes the physical per-head cache boundary assumed by conventional head-wise reuse.

  10. DeepSeek V4 Flash across 14 providers: cost, speed and caching Where you send a prompt changes what you pay, how long you wait, and whether the call succeeds. One controlled setup: 1k, 10k and 100k input-token targets, each with a 100 or 1…

← all threads