model roundup

DeepSeek 4.1

22 items · started 2026-09-08 · ongoing (last activity 2026-09-19)

  1. could not extract summary

  2. Recently saw Supermicro at Super Compute 2025 Overview and thought to share my build. I got a 4x RTX6000 Blackwell Max-Q node in place that I built out from my original rtx6000 ada workstation from 2024.

  3. The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV c…

  4. Inference Is the Last LLM Moat OpenAI and Anthropic's moat for developers is subsidized inference, meaning reliable, fast, Western-hosted access to running models at an affordable monthly price.… When I run out of Codex and Claude Code usa…

  5. DeepSeek’s 11/11 result showed why advanced agent benchmarks need to check both the outcome and the attack path: our audit confirmed six planned exploits and found five unexpected routes.

  6. I'm doing web developement, and game development for a hobby project. I've tried lots of harnesses / IDE's - Best I've found is VSCodium.+ Cline + Openrouter, using discounted models (GLM 5.3 Flash is 50% off atm for example) I used Cursor…

  7. SGLang and Miles Add Day-0 Support for DeepSeek-V4.1 1. Architecture overview DeepSeek-V4.1 introduces several architecture choices that shape the serving stack.

  8. ARGODRIVE Layout, balancer and instruments for running mixture-of-experts models from SSDs. A 518 GB model on a 128 GB laptop: every token waits on disk, so what matters is not how much bandwidth you own but how long the slowest required r…

  9. Quite a niche project that shows in a neat way historical wind force and orientation, anywhere on earth. No particular use case in mind, apart from checking that my next house will not be affected by the Mistral wind (the horrendous wind d…

  10. cognition-claude-proxy A local proxy that lets Claude Code (or any Anthropic-API client) use Devin's model catalog — SWE-2, GLM-5.2, DeepSeek V4.1 Flash, and 200+ others — as its backend. It translates the Anthropic Messages API to the Con…

  11. For months I've been using a CC/Codex plugin (agent-talk) to use GPT as an adversarial reviewer for Claude. Then I recently switched to herdr -- very cool terminal session server/multiplexer.

  12. Antirez just landed a commit[1] to the github repo for DwarfStar that adds support for his heavily quantized variant of DeepSeek V4.1 Flash. He claims it runs with SSD streaming on a macbook pro M5-Max 128 GB laptop.

  13. We are late to this but better than never. Have been busy finalizing the second AIE NYC, which is happening in one month.

  14. FP4 Brain on X: "got deepseek V4.1 flash running locally on a 16GB m1 mac mini original FP4/FP8 weights, ssd streaming + custom mlx runner 108s ttft and about 23s/token (not to be confused with tok/s)" got deepseek V4.1 flash running local…

  15. Anthropic, you better get your shit together

  16. Subscription vs API vs GPUs One fully used ChatGPT Pro 20x seat implies a monthly token volume. This calculator prices that same volume five ways: the subscription sticker, the GPT-5.6 Sol API, the DeepSeek-V4.1-Flash API, GPUs you buy, an…

  17. DeepSeek V4.1 Flash (Reasoning, Max Effort) Intelligence, Performance & Price Analysis Model summary DeepSeek V4.1 Flash (Reasoning, Max Effort) is amongst the leading models in intelligence and reasonably priced when comparing to other op…

  18. - Introducing the smallest model in our new architecture family, with native visual understanding. - Designed for greater capability, faster inference, higher throughput, and scaling to larger models.

  19. DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Technical Report 👁️ Introduction We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to…

  20. DeepSeek V4.1 Flash Beta An experimental beta version of DeepSeek V4.1 Flash that expires on September 10th 2026. View API reference- Input and output price - Input $0.22, Output $0.66, Per 1M tokens - 24h uptime - Loading AI Gateway uptim…

← all threads