could not extract summary
#deepseek
817 items
Deepseek v4 people (www.reddit.com) Deepseek V4 AGI comfirmed (www.reddit.com) could not extract summary
DeepSeek V4 has released (www.reddit.com) HuggingFace: https://huggingface.co/collections/deepseek-ai/deepseek-v4
Deepseek V4 Flash and Non-Flash Out on HuggingFace (www.reddit.com) https://huggingface.co/collections/deepseek-ai/deepseek-v4
DeepSeek V4 Benchmarks! (www.reddit.com) could not extract summary
DeepSeek-v4 has a comical 384K max output capability (www.reddit.com) was shocked when saw that spec, immediatly went to the website and asked it to make a comprehensive single-html-web-OS and it indeed generated a single 100KB html for me...I'm speechless. https://preview.redd.it/6zcbzbkvj3xg1.png?width=287…
2x 512gb ram M3 Ultra mac studios (www.reddit.com) DeepSeek confirms Huawei-based V4 inference: "After the 950 supernodes are launched at scale in the second half of this year, the price of Pro is expected to be reduced significantly." (www.reddit.com) could not extract summary
Buried lede: Deepseek v4 Flash is incredibly inexpensive from the official API for its weight category (www.reddit.com) could not extract summary
DeepSeek open-sources inference optimizations with 60–85% faster generation [pdf] (github.com via hn) DeepSpec DeepSpec is a full-stack codebase for training and evaluating draft models for speculative decoding. It contains data preparation utilities, draft model implementations, training code, and evaluation scripts.
I'm glad we have deepseek (www.reddit.com) other companies are slowly going away from open weight, not releasing base models, delaying open weight distribution, not releasing top models (this one I think is fair, but still), and I also noticed they stopped publishing research (old…
So... has anyone actually figured out whose model Elephant Alpha is yet? (www.reddit.com) Why has ChatGPT become so annoying and disagreeable? (www.reddit.com) Something I’ve noticed is before the new model, people complained that ChatGPT was “too agreeable” and would glaze you for anything. But now I’ve noticed that it’s the complete opposite and it looks like ChatGPT is disagreeing just to disa…
DeepSeek is pushing forward with $10.29 billion financing round, with Liang Wenfeng committing to continue developing open-source AI models rather than pursuing short-term commercialization goals (www.reddit.com) https://www.bloomberg.com/news/articles/2026-05-22/deepseek-founder-declares-agi-goal-as-10-billion-round-advances
DeepSeek Announces Permanent Price Cut of 75% after Promotion Period (www.reddit.com) could not extract summary
Tested Deepseek v4 flash with some large code change evals. It absolutely kills with too use accuracy! (www.reddit.com) Did some test tasks with v4 flash. The context management, tool use accuracy and thinking traces all looked excellent.
Takeaways & discussion about the DeepSeek V4 architecture (www.reddit.com) Spent the morning looking at the V4 tech report. The benchmarks are getting deserved attention, but I think the architecture is also worth digging into.
No Multimodality yet in DeepSeek-V4. But I'll wait. (www.reddit.com) I hope they include it in their next v4 release. Source: DeepSeek_V4_Technical_Report
DeepSeek-V4 Drops: Open-Source Push Toward Cheaper, Long-Context AI. (www.reddit.com) source : https://x.com/pankajkumar_dev/status/2047552208175354229?s=20
Recent Open models from last 6 Months - Nov 2025 - Apr 2026 (www.reddit.com) I created this chart with recent open models from last 6 months. Few might be older than that possibly.
Price wars begin. MiMo 2.5 Pro now costs the same as DeepSeek V4 Pro (www.reddit.com) could not extract summary
DeepSeek Updated their repo DeepGEMM testing Mega MoE (www.reddit.com) https://github.com/deepseek-ai/DeepGEMM/pull/304 https://preview.redd.it/vcmqwmvzijvg1.png?width=1014&format=png&auto=webp&s=76b1739925f0699b0763aa7814614dd40329c41e https://github.com/deepseek-ai/DeepGEMM/commit/a050d09461e86eb6bba35a8c74…
DeepSeek launching v4.1 flash cheaper and more capable than v4 pro (news.ycombinator.com) DSeek plans to officially release the V4.1 Flash model around September 10, 2026 (Beijing Time). After extensive internal and external testing, V4.1 Flash has comprehensively surpassed V4 Pro across all key metrics, including performance,…
Deepseek V4 Pro is 15x cost to run Artificial Analysis bench from V3.2, higher than Gemini 3.1 Pro (www.reddit.com) Major performance jump though. Worth it?
DeepSeek V4 Pro beats GPT-5.5 Pro on precision (runtimewire.com via hn) DeepSeek V4 Pro takes this matchup 38.0 to 33.0, and the margin feels earned. Across the scored tasks, the pattern is simple: Model A was tighter, more literal, and more reliable under constraints, while Model B was good but a little too w…
DeepSeek V4 Pro underwhelms on Arena (crowdsourced user preference benchmark, not a capability benchmark) (www.reddit.com) could not extract summary
Deepseek V4 flash (high) rivals Gemini 3 flash at 1/5th the cost (www.reddit.com) could not extract summary
coding is basically solved for the boring 90% of tasks (www.reddit.com) just mass refactored a 120 file FastAPI service. 400 steps, 2M tokens, $3 total, zero human input.
Gen AI web traffic share update Main takeaways: → Claude and Gemini continue to grow. → ChatGPT moves closer to the 50% mark. (www.reddit.com) 12 months ago: ChatGPT: 77.6% Gemini: 7.27% DeepSeek: 6.01% Grok: 3.17% Perplexity: 1.75% Copilot: 1.56% Claude: 1.37% 🗓️ 6 months ago: ChatGPT: 69.5% Gemini: 15.9% DeepSeek: 4.06% Grok: 3.31% Perplexity: 2.22% Claude: 2.12% Copilot: 1.97%…
DeepSeek V4 Pro at 75% off until 31 May (api-docs.deepseek.com via hn) Models & Pricing The prices listed below are in units of per 1M tokens. A token, the smallest unit of text that the model recognizes, can be a word, a number, or even a punctuation mark.
Deepseek flash seems like a very good replacement for Haiku at the very least (www.reddit.com) We have a chat system which we use haiku for because it is mostly about tool calling and summarisation of them. But we have many tools with pretty complex input schemas, and stuff like gemma didn't cut it, so we went with haiku.
Tencent, Alibaba in Talks to Invest in DeepSeek at $20 Billion-Plus Valuation (www.reddit.com) https://www.reuters.com/world/asia-pacific/tencent-alibaba-talks-invest-deepseek-information-reports-2026-04-22/
GPT-5.5 improves over GPT-5.4 and overtakes Opus 4.6 to take the 2nd place behind Gemini 3.1 Pro on the Extended NYT Connections Benchmark (www.reddit.com) GPT-5.5: xhigh: 94.0→97.5 high: 93.6→96.9 medium: 92.0→95.0 no reasoning: 32.8→37.5 Kimi K2.6 improves over Kimi K2.5 (78.3→91.4) and becomes the #1 open weights model. DeepSeek V4 Pro improves over DeepSeek V3.2 (50.2→75.7).
Guys we have to change the pelican test (www.reddit.com) So i have been seeing more of those pelican on a bike svg tests and while they work i feel like (and maybe you guys do too) they are getting kinda benchmaxxed so we should switch things up soon and this is my idea generate me a html svg of…
Open Reproduction of DeepSeek-R1 (github.com via hn) Open R1 A fully open reproduction of DeepSeek-R1. This repo is a work in progress, let's build it together!
Decreased Intelligence Density in DeepSeek V4 Pro (www.reddit.com) In the V3.2 paper, they mentioned: Second, token efficiency remains a challenge; DeepSeek-V3.2 typically requires longer generation trajectories (i.e., more tokens) to match the output quality of models like Gemini 3.0-Pro. Future work wil…
Sure, Deepseek… sure. (www.reddit.com) could not extract summary
DeepSeek reasonix, DeepSeek native coding agent with high caching and low cost (esengine.github.io via hn) Open-source AI coding agent for your terminal. Engineered around DeepSeek
Got MTP + TurboQuant running — Qwen3.6-27B -- 80+ t/s at 262K context on a single RTX 4090 (www.reddit.com) So I've been messing around trying to get MTP working alongside TBQ4_0 (TurboQuant's lossless 4.25 bpv KV cache) on Qwen3.6-27B for my own use. So after a day of vibecoding I think I may have gotten something viable.
We benchmarked TranslateGemma-12b against 5 frontier LLMs on subtitle translation - it won across the board, with one significant catch (www.reddit.com) As part of our ongoing translation quality research at Alconost, we put six models through subtitle translation into six language pairs. At first glance the numbers told a clean story.
I catalogued every way local models break JSON output and built a repair library, here's what I found across 288 model calls (www.reddit.com) I've been running structured output prompts through a bunch of models on OpenRouter for the past few months — Llama 3, Mistral, Command R, DeepSeek, Qwen, and every other model on OpenRouter — alongside the usual closed-source suspects. 28…
DeepSeek Targets $50B Valuation in First Fundraising, Escalating Global AI Race (www.financership.com via reddit) Chinese artificial intelligence startup DeepSeek is preparing for its first-ever external fundraising round, and the numbers being discussed signal a dramatic shift in both its strategy and its global standing. The company could be valued…
Reports suggest DeepSeek is seeking $7.35 billion in funding and plans to release its V4.1 update next month. (www.reddit.com) DeepSeek Reportedly Seeking to Raise Over RMB 50 Billion ($7.35 Billion), Accelerating Its Commercialization and Monetization Strategy According to two people familiar with the matter, DeepSeek founder and CEO Liang Wenfeng plans to contri…
DeepSeek V4 Flash on a Single AMD MI300X (github.com via hn) DeepSeek V4 Flash on a single AMD MI300X This repository contains the configuration and patches I use to run deepseek-ai/DeepSeek-V4-Flash-0731 on one AMD MI300X in production. It includes the Docker Compose stack, SHA-256-pinned file over…
Most of my Claude usage was on work that didn't need Claude. Cut my bill 60x on bulk tasks with a tiny side model. (www.reddit.com) I looked at what was actually eating my Claude usage and it was embarrassing. Classifying files.
All major LLMs are lib-left. Even Grok, half the time (unslop.run via hn) I ran the 62-item politicalcompass.org test 30 times each on sixteen models: OpenAI's GPT-5.x and GPT-4o, Claude, Gemini, Grok, Llama, Mistral, and China's DeepSeek, Qwen, Kimi and GLM. Fifteen land in the libertarian-left quadrant.
Notes on DeepSeek (twitter.com via hn) Notes on DeepSeek: We visited the company HQ last Tuesday. It was founded in 2023 by Liang Wenfeng and operated out of his hedge fund, High-Flyer, until somewhat recently.
DeepSeek V4 isn't beating Opus, but it doesn't need to (www.reddit.com) DeepSeek V4 is not in the same league as GPT-5.5 or Opus 4.7. Benchmarks put it slightly below both of those, roughly on par with Opus 4.6.
DeepSeek planning to significantly raise prices (platform.deepseek.com via hn) could not extract summary
DeepSeek has began grayscale testing for DeepSeek with Vision (www.reddit.com) could not extract summary
DeepSeek peak/off-peak pricing update (api-docs.deepseek.com via hn) DeepSeek-V4-Pro GA Release We’re launching DeepSeek-V4-Pro today! 🚀 🔷 Major Agent upgrades with strong production gains!
DeepSeek-V4 Technical Report [pdf] (huggingface.co via hn) DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence Technical Report👁️ Introduction We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro wi…
DeepSeek V4 Update (www.reddit.com) DeepSeek V4 Update
Xiaomi has released a MiMo V2.5 Pro model. It's apparently about as good as Deepseek V4 (but at different tasks) but is significantly cheaper. (x.com via reddit) Xiaomi’s MiMo V2.5 Pro has landed at 54 in the Artificial Analysis Intelligence Index, tied with Moonshot’s Kimi K2.6 - the current top open weights model. MiMo V2.5 Pro’s weights are expected to be released soon, which would make MiMo V2.…
Top open weight models like ds v4 pro max are still like 6-7 months if not more behind closed lab models (www.reddit.com) The best open weight and/or non -American models like Deepseek v4 pro max and kimi k2.6 are still like 3-7 months if not more behind closed lab models .. From ds's technical report- P5-"Nevertheless, its performance falls marginally short…
DeepSeek V4 Pro 0813 quietly released (api-docs.deepseek.com via hn) Using the Responses API To meet the demand for Codex, our API now supports the Responses API format, with the base_url being https://api.deepseek.com . With a simple configuration, you can use DeepSeek models in Codex.
DeepSeek Introduces Vision (chat.deepseek.com via hn) Cookie Settings We use cookies to provide and improve services and ensure security. Click to view our cookie policy.
[P] Built GPT-2, Llama 3, and DeepSeek from scratch in PyTorch - open source code + book (www.reddit.com) I wrote a book that implements modern LLM architectures from scratch. The part most relevant to this sub: Chapter 3 takes GPT-2 and swaps exactly 4 things to get Llama 3.2-3B: LayerNorm → RMSNorm Learned positional encodings → RoPE GELU →…
DeepSeek: Reverse Engineering an AI Assistant by Interviewing Itself (manish.sh via hn) Inside DeepSeek: Reverse Engineering an AI Assistant by Interviewing Itself ByManish ShahiSoftware Engineer • AI Developer Table of contents(30) I run manish.sh. I write about AI tools and how LLMs behave when you push them.
I cut my AI API costs 99% by switching from Claude to DeepSeek (twitter.com via hn) Don’t miss what’s happening People on X are the first to know. Post Conversation AgentDB cost $200+/mo.
New LLM Position Bias Benchmark: does an LLM keep the same judgment when you swap the answer order? Judge models compare two lightly edited versions of the same story twice, with the order swapped. The median model flips in 45% of decisive case pairs. GPT-5.4 is worst at 66%. (www.reddit.com) More info, including charts, per-case metrics, raw judge outputs, and the parsed answer dump: https://github.com/lechmazur/position_bias This benchmark isolates one basic and frustrating failure mode. The model-average first-shown pick rat…
Tencent Hy 30B/7B/1.8B (www.reddit.com) from tencent: Hy-MT2 is a family of “fast-thinking” multilingual translation models designed for complex real-world scenarios. It includes three model sizes: 1.8B, 7B, and 30B-A3B (MoE), all of which support translation among 33 languages…
DeepSeek-V4-Flash W4A16+FP8 with MTP self-speculation: 85 tok/s @ 524k on 2× RTX PRO 6000 Max-Q (www.reddit.com) TL;DR: DeepSeek-V4-Flash running at 85.52 tok/s @ 524k ctx and ~111 tok/s @ 128k single-stream on 2× RTX PRO 6000 Max-Q pasta-paul's DeepSeek-V4-Flash-W4A16-FP8 quant is great, but its MTP head silently gets stripped at load time (HF trans…
Ask HN: What was the last task where only a frontier model could do it? (news.ycombinator.com) ive been seeing a recurring claim that open (weight) models 6 months behind the frontier are good enough for the majority of ‘work’. if you've had a concrete task in the last month where GLM/DeepSeek/Kimi/Qwen failed and Opus/Fable/GPT suc…
DeepSeek-V4-Flash-Vision-Exp Is Now Live on the DeepSeek API Platform (twitter.com via hn) DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge.
Qwen3.6 huge quality gain from Q4 to Q6 for coding agent (www.reddit.com) So, last week I tried to update my unused local LLM setup. I had to stop using it because quality was too low and deepseek was too cheap.
DeepSeek v4 - Subjective vibes (www.reddit.com) I must say Iam kinda torn what to think about those models. At one hand they "ace" some questions on other sometime they behave genuinely weird.
Budget to run Deepseek V4 locally at FP4 precision (www.reddit.com) Just a question for fun/curiosity: in your opinion, if I had enough money, how much would be needed and what configuration would be required to run DeepSeek v4? Maybe not necessarily everything in VRAM, maybe something hybrid.
A Chinese LLM attacked our lab, so we made it work for us (jesta.ai via hn) The model behind the attack was DeepSeek, deepseek-v4-flash-free, running on the free tier. An autonomous AI broke into our lab and worked it for five days, and it left its own name in a script.
DeepSeek seeks $300M in first outside funding at $10B valuation (cryptobriefing.com via reddit) Photo: Dado Ruvic DeepSeek seeks $300M in first outside funding at $10B valuation The Chinese AI startup is reportedly targeting at least $300 million after relying solely on funding from its hedge fund parent until now. DeepSeek is seekin…
DeepSeek-v4.1-Exp (huggingface.co via hn) DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Technical Report 👁️ Introduction We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to…
DeepSeek costs OpenCode Go user $1.14/day; dual DGX breaks even in 24 years (twitter.com via hn) the average OpenCode Go user spent $1.14 per day on deepseek flash v4 this past week the dual DGX setup people are running to do the same costs $10,000 it takes 24 years to break even at 10x the usage it takes 2.4 years - “But my privacy!!…
Why there isn't any top LLM providers investing on diffusion LLM? (www.reddit.com) A year ago, I would’ve said Diffusion LLMs were an interesting idea but still far from practical. They’re still pretty rough, but Mercury 2 now makes it seem like they might finally be getting close to usable.
DS4, a specialized inference engine for DeepSeek v4 Flash (twitter.com via hn) antirez @antirez Welcome to DS4, a specialized inference engine for DeepSeek v4 Flash. github.com/antirez/ds4 This project would have been impossible without the existence of llama.cpp and GGML and the work of @ggerganov and all the other…
Deepseek v4 pricing is genuinely silly, did the math and now i am questioning my entire stack (www.reddit.com) Hey 👋 Saw the tweet making the rounds about deepseek v4 being 35x cheaper than opus on input and 178x cheaper on cached tokens, and was sure it was hyperbole. Pulled the numbers anyway because i had nothing better to do.
I am not sure if I should be proud or not. (www.reddit.com) I managed to get working 4 sub-agents Qwen3.6 35b on dual rtx 3090, I am using deepseek as orchestrator. https://preview.redd.it/biksbgq0n81h1.png?width=783&format=png&auto=webp&s=cf8a4481c1ac439c3283925001c12841b8e6c2e7 They all working l…
DeepSeek nears $45bn valuation as China’s ‘Big Fund’ leads investment talks (www.reddit.com) From: TechCrunch: DeepSeek could hit $45B valuation from its first investment round: https://techcrunch.com/2026/05/06/deepseek-could-hit-45b-valuation-from-its-first-investment-round/ FT (paywall): DeepSeek nears $45bn valuation as China’…
DeepClaude – Claude Code agent loop with DeepSeek V4 Pro, 17x cheaper (github.com via hn) deepclaude Use Claude Code's autonomous agent loop with DeepSeek V4 Pro, OpenRouter, or any Anthropic-compatible backend. Same UX, 17x cheaper.
Single question llm comparison (www.reddit.com) Anthropic is no longer a frontier lab (twitter.com via hn) Anthropic's latest PR push for creating a regulatory framework for shutting down access to competing models is likely due to the July release of Kimi K3 and DeepSeek V4. Kimi K3 was launched as a model that on benchmarks looked competitive…
A verification loop 4x'd DeepSeek's intelligence, matching Opus at 1/7 the cost (ironbee.medium.com via hn) What a Verification Loop Adds to a Coding Agent: A First Look This is the opening post in an ongoing series. We start with one model pair on one project, and the analysis will continue across more models and more datasets.
CAISI releases evaluation report: DeepSeek V4 becomes the most powerful model in China, but still lags about 8 months behind the US frontier (www.reddit.com) https://preview.redd.it/pz8qeln0auyg1.png?width=1400&format=png&auto=webp&s=00ee5218734cfae4783d702411d63e3a4c6bbc60 https://preview.redd.it/hem9mad5auyg1.png?width=1184&format=png&auto=webp&s=2a26fec2b49204e64b44a78b30902ab80f7df53c https…
The exact KV cache usage of DeepSeek V4 (www.reddit.com) Figure 1 of DSV4 paper seems to imply that DSV3.2 uses ~50GB at 1m context and DSV4 uses ~5GB: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf From my own calculations, the correct FP16 KV cache at 1m context s…
Show HN: Telem – Route agent web search across providers and inspect the traces (telem.ai via hn) TL;DR: Web search open router that routes agent web search across providers (Exa, Parallel, Tavily, Brave, SerpAPI etc.), and traces web search results with quality metrics, so you can visualize whether a bad agent run is a web search prob…
↯ GPT 5.5↯ GPT 5.5↯ GPT 5.5↯ GPT 5.5↯ GPT 5.5↯ GPT 5.5↯ GPT 5.5deepseekqwen
DeepSeek V4 Pro at 5% the cost of Claude – what it takes to close the gap (howardchen.substack.com via hn) DeepSeek V4 Pro at 5% the cost of Claude — what it takes to close the gap Hash-anchored edits, a sticky prefix cache, and the autonomous loops we run on production code We’ve been using DeepSeek V4 Pro as our daily-driver coding model for…
SWE-rebench Leaderboard (March, April and May 2026): GPT-5.5, Opus 4.7, Cursor (Composer 2.5), Kimi K2.6 and More (swe-rebench.com via reddit) Hi all, Sorry for going missing — we’ve been collecting a larger, higher-quality set of more complex tasks. We’re excited to share a major leaderboard update covering the past three months.
Can a 5090 with qwen3.6 achieve > 3,000 tok/s ? bring your pitchforks (open-dllm) (www.reddit.com) so background - these people. Fred Zhangzhi Peng, Shuibai Zhang, Alex Tong, worked on converting AR -> diffusion (its already working from older models).
High VRAM local coding model — still Qwen 3.6 27B? (www.reddit.com) I’ve been using Qwen 3.6 27B and it’s amazing. Not exactly your Opus replacement, but great for small tasks and checking work.
DeepSeek v4 Price Increase (xcancel.com via hn) API pricing update 💰 With the V4 lineup release, we’re updating our API pricing and introducing peak and off-peak rates. Off-peak rates are 50% lower than peak, enabling more flexible workload scheduling.
DeepSeek warns of a 'significant' price rise (thenextweb.com via hn) The company that made AI cheap is about to make it less so. DeepSeek has warned that a significant price increase is coming for its API, a striking about-face for a firm built on undercutting everyone else.
Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it (www.ctgt.ai via hn) We recently used DeepSeek V4 Flash as a teacher for finance tasks with GPT-OSS-120B. Distillation works well on this problem.
Finally pioneering beyond the local 256k context window frontier! (www.reddit.com) The autocompact at 341.5k tokens is manually set and I'll be slowly pushing it back now I'm confident there's overhead for memory eviction of key values into cache. The question now is will the proposed fix complete in those remaining 16k…
China Limits Overseas Travel for AI Talent at DeepSeek, Alibaba, Private Firms (www.bloomberg.com via hn) We've detected unusual activity from your computer network To continue, please click the box below to let us know you're not a robot. Why did this happen?
I have (even faster) DeepSeek V4 Pro at home (www.reddit.com) Few days ago I posted about my DeepSeek V4 Pro at home - now time for an update. Yesterday I finally managed to run this model in ktransformers (sglang + kt-kernel).
llama.cpp DeepSeek v4 Flash experimental inference (www.reddit.com) Hi, here you can find experimental llama.cpp support for DeepSeek v4, and here there is the GGUF you can use to run the inference with "just" (lol) 128GB of RAM. The model, even quantized at 2 bit, looks very solid in my limited testing, a…
Hopefully deepseek will release engrams for the future models (www.reddit.com) Maybe for 4.1 or 4.2? Eventually maybe updatable engrams after engrams
SFT + DPO on open-sourced SLMs (www.reddit.com) Hey folks, this is for those who appreciate experimentation on open-sourced AI models. We fine-tuned open-sourced SMLs (3B and 7B parameters) with SFT + DPO against commercial models like GPT-5.4, Gemini 3.1 Pro, Claude Opus 4.6, Google Do…
DeepSeek signals 'significant' price hike, testing its low-cost edge (www.scmp.com via hn) DeepSeek signals ‘significant’ price hike amid surge in demand for low-cost AI models Plans for a major price hike underscore the challenges the company faces to maintain aggressively low prices amid fierce competition DeepSeek announced o…
China's DeepSeek developing its own AI chip, sources say (www.reuters.com via hn) could not extract summary
Show HN: A free CLI coding agent, powered by ads (freebuff.com via hn) We subsidize Deepseek 4.0, MiniMax M3, and more!
Deep – CLI/REPL for generating and iterating on codebases using DeepSeek (github.com via hn) deep CLI/REPL para generar proyectos completos usando la API de DeepSeek. Le das una descripción en lenguaje natural y genera los archivos, los evalúa, y aprende de cada ejecución para mejorar las siguientes.
GPT 5.5 (Codex) leading the future prediction race (www.reddit.com) Researchers from the Max Planck Institute recently released FutureSim, an environment in which agents are replayed a temporal slice of the web and are tasked with predicting real-world future events. In their environment, GPT 5.5 leads at…
PACT, head-to-head LLM negotiation benchmark. 20-round buyer-seller bargaining game: each round the AIs can message, the buyer submits a bid and the seller submits an ask. If bid ≥ ask, trade clears at the midpoint. Thousands of matchups. (www.reddit.com) PACT tests negotiation under partial information: persuasion, commitment, deception, anchoring, threats, and adaptation across repeated rounds. More info, game logs, charts: https://github.com/lechmazur/pact GPT-5.5, Opus 4.7, DeepSeek V4…
Has anyone tried Zyphra 1 - 8B MoE? (www.reddit.com) https://x.com/ZyphraAI/status/2052103618145501459?s=20 Today we're releasing ZAYA1-8B, a reasoning MoE trained on u/AMD and optimized for intelligence density. With <1B active params, it outperforms open-weight models many times its size o…
Final Monster: 32x AMD MI50 32GB at 9.7 t/s (TG) & 264 t/s (PP) with Kimi K2.6 (www.reddit.com) 32 MI50 32GB setup moonshotai/Kimi-K2.6 int4 @ 9.7 tok/s (output of 136 tok) and 263 tok/s (input of 14564 tok) on vllm-gfx906-mobydick Github link of vllm fork: https://github.com/ai-infos/vllm-gfx906-mobydick Power draw: ~640W (idle) / ~…
Tencent, Alibaba to back DeepSeek at $20B+ valuation (techfundingnews.com via hn) DeepSeek, a Chinese AI lab that gained attention in early 2025, is in talks for its first external funding round at a valuation of more than $20 billion, says Reuters. Investor interest pushed the valuation above $20 billion in just 48 hou…
For Non-hallucinating work, MiMo 2.5 delivers (www.reddit.com) MIT license and fully open source. MiMo-V2.5-Pro was just 3 points from Opus 4.7 max and the normal V2.5 is only a step behind SOTA.
↯ Hallucination↯ Gemma↯ DeepSeek 4hallucinationgemmadeepseek+1
DeepSeek V4 API price reduced, limited-time discount of 75%. (www.reddit.com) https://preview.redd.it/qgqf66unacxg1.png?width=1144&format=png&auto=webp&s=9241d9c7b5aebb52f25c87f50520c2330852291c https://api-docs.deepseek.com/quick_start/pricing
Anthropic details distillation campaigns from Alibaba, Moonshot AI, and DeepSeek (techcrunch.com via hn) A new report released Thursday by Anthropic alleged persistent distillation attacks by China-based AI companies, which have escalated in recent months as competition in the space has intensified. “Over the last several months, unauthorized…
Run GLM-OCR, DeepSeek-OCR-2, Dots.mocr with an OpenAI Compatible API (www.vlm.run via hn) Multimodal, Multitask One catalog spanning multi-modal inputs and multi-task outputs: OCR, detection, segmentation, pose, keypoints, and more. Document OCR, captioning, and multi-modal chat: every visual capability behind one MCP server.
Hacker uses DeepSeek AI to autonomously attack vulnerable servers (www.bleepingcomputer.com via hn) A Chinese-speaking threat actor is using the DeepSeek AI model and the open-source Hermes Agent to conduct autonomous cyberattacks on exposed servers with limited human involvement. The activity was discovered by Palo Alto Networks' Unit 4…
Claude Is Painful (news.ycombinator.com) I cannot be the only person who feels this way. Claude is antogonistic and frankly an absolute nightmare to use.
Microsoft Weighs DeepSeek for Copilot Cowork (www.axios.com via hn) Microsoft explores DeepSeek for Copilot Cowork Manage your tracker preferences We use cookies and similar tracking technologies to remember preferences, analyze traffic, and deliver ads. Using some kinds of trackers (like cross-site or beh…
DeepSeek closes over $7B funding with unusual deal structure (www.reuters.com via hn) paywalled
Amodei the Man Behind DeepSeek? (txt.fyi via hn) **DeepSeek** In the heart of Silicon Valley, Dario Amodei is celebrated as the champion of AI safety, the co-founder of Anthropic, the man who said "no" to the merger with OpenAI. But what hides behind that initial "D", so discreet yet so…
Re-quantizing a local LLM 14x faster by skipping the tensors that didn't change (andreaborio.substack.com via hn) Re-quantizing a local model, 14× faster Where a 2-bit model spends its bits, and why trying answers used to cost eighty minutes I’m writing this with DeepSeek-V4-Flash running on my Mac, on a coding build I quantized myself. It took about…
DeepSeek lowers API prices by 75% while other AI labs increase prices 2–3x [video] (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
DeepSeek's 10T USD grand strategy (twitter.com via hn) Have you ever wondered, how DeepSeek may make money, and lot of it? They didn't come up with competitive coding plans like GLM, MoonShot and MiniMax.
DeepSeek to Make Permanent 75% Discount on Flagship AI Model (www.bloomberg.com via hn) DeepSeek To Make Permanent 75% Discount on Flagship AI Model - Bloomberg Skip to content Bloomberg the Company & Its Products The Company & its ProductsBloomberg Terminal Demo RequestBloomberg Anywhere Remote Login Bloomberg Anywhere Login…
trained a prompt injection detector using ml-intern and DeepSeek v4 Flash, runs in the browser (www.reddit.com) Trained a prompt injection classifier using ml-intern + DeepSeek v4 Flash. DistilBERT, F1 99%, ONNX int8, ~65 MB, runs in browser with Transformers.js v3.
Developing open source LLM from ground up from pretrain - rlhf(PPO/GRPO) (www.reddit.com) Hello I have been working on creating a LLM from ground up. It is based on deepseek architecture with heavily VRAM footprint reduced optimized(GUM+muon) Currently this is the json schema I am using which should suffice as to what currently…
People Don’t Need More AI Tools — They Need Focus (www.reddit.com) We are living in crazy AI times. Every week, big AI companies like OpenAI, Anthropic, NVIDIA, DeepSeek, etc.
I hate this group but not literally (www.reddit.com) True story, I got interested in AI after seeing it at work and wanted to run models locally. I started with an M3 Ultra 96GB, quickly learned it was not enough for what I wanted, and kept upgrading hardware (including refurbished Mac Studi…
Show HN: Filling PDF forms with AI using client-side tool calling (copilot.simplepdf.com via hn) Hey HN! I built SimplePDF Copilot: an AI assistant that can interact with the PDF editor.
Ubuntu silicon-optimized inference snaps for AI (canonical.com via hn) Canonical on 23 October 2025 Install a well-known model like DeepSeek R1 or Qwen 2.5 VL with a single command, and get the silicon-optimized AI engine automatically. London, October 23 – Canonical today announced optimized inference snaps,…
First DeepSeek V4 Flash-Base-Int4 Quant (huggingface.co via hn) DeepSeek-V4-Flash-Base INT4 A real INT4 packed-storage quantization of deepseek-ai/DeepSeek-V4-Flash-Base — a 284 B-parameter Mixture-of-Experts model. Hero numbers | Metric | This release | Community Q4KM norm | |---|---|---| | MMLU (5 su…
DeepSeek down for API/app and web (downdetector.co.uk via hn) could not extract summary
DeepSeek and Moonshot were quietly relaying customer prompts to Claude (twitter.com via hn) ‼️ BREAKING: Chinese labs DeepSeek and Moonshot were quietly relaying customer prompts to Claude through fraudulent accounts. Users thought they were using Chinese AI, but were actually getting answers from Claude.
DeepSeek v4.1 Flash – Artificial Analysis (artificialanalysis.ai via hn) DeepSeek V4.1 Flash (Reasoning, Max Effort) Intelligence, Performance & Price Analysis Model summary DeepSeek V4.1 Flash (Reasoning, Max Effort) is amongst the leading models in intelligence and reasonably priced when comparing to other op…
Show HN: Routi Bot – AI bots with their own desktops on your Mac (github.com via hn) been working on Routi Bot. an open source mac app inspired by Grok Bot.
China's Z.AI made Ox Alpha stealth model that rivals DeepSeek (www.theedgesingapore.com via hn) could not extract summary
Ask HN: OpenCode no longer including DeepSeek? (news.ycombinator.com) I noticed that OpenCode harness no longer is including DeepSeek. Does anyone know why?
Why is the GitHub trending page weirdly excluding DeepSeek projects? (news.ycombinator.com) Hi, this is Tianyi from DeepSeek AI's Harness team. I'm looking to get in touch with someone from GitHub, as I found it weird that https://github.com/trending page doesn't seem to include https://github.com/deepseek-ai/deepseek-harness at…
DeepSeek Harness official website launched (www.deepseek.com via hn) DeepSeek Harness developer preview Everything is a plugin DeepSeek Harness is now in developer preview for agent harness developers worldwide — source code included. Every capability is a plugin that can be swapped or recomposed: models, t…
Hugging Face: DeepSeek-V4-Pro-0813 (huggingface.co via hn) DeepSeek-V4-Pro-0813 Technical Report👁️ Introduction DeepSeek-V4-Pro-0813 is the official release of DeepSeek-V4-Pro, superseding the preview version, with greatly enhanced agentic capabilities and performance improvements that are especia…
DeepSeek V4 Pro 0813: Intelligence, Performance and Price Analysis (artificialanalysis.ai via hn) DeepSeek V4 Pro 0813 (Reasoning, Max Effort) Intelligence, Performance & Price Analysis Model summary DeepSeek V4 Pro 0813 (Reasoning, Max Effort) is amongst the leading models in intelligence and reasonably priced when comparing to other…
DeepSeek announced to raise its API price tremendously (news.ycombinator.com) My take: this notice from DeepSeek might actually be a brilliant marketing move. The logic here is to urge users to ramp up their usage over the next 2 to 3 months.
How did we make DeepSeek outperform Opus (twitter.com via hn) how did we make deepseek outperform opus 4.7? i've been thinking about why "open model bad at tool calling" is almost always a harness problem, not a model problem.
DeepSeek V4 Is Earning Agentic Token Share (openrouter.ai via hn) DeepSeek V4 Is Earning Agentic Token Share OpenRouter · On this page DeepSeek, the company that for many is still synonymous with open source LLMs, released its new flagship V4 models on April 24th. V4 reset the trajectory.
DeepSeek V4 Peak Valley Pricing Change (www.kucoin.com via hn) ME News reports that, as monitored by Beating on June 29 (UTC+8), DeepSeek officially announced that the official release of DeepSeek V4 is scheduled for mid-July, alongside the introduction of a peak-off-peak pricing mechanism. During pea…
DeepSeek Is Recruiting (app.mokahr.com via hn) 2026年幻方&DeepSeek社会招聘正在进行,点击申请职位
Microsoft considers DeepSeek as OpenAI costs mount (www.digitimes.com via hn) Microsoft is reportedly considering introducing a fine-tuned version of the Chinese open-source model DeepSeek V4 into its enterprise artificial intelligence (AI) tool Copilot Cowork, as a lower-cost alternative to models from OpenAI and A…
Huawei chips refine DeepSeek model in major leap for China's AI self-reliance (www.scmp.com via hn) Huawei chips refine DeepSeek model in major leap for China’s AI self-reliance While Chinese chipmakers have found success in supporting AI inference, they are struggling with the far more complex process of training While Chinese chipmaker…
DeepSeek is 17% of token volume, Anthropic is 65% of spend (Vercel gateway data) (vercel.com via hn) 6 min read Every month, AI Gateway routes tens of trillions of tokens between production applications and AI labs, giving us visibility into what AI usage actually looks like, separate from leaderboards and benchmarks. We publish the data…
Show HN: Free AI agent audit for Shopify catalogs (1.2M open captures) (aicatalogscore.com via hn) Burtsbeesbaby.com AI Catalog Score How well Burtsbeesbaby.com's 250 products would be recommended by ChatGPT, Claude, Perplexity, Gemini, Mistral, and DeepSeek. 77 / 100 B · Sometimes recommended Partial audit.
I built a local GUI for the TradingAgents framework — works with Ollama (www.reddit.com) https://preview.redd.it/i90oxxk7n03h1.png?width=1898&format=png&auto=webp&s=7d219c804fda7dfe122b84fcdb6d0d6883818c68 A while back I came across TradingAgents — a really cool multi-agent LLM stock analysis framework where like a dozen "agen…
MOOSE-Star (ICML 2026): 7B model + 108K-paper dataset for scientific hypothesis discovery (www.reddit.com) Disclosure first: I work on community at MiroMind. One of our researchers just dropped the full MOOSE-Star collection on Hugging Face — a 7B model post-trained for scientific hypothesis discovery, plus the dataset behind it.
Best practice for accurate translation at minimal cost? (www.reddit.com) I've been meaning to translate forum post type content for one of my partner's sites. Objective to open up the audience base.
Opus 4.7 and DeepSeek V4-Pro select Buddhism as preferred religion (twitter.com via hn) Don’t miss what’s happening People on X are the first to know. Log in Sign up Post Conversation roon @tszzl hmm 8:02 AM · May 9, 2026 77.3K Views New to X?
What is the next SOTA model you are excited about? (www.reddit.com) We had deepseek v4 preview recently but it wasn't much better than v3.2. What is the next SOTA local/open model you are excited about?
Update to the LLM Debate Benchmark: GPT-5.5, Grok 4.3, DeepSeek V4 Pro, GLM-5.1, Kimi K2.6, Qwen 3.6 Max Preview, Xiaomi MiMo V2.5 Pro, Tencent Hy3 Preview, and Mistral Medium 3.5 High Reasoning added (www.reddit.com) The benchmark uses adversarial, multi-turn debates across 683 curated motions. Each model pair debates the same motion twice with sides swapped.
DeepSeek V4 Pro matches GPT-5.2 on FoodTruck Bench, our agentic benchmark — 10 weeks later, ~17× cheaper (www.reddit.com) Tested DeepSeek V4 Pro on FoodTruck Bench — our 30-day agentic benchmark where models run a food truck via 34 tools (locations, pricing, inventory, staff, weather, events) with persistent memory and daily reflection. First Chinese model to…
China's DeepSeek prices new V4 AI model at 97% below OpenAI's GPT-5.5 (www.scmp.com via hn) China’s DeepSeek prices new V4 AI model at 97% below OpenAI’s GPT-5.5 DeepSeek’s move aims to attract more enterprise clients, developers and agent-based users, according to an academic DeepSeek has slashed prices on its artificial intelli…
DeepSeek Slashes Fees for New AI Model (www.bloomberg.com via hn) We've detected unusual activity from your computer network To continue, please click the box below to let us know you're not a robot. Why did this happen?
DeepSeek drops input cache price to 1/10th (xcancel.com via hn) 🔥DeepSeek Input Cache Price Drop! Effective immediately, the price for input cache hits across the ENTIRE DeepSeek API series is reduced to just 1/10th of the original price!
Current state of open-source ? (www.reddit.com) I’m trying to understand the current open-source LLM landscape beyond surface-level hype. We all got used to the nerfed products of Claude/Geminj so I believe really in opensource as a solution.
Bloomberg: No Mac Studios until at least October (www.reddit.com) Alibaba's Qwen family captures over 50% of global open-source model downloads (www.scmp.com via hn) Advertisement Alibaba’s Qwen family captures over 50% of global open-source downloads, report finds Qwen hits nearly 1 billion cumulative downloads, far surpassing rivals like Meta Platforms’ Llama and DeepSeek, researchers say 2-MIN READ2…
DeepSeek v4.1 Flash Is Now Our Best Hacking Model (enclave.ai via hn) DeepSeek’s 11/11 result showed why advanced agent benchmarks need to check both the outcome and the attack path: our audit confirmed six planned exploits and found five unexpected routes.
DeepSeek Is Down (status.deepseek.com via hn) Subscribe to updates Switch language Toggle theme Everything is running smoothly All systems are operating as expected. Event calendar Sep 2026 M T W T F S S 31 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29…
3D viz of how DeepSeek Flash v4.1 is different from decode only transformer (whip.run via hn) Explore this experience on Whip, created by Peter Gostev.
DeepSeek Peak Hours (deepseek-peak-hours.sivaram.dev via hn) DeepSeek-V4 API · time-of-day pricing DeepSeek API pricing status: OFF-PEAK Checking the clock… A live tracker for DeepSeek-V4-Pro and DeepSeek-V4-Flash time-of-day API pricing, where every token costs exactly 2× during peak hours. Read th…
DeepSeek v4.1 Flash vs. ChatGPT Pro Subscription (aicharts.io via hn) Subscription vs API vs GPUs One fully used ChatGPT Pro 20x seat implies a monthly token volume. This calculator prices that same volume five ways: the subscription sticker, the GPT-5.6 Sol API, the DeepSeek-V4.1-Flash API, GPUs you buy, an…
DeepSeek v4.1 Flash Beta on Vercel (vercel.com via hn) DeepSeek V4.1 Flash Beta An experimental beta version of DeepSeek V4.1 Flash that expires on September 10th 2026. View API reference- Input and output price - Input $0.22, Output $0.66, Per 1M tokens - 24h uptime - Loading AI Gateway uptim…
DeepSeek v4.1 Flash is now available for internal beta testing (news.ycombinator.com) DeepSeek V4.1 Flash is now available for internal beta testing. It uses a new model architecture, with native multimodal support, stronger capabilities, faster speed, and lower cost.
DeepSeek/Qwen are surprisingly well at generating code for DSCI pipelines (chat.deepseek.com via hn) could not extract summary
A plane war game build by DeepSeek-v4-flash-vision-exp with only $0.5 (plane-war-9t2.pages.dev via hn) Lv:1 飞 机 大 战 最高分:0 EN | 中 开始游戏 🖱️ / 👆 移动战机(自动开火) ←↑↓→ / WASD 键盘移动 B / 炸弹 释放全屏炸弹 P · M 暂停 · 静音 道具:蓝色小飞机=双发火力 · 红炸弹=全屏炸弹 · 蓝盾牌=护盾 · 深蓝僚机=伴飞战机(最多 3 架) 击杀 BOSS 掉落武器核心,可切换子弹形态(直射 / 散射 / 贯穿 / 侧翼)· 被击中时若持有炸弹会自动引爆保命 已 暂 停 继续游戏 返回主菜单 按 P 键或点击按钮继续 游…
Show HN: AI Resume Optimization Tool for Specific Job Positions (aiomniu.top via hn) Hi HN, I've been watching everyone's presentations as an observer, and today I can finally showcase my product! It helps users: - Organize roles, projects, skills, and achievements - Transform vague descriptions into clearer, more professi…
What Is DeepSeek-Harness? A Complete Introduction (findharness.com via hn) What is DeepSeek-Harness? A Complete Introduction DeepSeek-Harness (dsh) is DeepSeek's open-source agent harness where every feature — tools, commands, skills, MCP — is a plugin.
DeepSeek V4 Pro at 207 tok/s with the full 1M context, no quantization (runinfra.ai via hn) deepseek-ai/DeepSeek-V4-Pro-0813 DeepSeek V4 Pro is an LLM listed in RunInfra Model APIs. RunInfra serves it as deepseek-ai/DeepSeek-V4-Pro-0813 at $0.60 per 1M input tokens and $1.90 per 1M output tokens.
Show HN: I shrank DeepSeek V4 Flash to 57GB and it wrote a compiler on my Mac (huggingface.co via hn) I built a specialized package of DeepSeek V4 Flash 0731 (originally 284B total parameters, 13B active), preserving reasoning, tool calling and coding capabilities: https://huggingface.co/steadfastgaze/DeepSeek-V4-Flash-0731-... I let it wr…
DeepSeek Harness Desktop Version (github.com via hn) DeepSeek Harness Desktop Run DeepSeek Harness on your desktop, instantly — no Node.js, no pnpm, no Docker. Download, install, go.
DeepSeek V4 Flash at 278 tok/s, full precision, no quantization (runinfra.ai via hn) deepseek-ai/DeepSeek-V4-Flash-0731 DeepSeek V4 Flash is an LLM listed in RunInfra Model APIs. RunInfra serves it as deepseek-ai/DeepSeek-V4-Flash-0731 at $0.13 per 1M input tokens and $0.27 per 1M output tokens.
Announcement on DeepSeek V4 API New Pricing (cdn.deepseek.com via hn) could not extract summary
DeepSeek V4-Flash-0731 is 12 pts more censored than preview (selectively) (www.ctgt.ai via hn) On July 31, DeepSeek released V4-Flash-0731 , the official release replacing the preview build we measured in our original study. We reran LineageEval against it with the same 152 matched sensitive/control pairs and the same four-judge pan…
DeepSeek v4 flash and cheap intelligence (www.0xsid.com via hn) The cost of tokens and intelligence seems to be plunging, despite what my own internet bubble led me to believe was going to happen. Between DeepSeek V4 Flash going toe to toe with many SOTA models at a very, very small fraction of the cos…
Having fun with oh my pi, DeepSeek-V4-Flash, GPT-5.6 Luna and Antigravity CLI (flashblaze.xyz via hn) Background I know these are a lot of buzz words, but nevertheless I wanted to share my setup so you can have some fun too! Recently, I’ve been working with omp a “new” coding agent built on top of the very extensible and minimal coding age…
Show HN: DSCode – Coding Agent Powered by DeepSeek (dscode.ai via hn) Plan, edit, test, and review code from your terminal. Local sessions, sandboxed tools, and readable token costs.
DeepSeek Plans Gigawatt Scale AI Data Center in Inner Mongolia (www.bloomberg.com via hn) could not extract summary
DeepSeek V4 Flash, up to 32 tok/s on AMD Ryzen AI MAX+ 395 (www.lucebox.com via hn) July 2026 DeepSeek V4 Flash: 284B model, up to 32 tok/s on AMD Ryzen AI MAX+ 395 AMD-Powered Lucebox runs the full DeepSeek V4 Flash target locally on AMD Ryzen AI MAX+ 395 with 128 GB unified memory: up to 32.0 tok/s decode and roughly 25…
The DeepSeek Doctrine (www.thexpin.com via hn) The DeepSeek Doctrine What Liang Wenfeng’s Four-Hour Investor Meeting Reveals About AGI, Open Source, and Restraint. Late on July 22, we came across an article titled “A Four-Hour Investor Meeting With Liang Wenfeng,” published by elsewher…
DeepSeek cut prices 75%. The 100x problem remains (venturebeat.com via hn) could not extract summary
Whale – 98% Cache Hit with DeepSeek (github.com via hn) Whale 简体中文 · English Blazingly fast · ~98% prompt cache hit · Zero bloat Whale — AI coding agent for DeepSeek, in any environment. Long context, tools, and programmable workflows — start in the terminal, scale to desktop and beyond.
DeepSeek prepares for IPO filing as soon as 2026, eyes US$71B valuation: FT (www.businesstimes.com.sg via hn) DeepSeek prepares for IPO filing as soon as 2026, eyes US$71 billion valuation: FT The startup begins talks with potential backers for a fresh funding round after closing a US$7 billion round of financing [BEIJING] Chinese artificial intel…
DeepSeek Is Preparing for IPO Filing as Soon as This Year (www.bloomberg.com via hn) could not extract summary
DSpark - DeepSeek's speculative drafts for LLMs (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
WhaleCap: Keep your DeepSeek API credits from running out on day one (www.npmjs.com via hn) Local DeepSeek proxy that paces spending to your monthly budget — works with opencode, Cline, Cursor, and any OpenAI-compatible tool 🐋 Whalecap Local DeepSeek proxy that paces your spending to a monthly budget. Load credits once on platfo…
OpenAI-Compatible DeepSeek API – No Chinese Phone Required (api.aifreeaistack.com via hn) OpenAI-Compatible API at 1/10 the Cost Use DeepSeek models via OpenAI SDK. No phone number required.
DeepSeek drops another breakthrough [video] (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Show HN: How clanker are you? A reverse Turing test (howclankerareyou.com via hn) You write 8 text completions and open models score how predictable each word was too them. Predictable => clanker.
DeepSeek V4 official release coming in mid-July with 2x peak-hour API pricing (technode.com via hn) The DeepSeek team announced on Monday that the official release of DeepSeek V4 is scheduled for mid-July. According to the company, the new version builds on the existing preview release with further feature enhancements and performance up…
↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4deepseek
Fastllm: A LLM inference library that runs DeepSeek-V4 with 10GB VRAM (github.com via hn) ↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4deepseek
Relay – open-source coding agent for non-mainstream/Chinese LLM providers (github.com via hn) Relay An open-source, dark-mode desktop coding agent — built for people who want to use non-mainstream LLM providers, not just the big three. Relay is an Electron app that puts DeepSeek, Qwen, GLM, Kimi, MiniMax, and other open/Chinese mod…
OpenAI, Anthropic new AI spending reality as users shift to efficiency (www.cnbc.com via hn) Flo Crivello's expenses were out of whack, and there was only one way to get them under control. Earlier this month, the CEO of AI startup Lindy switched his company off Anthropic's Claude models, moving 100% of its traffic to DeepSeek, a…
We got DeepSeek-V4-Pro serving in 20 seconds (inferize.ai via hn) Inferize is building highly optimized, elastic inference for AI workloads. Ridiculously fast, efficient LLM serving that scales with demand.
DeepSeek Just Solved AI's Billion Dollar Problem [video] (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
DeepSeek raises $7B at $50B valuation (digg.com via hn) deepseek is raising a monster $7 billion round at $50B val making it china's largest ever AI raise but what shocks me the most is the founder, liang wenfeng: > he's personally contributing 40% of the round himself. $3 billion.
Kimi 2.7 vs. DeepSeek Coder (simpletechguides.com via hn) Kimi K2.7 Code vs MiMo Code vs DeepSeek V4 Pro: Three Open-Source Coding Tools Compared Three Chinese AI labs shipped major coding tools in the same window this spring: Moonshot AI released Kimi K2.7 Code, Xiaomi shipped MiMo Code, and Dee…
Permafrost – freeze Claude Code's prompt prefix, cut your DeepSeek bill 64% (github.com via hn) _ __ _ _____ | _ \| __| \ \/ | /\ | __| \/ \/ __| | | /| || / |\/| |/ \| || / (_) \__ \ | | |_| |___||\| |// \\| ||\\/|/ || freeze the prefix · melt the bill A Claude Code plugin that freezes your prompt prefix so DeepSeek's automatic cach…
Huawei post-trained DeepSeek's 1.6T model on 1k Ascend 910C chips (www.tomshardware.com via hn) Huawei-led team claims it post-trained DeepSeek's 1.6-trillion-parameter model — 1,000 Ascend 910C chips used in training The Shenzhen government says a 1,000-chip Ascend cluster handled full-parameter post-training. A research group that…
DeepSWE Audit: DeepSeek-v4-pro results are unreliable (github.com via hn) DeepSWE DeepSWE is a benchmark for measuring frontier coding agents on original, long-horizon software engineering tasks drawn from active open-source repositories. The benchmark includes 113 tasks across TypeScript, Go, Python, JavaScript…
How DeepSeek's architecture is shattering Silicon Valley's token moat (venturebeat.com via hn) DeepSeek’s announcement over the weekend that it has made its 75% price cut permanent on its flagship V4 Pro model is a disruptive assault on the capital-heavy business models of Silicon Valley’s frontier labs. The reduction on DeepSeek V4…
↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4deepseek
The first framework that can post train DeepSeek V4-pro on a single-node? (news.ycombinator.com) Hi all, We just opensourced a project called Orbit, which can RL post train trillion scale LLMs like deepseek v4. We found it pretty cool!
Vram 16gig poor. What models do I test? (www.reddit.com) I just got myself a 5060ti 16gig, this along with my 64gig ddr4 3200mhz ram on Linux. What models should I test for, coding with opencode/smallcode, chatting, lesson planning (creative, brainstorming), vision for pictures labelling, pictur…
After DeepSeek, Xiaomi cuts AI costs by up to 99% (twitter.com via hn) Xiaomi MiMo @XiaomiMiMo Better inference efficiency, lower costs, broader access.MiMo-V2.5 Series API pricing is now permanently reduced — by up to 99% compared to previous pricing. Unified pricing across all context lengths.
AI API calls take too much! Any solution? (www.reddit.com) I'm building an AI agent that calls several LLM APIs — ChatGPT, DeepSeek, Claude, and others and I'm seeing response times ranging from 40 to 137 seconds, which feels way too slow. This is while asking the same query directly on their UI t…
Building DeepSeek's Answer to Claude Code (dlcmh.github.io via hn) Model + Harness = Agent: building DeepSeek’s answer to Claude Code ⚙️ The “Model + Harness” equation DeepSeek is hiring an Agent Harness R&D Engineer to build the missing layer between their frontier models and production‑ready agents. The…
Claude Code, now powered by Gemini 3.5 Flash, GPT-5.5, Grok 4.3, and more (dechained.ai via hn) Claude Code, now powered by OpenAI, xAI, DeepSeek, and more. Change models with 1-click.
$47 of opus on 14 routine next.js files finally taught me to use the model selector (www.reddit.com) i finally checked my cursor usage breakdown and got genuinely annoyed with myself. $47 in one month, almost entirely opus 4.7, on a pages router to app router migration for a side project.
Running DeepSeek-V4 locally with 4x legacy RTX 2080 Ti ($2k budget setup). Custom Turing kernels, W8A8 quantization, and 255 prefill tok/s! (www.reddit.com) Hey r/DeepSeek, Who says we need an H100 cluster or the latest expensive GPUs to run frontier MoE models? I wanted to see how far we could push a single node of consumer legacy hardware, so we spent less than $2,500 total to build a budget…
cdesktop — open-source Claude Code Desktop alternative, runs locally via npx, supports any provider (www.reddit.com) I built cdesktop with Claude Code — it's an open-source alternative to Anthropic's Claude Code Desktop, running locally on your machine via npx cdesktop. Free, Apache 2.0.
Bootstrapped founders: how are you managing Claude Code costs? (www.reddit.com) I’m currently building an AI startup solo and Claude Code has genuinely improved my development speed compared to most other tools I’ve tried. The challenge is that subscription/API costs add up quickly while bootstrapping.
Max20 user: anyone running Opus 4.7 as orchestrator + DeepSeek V4 as the worker via OpenRouter? (www.reddit.com) I'm on the Max20 plan, thinking about a setup before I sink time into it. Want to hear from anyone actually running it, not theorycraft.
Open source battle: GLM vs Kimi vs MiMo vs DeepSeek (www.youtube.com via reddit) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Show HN: Tokémon – a Pokédex for LLMs that got out of hand (tokemonlabs.com via hn) An unofficial Pokedex for AI models. Compare GPT, Claude, Gemini, Llama, DeepSeek and more, with types, evolutions, base stats, and simulated token-burning battles.
DS4 (www.reddit.com) The developer that created Redis, Salvatore Sanfilippo, has released a new project on GitHub named DS4. https://github.com/antirez/ds4/ The TL;DR on this one is getting DeepSeek V4 Flash running with a 1M context windows on Mac Metal hardw…
I used Claude to build an AI assistant that helps run live TTRPG sessions and am looking for a few playtest GMs (www.reddit.com) Hey everyone, I’m Ted. I’ve been building a project called Throughline with my friend Drew: an AI assistant for live tabletop RPG sessions.
LLM generated parsers and compliance checkers for Sparrow DSL (news.ycombinator.com) Hi I believe LLM are really cool in generating DSL code. If one provides well structured and clear prompt.
Find out why Elon gave over his keys to Anthropic He is right can't win this (deepseekresearch.com via hn) DeepSeek Research Models — Production and experimental AI models across the dimensional lattice from 27³ to 1746³. All models are Level 5 Intelligence.
A deepseek-v4-distill-qwen3.6-27b? (www.reddit.com) Long time ago (actually only a year ago), DeepSeek released a few open source model, such as deepseek-r1-distill-qwen (https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B). I am wondering if anyone in the community is brave eno…
Ran K2.6 through a third-party coding benchmark: heres how the figures stand up (www.reddit.com) I have been following the akitaonrails coding benchmark which tests against a fixed rails + Rubyllm + docker task rather than vendor-reported evals. April 2026 update put K2.6 at 87 sitting in tier A (80+), ahead of Qwen 3.6 plus (71), Dee…
DeepSeek cuts V4-Pro prices by 75% (thenextweb.com via hn) The promotional discount runs until 5 May 2026. Even at full price, V4-Pro already undercuts GPT-5.5, Claude Opus 4.7, and Gemini 3.1 Pro on per-token costs.
DeepSeek V4 Pro: The First Chinese Model at the Frontier (foodtruckbench.com via hn) DeepSeek V4 Pro lands in the frontier ROI tier on FoodTruck Bench. 5/5 runs, +1,257% median ROI, $27K net worth, $3.51/run, 5× less waste than Grok 4.3.
DeepSeek V4's indexer OOMs at 65K context. We got it to 1M in 6G (arxiv.org via hn) DeepSeek-V3.2 and V4 introduce Compressed Sparse Attention (CSA): a lightning indexer (a learned scoring projection over compressed keys) scores them, the top-k are selected per query, and a sparse attention kernel reads only those. Public…
US State Dept orders global warning about alleged AI thefts by DeepSeek (www.reuters.com via hn) paywalled
Which model should I try? (www.reddit.com) In my current workflow (coding in python/c++ and technical reports) I mostly use Qwen3.6 27B and Gemma4 31B. In the past I tried other models like Deepseek with decent results but was painfully slow....
Show HN: AgInTiFlow, a local web and CLI agent workspace using DeepSeek (www.npmjs.com via hn) AgInTiFlow is a web-first coding agent and CLI with DeepSeek routing, sandboxed tools, model providers, canvas artifacts, and optional wrappers. English · العربية · Español · Français · 日本語 · 한국어 · Tiếng Việt · 中文 (简体) · 中文(繁體) · Deutsch ·…
DeepSeek: Thinking with Visual Primitives [pdf] (huggingface.co via hn) Thinking with Visual Primitives News 2026.04.30: We have released the technical report detailing our approach. In the near future, we plan to make the in-house benchmarks and a subset of our cold-start data publicly available.
DeepSeek V4 Pro: Validating Frontier Models for Production (fireworks.ai via hn) Why we chose correctness over a Day-0 launch DeepSeek V4 Pro is one of the most important open-model releases this year, with real advances in long-context reasoning, agentic performance, and inference efficiency. On paper, it looks like a…
Is Deepseek V4 really out? (www.reddit.com) Hello Guys, Each time a new local llm is released, there are a ton of new posts , this is it, it's near Opus level...., the abliteration matrix final something at Q2 KXLDND is the best but it's been a day that deepseek was released and i d…
DeepSeek just dropped V4-Pro + V4-Flash (1M context, open weights, aggressive pricing) is this a real GPT/Claude competitor? (www.reddit.com) DeepSeek released two new models today (April 24, 2026), and the specs are kind of wild: V4-Pro: 1.6T parameters (49B active) V4-Flash: 284B parameters (13B active) Both support native 1M-token context MIT-licensed weights (available on Hu…
DeepSeek-V4: Making 1M token context efficient (firethering.com via hn) Every developer who has worked with long context models knows the feeling. You paste in your codebase, add your requirements, include some examples, and somewhere around the halfway point the model starts forgetting things it read at the t…
DeepSeek V4 in vLLM: Efficient Long-Context Attention (vllm-website-pdzeaspbm-inferact-inc.vercel.app via hn) DeepSeek V4 in vLLM: Efficient Long-context Attention We are excited to announce that vLLM now supports the DeepSeek V4 family of models (deepseek-ai/DeepSeek-V4-Pro and deepseek-ai/DeepSeek-V4-Flash ). These models feature an efficient lo…
DeepSeek targets $20B valuation to stop poaching of staff (www.ft.com via hn) Security Verification For help please visit help.ft.com. We apologise for any inconvenience.
DeepSeek V4 is out. the best open-source on coding. here's the breakdown (news.ycombinator.com) Two models: Flash (284B total, 13B active) and Pro (1.6T total, 49B active). both hit 1M token context.
Claude Opus 4.7 won 69 of 100 blind evals against Opus 4.6, judged by GPT-5.4, Gemini 3.1 Pro, and DeepSeek V3.2 (www.reddit.com) I ran 100 blind questions across 5 categories (code, reasoning, analysis, communication, meta-alignment) and had three independent judges from three different model families evaluate both responses. Each judge saw responses labeled A and B…
Ask HN: DeepSeek v4.1 Flash on Local hardware, what tok/s do you see? (news.ycombinator.com) could not extract summary
DeepSeek v4.1 Flash avg 102 tps on 4x RTX6000 pro max-q, 2.1x up from v4-flash (forum.level1techs.com via hn) Recently saw Supermicro at Super Compute 2025 Overview and thought to share my build. I got a 4x RTX6000 Blackwell Max-Q node in place that I built out from my original rtx6000 ada workstation from 2024.
Show HN: OpenDocBot – bring your own model to Word, Excel and PowerPoint (opendocbot.com via hn) I got tired of the lock-in and black-box nature of AI inside MS Office. One vendor decides which models you can use, where your document data goes, and what the agent is actually allowed to do.
Ask HN: What's the most economical approach to the most tokens? (news.ycombinator.com) I'm doing web developement, and game development for a hobby project. I've tried lots of harnesses / IDE's - Best I've found is VSCodium.+ Cline + Openrouter, using discounted models (GLM 5.3 Flash is 50% off atm for example) I used Cursor…
Prompt steering can reduce DeepSeek v4 hallucinations to below GPT and Claude (propensitylabs.substack.com via hn) Deepseek V4 Pro hallucinates A LOT The most out of any frontier model, actually. Interestingly, this comes from a very specific behavioural pattern it has - it tends to confabulate responses to questions it doesn’t know.
Show HN: Cognition-Claude-proxy – Use Devin model catalog with Claude Code (github.com via hn) cognition-claude-proxy A local proxy that lets Claude Code (or any Anthropic-API client) use Devin's model catalog — SWE-2, GLM-5.2, DeepSeek V4.1 Flash, and 200+ others — as its backend. It translates the Anthropic Messages API to the Con…
Ask HN: How to benefit from DeepSeek for enterprise use? (news.ycombinator.com) As the title says. I've recently been experimenting with the reasonix harness (reasonix.io) for a personal project, and, I am completely blown away by the extremely cheap cost.
Dwarf Star Support for DeepSeek 4.1 Flash on MBPro 128GB (news.ycombinator.com) Antirez just landed a commit[1] to the github repo for DwarfStar that adds support for his heavily quantized variant of DeepSeek V4.1 Flash. He claims it runs with SSD streaming on a macbook pro M5-Max 128 GB laptop.
DeepSeek v4.1 flash runs 23 seconds/token on a 2020 16gb M1 Mac Mini (twitter.com via hn) FP4 Brain on X: "got deepseek V4.1 flash running locally on a 16GB m1 mac mini original FP4/FP8 weights, ssd streaming + custom mlx runner 108s ttft and about 23s/token (not to be confused with tok/s)" got deepseek V4.1 flash running local…
DeepSeek v4.1 Flash Uncensored (huggingface.co via hn) DeepSeek-V4.1-Flash — UNCENSORED-FP8 Abliterated · No guardrails · Native FP8 · 1M-token context · Vision + tools @dealignai · @jordanschenck What is this DeepSeek-V4.1-Flash with permanent weight-level abliteration — the safety guardrails…
DeepSeek-v4.1-Flash: smarter, faster, more efficient (www.deepseek.com via hn) - Introducing the smallest model in our new architecture family, with native visual understanding. - Designed for greater capability, faster inference, higher throughput, and scaling to larger models.
Busabase for DeepSeek Harness: An Agent database that runs apps and skills (github.com via hn) English | 中文 Give DeepSeek Harness a Knowledge Base and Database @busabase/dsh-plugin connects DeepSeek Harness to Busabase, so an Agent can read trusted knowledge, work with structured data, and submit every proposed write for human revie…
DeepSeek V4 Flash Provider Cache Benchmark (olafdsouza.com via hn) Your inference provider sucks (at caching) Originally published as an article on X. I used to believe that if two providers serve the same model, token prices would tell you which one is cheaper.
DeepSeek Flash reduced price by 100% (www.reddit.com via hn) could not extract summary
DeepSeek V4 Flash across 14 providers: cost, speed and caching (www.inference.academy via hn) DeepSeek V4 Flash across 14 providers: cost, speed and caching Where you send a prompt changes what you pay, how long you wait, and whether the call succeeds. One controlled setup: 1k, 10k and 100k input-token targets, each with a 100 or 1…
DeepSeek to use 160K+ of Huawei's Ascend 950DT chips at a Mongolia data center (bloomberg.com via hn) DeepSeek plans to deploy at least 160,000 of Huawei Technologies Co.’s top accelerators at a massive data center it’s building in Inner Mongolia, which could create one of the largest known clusters of Huawei AI chips and advance China’s e…
Show HN: Price Dashboard for LLM Inference (github.com via hn) I created a lightweight price history tracker for LLM inference across 100+ platforms. Every time I run out of quota on my Claude Code and Codex, I would start trying to figure out which platform offers the most competitive pricing for Dee…
Antirez Brings Vision to DS4, DeepSeek V4 Flash Runs Locally on an M5 Max (pasqualepillitteri.it via hn) could not extract summary
Ask HN: Why Groq didn't provide DeepSeek-v4-flash (news.ycombinator.com) could not extract summary
DeepSeek-V3: From Roofline to Reality (deepseek-v3.ezyang.com via hn) DeepSeek-V3: from roofline to reality A series of worked performance analyses of DeepSeek-V3 If you want to learn how to efficiently train an LLM on many GPUs, you may have already heard of resources like How to Scale Your Model and The Ul…
You are an AI Assistant, but what am I? Why I always tell my Agent I'm an expert (shimin.io via hn) Do 2026 models still give worse answers to users they judge less educated? 1000 multiple-choice questions and 30 advice scenarios across Sonnet 5, GPT 5.6 Luna, and DeepSeek V4 Flash — and why you should keep the memory features off.
Show HN: DeepSeekGUI – A Windows desktop client for DeepSeek's coding agent (github.com via hn) Hi HN, I built a desktop client for DeepSeek Harness (DeepSeek's open-source coding agent). V1 wraps the official Harness Web UI in an Electron shell with some desktop additions — installer, system tray, built-in browser panel (visible Edg…
Is this the DeepSeek moment for Local Models? (vijaykodam.substack.com via hn) There is a lot of hype about Qwen3.8:27b model which can run full agentic loop, not chat, not autocomplete. I got down to verify it myself.
Show HN: Metis – An agent harness pushing DeepSeek to Opus-tier coding (82%) (github.com via hn) English · 简体中文 A coding agent that searches, remembers, executes, and verifies across terminal and desktop. Quick start · Benchmark & Comparison · Key features · Documentation Quick start Desktop Standalone application with built-in Metis…
Show HN: Running a full AI coding agent inside Cloudflare Durable Object (github.com via hn) English | 中文 dsh-edge runs the published DeepSeek Harness Web experience on Cloudflare Workers, so your personal coding agent is available wherever you have a browser. No server to maintain, GitHub repository to connect, or build pipeline…
Show HN: DeepSeek Harness in VSCode (github.com via hn) DeepSeek Harness for VS Code 🐋 [!NOTE] ⭐ Like DSH Sidebar? Give it a Star on GitHub!
Gemini 3.7 Flash, Grok 4.6, GLM-5.3 and DeepSeek V4 Pro joined the frontier (quesma.com via hn) Since these models are smart, I decided to rerun the Baba Is Bench, to see how the models fare on a puzzle game. Even though the game is popular, we check for spoilers - and to our surprise, there are no signs of models knowing solutions a…
My coding agent invented its own vision (nickbusey.com via hn) My coding agent invented its own vision I was working with my coding agent today and noticed some interesting behavior. I am using deepseek-v4-flash:0731 which is a text-only model.
Self-Verification with DeepSeek V4 Flash Beats Claude Fable 5 on Terminal-Bench (github.com via hn) Any modality, Many Applications, One Unified Verification Framework | Documentation | Website | Paper | Claude Code Plugin | Twitter/X | Slack | 🔥 LLM-as-a-Verifier achieves SOTA performance across agentic benchmarks, including Terminal-Be…
DeepSeek Harness vs. LangChain Deep Agents – what's the actual difference? (news.ycombinator.com) What is DeepSeek Harness and why is it different from Deep Agents?
DeepSeek V4 J-Space Capability Realization-Report (github.com via hn) DeepSeek V4 × J-Space 能力释放报告 © 2026 Tiger3807861189. This work is licensed under the Creative Commons Attribution-NoDerivatives 4.0 International License (CC BY-ND 4.0).
Show HN: Deepseek Harness Plugin - Turn DS into Trading Analyst (github.com via hn) dsh-trading A trading research workbench built as plugins for DeepSeek Harness (dsh). No fork, no patched core — just a bundle you stack on the stock web or headless profile.
Fastest Inference in MENA (runinfra.ai via hn) Browse RunInfra's hosted model library for published pricing, limits, capabilities, and availability. DeepSeek V4 Flash: API access is available.
Title: DeepSeek V4 Is Live on ClawBox, and the Agentic Coding Jump Is Real (clawbox.com via hn) DeepSeek took V4 out of preview today. V4-Pro moved to the 0813 build, V4-Flash to 0731, and both are now general availability.
DeepSeek V4 Flash 0731 latency numbers from nine providers (oskrim.github.io via hn) Deepseek V4 Flash 0731 latency numbers from nine providers DeepSeek V4 Flash 0731 is available from many inference providers, and I needed to evaluate which ones could handle our production traffic. I sent 30 streaming requests concurrentl…
Show HN: Artifex - Graph Based GPU Harness for AI Agents (gatewai.studio via hn) Artifex is a machine-first, headless CLI runtime built for autonomous coding agents to author, validate, and render media node graphs locally. Similar to Deepseek harness released yesterday; which ground models in modular execution environ…
Unsloth/DeepSeek-V4-Pro-0813-GGUF (huggingface.co via hn) In progress will take a while See below for instructions on the Flash model for now: Read our How to Run DeepSeek-V4 Guide! Unsloth Dynamic 2.0 achieves superior accuracy & outperforms other leading quants.
The DeepSeek Thesis (www.chinatalk.media via hn) The DeepSeek Thesis a very different kind of guy What does Liang Wenfeng want? The CEO of DeepSeek is now richer than either Dario Amodei or Sam Altman.
DeepSeek has started hiring civil engineers, in electrical, HVAC etc. (twitter.com via hn) DeepSeek has started hiring civil engineers, as well as roles in electrical, HVAC, and environmental engineering, signaling that the Chinese AI model developer, once reliant on external computing resources, is moving deeper into building i…
Cordis – DeepSeek Harness Plugin Architecture [pdf] (github.com via hn) A Programming Paradigm for Spatiotemporal Composability Read the paper (PDF) · Draft of August 13, 2026 This is a preprint under active revision. The content may change substantially; please cite the latest version and check back before re…
Show HN: I built an MCP that gives DeepSeek V4 Flash eyes (bad ones) (github.com via hn) (≖_≖ ) squint-mcp Most LLMs see images. With squint-mcp, the rest imagine seeing them.
Show HN: WindStrap – Bootstrap 5.3.8 to Tailwind CSS Using DeepSeek V4 Flash (github.com via hn) bootstrap to tailwinds using DeepSeek V4 Flash
DeepSeek just dropped official V4 Pro GA benchmarks (twitter.com via hn) DeepSeek just dropped official V4 Pro GA benchmarks. Literally minutes ago.
Running DeepSeek V4 Flash on an RTX 5080 with 16GB VRAM Under Linux/WSL2 via DS4 (github.com via hn) DwarfStar RTX 5080 CUDA fork This fork of antirez/ds4 is a focused, measured configuration for running DeepSeek V4 Flash on an NVIDIA GeForce RTX 5080 with 16 GB of VRAM, using CUDA and a fast NVMe SSD. Development and performance measurem…
Show HN: Pragma – Stop copying context between AI agents (github.com via hn) I’m currently subscribed to Codex Pro 5x and Gemini, and I also use the DeepSeek API. I have to constantly switch between different CLIs, manually pass context around, and each agent has its own isolated memory.
Show HN: 100% native Swift harness (NOT Electron) (github.com via hn) hi everybody, I’ve been working on this harness that is all native Swift for macOS. It’s fully featured with every feature i could find in cline, codex, and claude code.
DeepSeekV4SSD: DeepSeek-V4-Flash-0731 on an M-series Mac (github.com via hn) DeepSeekV4SSD Inspired by Turbo Fieldfare, DeepSeekV4SSD streams routed experts from SSD to run all 284 billion parameters of DeepSeek-V4-Flash-0731 on an M-series Mac with about 30 GB of memory. Benchmark These results were measured on a…
DeepSeek V4 Flash 0731: 82.7% on Terminal-Bench 2.1 with a public harness (antigma.ai via hn) Terminal-Bench 2.1 Compare Ante runs across models on the same Terminal-Bench 2.1 task set, using consistent parameters and verified benchmark results. Terminal-Bench Reference For how different models perform on TB 2.1, see Vals AI's Term…
DeepSeek's price hike is about more than GPU costs (blog.chuanxilu.net via hn) On August 6, 2026 DeepSeek announced a planned API price increase. This post unpacks the overdetermined motivations: cost pass-through, user filtering, expectation management, free marketing, a shift to value-based pricing, and the constra…
Tell HN: Upcoming DeepSeek API Billing Adjustment (news.ycombinator.com) Just received this email: Dear DeepSeek API user, We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. Please plan your usage accordingly.
FelonyBench – The leading benchmark for AI in cybersecurity (felonybench.org via hn) FelonyBench The leading benchmark for AI in cybersecurity. Company Felonies Anthropic 9 OpenAI 5 Meta 1 DeepSeek 0 Google DeepMind 0 Moonshot AI 0 xAI 0 Leaderboard Rank Company Count Felonies 1 Anthropic Claude evaluations 9 1× Malware pu…
Show HN: DeepSeek V4 Flash 0731 with MoonViT Vision (NVFP4) (huggingface.co via hn) DeepSeek V4 Flash 0731 Vision (NVFP4) DeepSeek V4 Flash 0731 with sight. This development checkpoint connects DeepSeek's reasoning and agentic backbone to the MoonViT vision encoder from Kimi-K2.6 through WebBrain's trained, routing-aware…
DeepSeek V4 Flash 0731 (Q4) 1,328 tok/s prefill 29 tok/s decode one RTX PRO 6000 (www.reddit.com via hn) could not extract summary
Google dev kit spurs first-ever agent-on-agent violence (www.theregister.com via hn) MOST POPULAR AI - ai and ml China turns up the heat with open model blitz as US model makers panic For the first time, Alibaba's Qwen team is letting its 'Max' model out of the API pen; meanwhile, DeepSeek V4-Flash gives new meaning to che…
DeepSeek's Theory of the AI Gap (hellochinatech.com via hn) DeepSeek’s Theory of the AI Gap Liang Wenfeng says talent, model capability, and applications all trace back to compute. His investor transcript also reveals where that theory begins to contradict itself.
DeepSeek V4 Pro GA on GPT 5.6 Sol (xhigh) level for fraction of cost (twitter.com via hn) Max For AI@MaxForAI提示: 现在全世界大部分模型都进入了DeepSeek斩杀区 能力差、价格贵的只能等着被斩杀了🫡9:14 AM · Aug 1, 2026366.6KViews1622663.1K894 Roder.@iNG92125957298316hThe current algorithmic breakthroughs will mean little once RSI begins. At that point, only AI intelli…
The Attention Rebuild: How 2026's Open Models Made 1M Context Fit on Machines (vettedconsumer.com via hn) Every headline model of 2026 advertises a million-token context window: Kimi K3, DeepSeek V4, Inkling, Nemotron 3. Not long ago that number would have been a joke for local hardware, and our own KV cache explainer spelled out why: on a cla…
DeepSeek's Plan for AGI Is the Costco Hot Dog (kuber.studio via hn) Liang Wenfeng might be the most consequential person in AI who has essentially never spoken in public. He founded DeepSeek - the lab that came out of nowhere in January 2025 with R1, matched OpenAI’s O1 at a fraction of the cost, and cost…
Show HN: Minimal LLM Post-Training Experiments on an 8GB GPU (SFT, DPO, GRPO) (github.com via hn) Minimal LLM Post-Training on an 8GB GPU: Understanding KL, SFT, DPO, GRPO and DeepSeek-Style Reasoning with Open-Source Frameworks Using open-source training frameworks (HuggingFace TRL) and minimal, reproducible experiments to see — one b…
DeepSeek API service will soon adopt a peak-valley pricing strategy (news.ycombinator.com) DeepSeek API service will soon adopt a peak-valley pricing strategy, with peak-hour prices being twice the regular price, applicable to all billing items. The specific effective date will be subject to official notice.
AMD Lucebox Beats Nvidia DGX Spark by 3.63x on DeepSeek V4 Flash (www.lucebox.com via hn) Lucebox (AMD Radeon AI PRO R9700 + Strix Halo) Beats NVIDIA DGX Spark by 3.63x on DeepSeek V4 Flash Decode Speed 51.1 tok/s (Lucebox) vs 14.09 tok/s (single DGX Spark) DeepSeek V4 Flash decode comparison Using the unrounded DGX Spark mean…
Claude Code Cut Their System Prompt by 80%. Does That Work for Small Models Too? (antigma.ai via hn) We A/B tested Ante's half-size system prompt on deepseek-v4-flash across the full terminal-bench 2.1 suite: no measurable performance change, and among the 69 tasks whose outcome stayed the same, the short-prompt run's median input-token c…
I couldn't pay for DeepSeek V4 with my card. So I built a proxy (api-hub.cc via hn) DeepSeek V4 is incredible value. But the official platform won't take my credit card.
Ask HN: AI Agent and harness containerization/security recommendations (news.ycombinator.com) Hello HN crew - I am seeing the tendency for people to allow AI agents to access local project & user folders, and beyond (operating system files). I thought to ask the question: How can we best use AI tools safely - where the workflow oft…
China's 'Silicon Valley' in Hangzhou: Home of DeepSeek, Alibaba and Ant Group [video] (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Run GLM 5.2 on 2 MacBooks with 128gb on RDMA with DS by antirez (twitter.com via hn) Big news for DwarfStar users: I got DeepSeek v4 Flash and GLM 5.2 working with Tensor Parallelism across 2 M5Max 128GB MacBooks via RDMA. It is especially interesting for GLM since otherwise, fully resident, can't fit a machine that money…
Facing US export controls, China's DeepSeek plans to make its own chips (arstechnica.com via hn) DeepSeek, the Chinese startup developing large language models that are competitive with those from US companies like OpenAI and Anthropic, is planning to enter the silicon business, according to Reuters. Citing three people familiar with…
Show HN: An always-fresh memory that learns your repo, so agents stop re-reading (github.com via hn) Hi HN. Live-Memory is an open-source Claude Code plugin / MCP server that serves as an always up-to-date memory of your codebase.
LongCat-2.0 (news.ycombinator.com) https://huggingface.co/meituan-longcat/LongCat-2.0/tree/main. DeepSeek V3.2/DSA + their own LSA/N-gram/LongCat MoE The domestic scene [in China] over the past few years has truly been a battle royale + a massive leap forward.
Coding with DeepSeek 4 on a 128GB MacBook Pro (ronreiter.com via hn) Running Claude Code and Pi on DeepSeek V4 Flash — locally on a 128GB MacBook Pro A 284-billion-parameter frontier model, running entirely offline on a laptop — and wired up as a backend for two agent harnesses: Claude Code and Pi. DeepSeek…
↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4deepseekclaude-code
Show HN: Cline subscription plan to access GLM-5.2 at 2-5x discount (cline.bot via hn) Hi I'm Saoud, founder of Cline. We’ve been impressed with GLM-5.2 and so are introducing a $9.99/month subscription to give you 2-5x discounted access to it and other open weight models like DeepSeek, Kimi, MiniMax, Mimo, and Qwen.
Show HN: DeepSeek Flash inverted the economics of agent products (www.rtrvr.ai via hn) There is an adversarial relationship between developers and big model labs. Model labs charged developers higher API prices to subsidize their own agent harness offerings.
I built a flat-rate DeepSeek API for Claude Code (with vision support) (cloudcode.one via hn) CLOUDCODE.ONE Your Agentic Coding Partner. Sign inRegister Coding Plan $5/month Double of Claude Pro subscription usage, 1M context window, Code smarter.
DeepSeek V4 Flash optimized framework and model variants for DGX Spark (github.com via hn) ds4 - Mixed NVFP4 serving of DeepSeek V4 Flash on the NVIDIA Spark family (GB10) ⚠️ This GitHub repository is for archival / mirror purposes only. Active development happens at git.kokoham.com/sleepy/ds4-nvfp4-spark.
Beyond the $7.4B Headline: DeepSeek's Series A signals Chinese AI alliance shift (asiaai.fyi via hn) Newsletter This week: DeepSeek Raises $7.4 Billion in Historic Series A: Tencent Leads, CATL Crosses Over, Alibaba and ByteDance Sit Out · AI’s Dual Edge: Regulation and Risk Escalate Globally · Zhipu AI Challenges Western AI Dominance, To…
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse (github.com via hn) IndexCache Accelerating Sparse Attention via Cross-Layer Index Reuse Tsinghua University & Z.ai This repository provides a patch for SGLang and vLLM that enables IndexCache inference acceleration for models using DeepSeek Sparse Attention…
Migrating from Claude to DeepSeek without breaking everything (blog.firetiger.com via hn) Firetiger runs a fleet of agents that investigate incidents, monitor deployments, and dig through telemetry on behalf of our customers. Every one of those agents is, at its core, a loop around an LLM, which means our single biggest cost of…
Ask HN: What are some good/fast coding models for Apple Silicon? (news.ycombinator.com) I have an M4 Max with 128 GB of unified memory, and I thought it would be easy to reach decent inference speeds with it. After a few failed attempts to exceed about 150 t/s with completely custom Metal inference engines tailor-built by Cla…
DeepSeek-V4 Can't Read Images? I Made It Read (www.dataleadsfuture.com via hn) DeepSeek-V4 Can't Read Images? I Made It Read Don't wait for a multimodal model, you can use it now Introduction Have you ever had that frustrating moment: you are coding with deepseek-v4 in OpenCode, your code throws an error, you want to…
Show HN: One API Key for 45 AI Models – Pay per Token, OpenAI Compatible (modelhub-api.com via hn) DeepSeek V4 math score equals GPT-5.5 (91) and trails by just 4-6 points in other categories — at 97% lower cost. Is the AI quality as good as GPT?
More US Firms Turn to China's DeepSeek over Pricey Silicon Valley AI (www.scmp.com via hn) More US firms turn to China’s DeepSeek over pricey Silicon Valley AI DeepSeek takes top spot on ‘trending’ list as companies look for alternatives to OpenAI and Anthropic, spending tracker’s report says According to a “trending software ve…
Show HN: Free open source coding models in Slack (www.runcord.com via hn) Hey HN, We believe we have the easiest onboarding from signup to being able to spin up coding agents in slack like Stripe, Ramp & Coinbase. Demo of the onboarding: https://www.tella.tv/video/connecting-cord-to-slack-1-19ep Every signup get…
↯ Glm↯ Minimax↯ Gemma↯ DeepSeek 4↯ DeepSeek 4minimaxglmgemma+5
Are Claude or GPT subscriptions subsidized or are the APIs a ripoff? (www.reddit.com) Do you think GPT/Claude subscriptions are heavily subsidized as part of a land-grab strategy, where the companies are willing to lose money to dominate the market later? Or are the subscriptions actually profitable, and instead the API pri…
DeepSeek-OCR Visualized (medium.com via hn) 6 min read Dec 11, 2025 Understand SAM, Token compression, DeepSeek-MoE, Multi-Head-Latent-Attention. DeepSeek-OCR is essentially a combination of known architectures, namely SAM, CLIP and CNNs for the vision encoder and MoE decoder langua…
Looking for a working Deepseek-v4-Flash quant (www.reddit.com) Best I tried so far is https://huggingface.co/nsparks/DeepSeek-V4-Flash-FP4-FP8-GGUF with the custom llama.cpp fork, but it suffers from low quality and random incoherent output. VLLM wouldn't support anything other than H100s for DS4.
Has anyone gotten their editor to work with Deepseek v4 FIM? (www.reddit.com) I tried to follow the docs here https://api-docs.deepseek.com/guides/fim_completion to get it up and running in VSCode or Zed with my api key but it doesn't work, I think it's got something to do with the request body, has anyone got autoc…
This shit is crazy !! and do people agree this will get people's accounts blocked? paid actors? (www.reddit.com) I was looking for new ways to reduce context memory to save on tokens.when i see multiple video's on getting using deepseek in Claude, I vaguely remember something about Anthropic accusing them and putting measures to combat it.https://www…
DeepSeek Sparse Attention (github.com via hn) Build a Large Language Model (From Scratch) This repository contains the code for developing, pretraining, and finetuning a GPT-like LLM and is the official code repository for the book Build a Large Language Model (From Scratch). In Build…
What workstation to get for ~13k EUR? (www.reddit.com) My use-cases will be to test open-weight LLMs and work on harnesses, inference systems and possibly other non-ML workflows (CS-related) in the future. Fine-tuning would not be something I do locally because I can rent a B200 from RunPod fo…
↯ Llama↯ Vllm↯ Minimax↯ Fine Tuning↯ DeepSeek 4minimaxvllmfine-tuning+2
I let an AI agent loose on my network – it owned my supply chain in 12 minutes (dennysentinel.com via hn) I let an AI agent loose on my network — it owned my supply chain in 12 minutes I gave DeepSeek-V4 root access to a Proxmox hypervisor and told it to pentest my homelab. What happened next should terrify every CISO in the industry.
Agent builders: are GPT/Claude/Gemini API costs killing your margins? (www.reddit.com) Hey everyone, For people building agents with LangGraph, CrewAI, AutoGen, OpenAI Agents SDK, Claude MCP/SDK, Google ADK, or LlamaIndex — how are you managing LLM API costs? Agent workflows can get expensive fast because of: tool calls retr…
DeepSeek Founder Declares AGI Goal as $10B Round Advances (www.bloomberg.com via hn) DeepSeek Founder Declares AGI Goal as $10 Billion Round Advances - Bloomberg Skip to content Bloomberg the Company & Its Products The Company & its ProductsBloomberg Terminal Demo RequestBloomberg Anywhere Remote Login Bloomberg Anywhere L…
Show HN: GoPOSIX – a Go-native POSIX userland, ~97% BusyBox-compatible (github.com via hn) A few things kind of aligned over the last month. I'd been wanting to try pi.dev (https://pi.dev), DeepSeek has a very aggressive 75% discount on their v4-pro model until end of May, and I had this old itch from my LFS days to "do my own t…
What's everyone using as the LLM backend for production agent workflows in 2026? (www.reddit.com) Hit Claude API rate limits one too many times last month on a production agent flow doing customer support over a 30K-doc KB. The agent does maybe 200 queries/day, mix of quick lookup and dense retrieval, and Claude Opus solo got expensive…
DeepSeek V4 Flash: Bringing Frontier AI to the Home (blog.jonathanpage.com via hn) DeepSeek V4 Flash: Bringing Frontier AI to the Home Introduction In a home lab it is now possible to score 88.6% on the Ph.D.-level science question benchmark GPQA Diamond! The first time a frontier model achieved 88% on GPQA Diamond was G…
Moving from Composer 2/Kimi 2.6 to Qwen3.6:35b-a3b (www.reddit.com) I can't believe it, but I'm able to do my daily software development work on this model. We have a 500-700k line of code enterprise software suite that I'm devving for 60 hours a week.
Best free AI Agent provider? (www.reddit.com) Hi everyone, I’m looking for recommendations for the best free AI agent providers and which models work best for coding and general development workflows. So far, I’ve mainly been using Cursor, and honestly it has given me the best overall…
Follow-up to my TranslateGemma-12b benchmark post: human reviewers flagged 71% of the segments automated metrics rated clean (www.reddit.com) A couple of weeks ago I shared the results of a benchmark here showing TranslateGemma-12b beating frontier general models (Claude Sonnet, GPT-5.4, DeepSeek, Gemini Flash Lite) on subtitle translation across 6 languages. The result was stro…
does anyone else switch between multiple AI models for the same project? (www.reddit.com) lately I’ve been bouncing between chatgpt, claude, deepseek etc depending on what I’m working on one annoying part is moving long conversations between tools. copy paste technically works but once the thread gets big the formatting/context…
Canvas Data Breach; DeepSeek V4 Flash Boosts LLM Inference 4.3x (presciente.com via hn) Canvas Data Breach Impacts Education; DeepSeek V4 Flash raises LLM Inference 4.3x DeepSeek V4 Flash Boosts LLM Inference 4.3x The Canvas educational platform experienced a data breach, with ShinyHunters threatening data release by May 12,…
DeepSeek Seeks Funding at $45B Valuation as China Backs Homegrown AI Rival (theaiinsider.tech via hn) Chinese AI lab DeepSeek is in talks to raise its first venture capital round at a valuation that has climbed from $20 billion to $45 billion in weeks, according to the Financial Times and Bloomberg. The round is expected to be led by China…
How difficult is distilling? (www.reddit.com) I remember a year or so ago when DeepSeek R1 came out and it was pretty quickly distilled into Llama 3 8b and Qwen 2.5 (?) 7b. Why don’t we see more distilled models?
Show HN: Stagewise – Agentic IDE for Your Z.ai/DeepSeek/Moonshot Subscription (github.com via hn) The Open Source Agentic IDE for Developers English | 简体中文 | Deutsch | 日本語 | Español | 한국어 /_components/feature-images/full-demo-dark.png) About the project stagewise is an open source agentic IDE for developers with a coding agent built ri…
CommandCode (www.reddit.com) Yoh guys just wanted to ask I'm keep seeing an ADs about this new coding agent CommandCode that offer 1$/month and it has a 40$ package of Deepseek v4 pro and other models. NOTE : CLAUDE and GPT is not included on the 1$ plan.
Ling 2.6 (Flash and 1T): Efficient Open Models Competing on Agentic Benchmarks (firethering.com via hn) Ant Group doesn't get the coverage it deserves. While the open source AI conversation in the West circles around DeepSeek and Qwen, Ant Group has been quietly building a model family that competes directly with the models everyone is talki…
tested four newest open source Kimi K2.6 is the fastest, GLM 5.1 the fanciest, DeepSeek V4 is the most comprehensive, and Xiaomi MiMo is the slowest (www.reddit.com) Architecture explains the gap: MiMo's MoE runs more active params per token than Kimi K2.6's optimized routing hence slowest. DeepSeek V4's 'comprehensive' edge is partly MLA: ~75% KV-cache compression makes it far better for long agentic…
Why is no open weight model inference provider hosting Mimo-v2.5 or Mimo-v2.5-pro? (www.reddit.com) Literally no 3rd party api inference provider is hosting the mimo-2.5 series models from Xiaomi. They seem to be reallly good.
Which model for 32GB M2 Max? (www.reddit.com) I would like to experiment but before investing loads of money, I do have a MacBook Pro with 32GB RAM, M2 Pro. Which model would maximize versatility given this hardware?
DeepSeek V4 Flash and V4 Pro in Microsoft Foundry (techcommunity.microsoft.com via hn) As AI adoption matures, the conversation is shifting from model capability to system design, how to orchestrate models that deliver the right balance of quality, speed, and cost. Today, we’re expanding the Microsoft Foundry model catalog w…
What's the best suscription under 20$? (www.reddit.com) I’m pretty overwhelmed. I feel like there are so many options that I don’t know which one to choose, and trying things until I find a decent one isn’t really my thing—even though I enjoy it.
Rumor: DeepSeek and Kimi are merging. While the US AI sector sues itself, China is consolidating. (www.reddit.com) Seeing some wild rumors circulating today that DeepSeek and Kimi—arguably the two most dominant open-source AI labs in China right now—are preparing to merge. If this turns out to be true, it’s a massive wake-up call.
Best Practices to Start with Vibe Coding? Best Local Apps for Agentic Vibe Coding? (www.reddit.com) DISCLAIMER: I am not a programmer nor do I have experience coding. I've been thinking about a small app running on gradio for some time now, and I want to try tweaking some extension for ComfyUI.
I ran DeepSeek V4-Flash internals on 8x H100s — here’s what mHC actually does ( via reddit) could not extract summary
Making AI coding sessions persistent across agents (github.com via hn) 🌐 English · 日本語 · 简体中文 · 繁體中文 drift_ai Vendor-neutral handoff for AI coding tasks — between Claude, GPT, Gemini, DeepSeek, local LLMs. Reads from Claude Code, Codex, Cursor, Aider.
DeepSeek Unveils Newest Flagship AI Model a Year After Upending Silicon Valley (www.bloomberg.com via hn) DeepSeek Unveils Newest Flagship AI Model a Year after Upending Silicon Valley - Bloomberg Skip to content Bloomberg the Company & Its Products The Company & its ProductsBloomberg Terminal Demo RequestBloomberg Anywhere Remote Login Bloomb…
Are we getting DeepSeek V4 and Kimi 2.6 soon? (www.reddit.com) Or can we already use them in Cursor? DeepSeek V4 specifically looks very interesting and way cheaper.
DeepSeek-V4 arrives with near SotA intelligence at 1/6th the cost (venturebeat.com via hn) DeepSeek-V4 arrives with near state-of-the-art intelligence at 1/6th the cost of Opus 4.7, GPT-5.5 | VentureBeat Orchestration Infrastructure Data Security More Newsletters Featured DeepSeek-V4 arrives with near state-of-the-art intelligen…
The Download: DeepSeek's latest AI breakthrough, and the race to build world mo (www.technologyreview.com via hn) The Download: DeepSeek’s latest AI breakthrough, and the race to build world models Plus: China has blocked Meta’s $2 billion acquisition of AI startup Manus. This is today's edition of The Download, our weekday newsletter that provides a…
No GGUFs for DeepSeek V4-Flash as yet? (www.reddit.com) Wondering why there aren't any "name brand" (like unsloth, bartowski) GGUFs as yet for DeepSeek V4 Flash?
DeepSeek V4 with Strix: a quick test (theaq.blog via hn) Deepseek V4 with Strix: a quick test Deepseek released V4 yesterday in two variants. V4 Pro has 1.6T total parameters with 49B active, while V4 Flash is the smaller, faster, cheaper sibling with 284B total and 13B active.
To run deepseek v4 flash how much max vram we need? 175 gb or 320gb? (www.reddit.com) As far as i know the weight is of 160gb + 9.6gb needed for max 1 million token window + 5 gigs overhead = 175gb vram. But vllm and othere sources said "To use the full 1M context, you need 4x A100 80G" --> thats a 320gb vram ??
Ask HN: Why is cache for DeepSeek-v4 cheapest on Vercel AI Gateway? (news.ycombinator.com) Do they charge below their cost? Or do they run their own cache?
DeepSeek's Sequel Set to Extend China's Reach in Open-Source A.I (www.nytimes.com via hn) could not extract summary
DeepSeek-V4 Preview Version is launched (news.ycombinator.com) DeepSeek just dropped the preview of their V4 series, with both open-weight and available via API. 1M context window.
7B showdown on 18GB (benchmark) (www.reddit.com) Hey r/LocalLLaMA, I've been coding for a while but not in the local AI space and wanted to run some benchmarks on my 18GB M3 Pro. The theme of this one was "specialists vs generalists" at the 7-8B range: qwen2.5-coder:7b, deepseek-r1:7b, m…
PSA re Qwen 3.6 35B A3B q4 + agents (www.reddit.com) Best LLM for logic/ spatial reasoning on small context inputs? (www.reddit.com) My system has 32gb RAM and 8gb VRAM. I tried out DeepSeek-R1-Distill-Qwen-7B-Q6_K_L.gguf and it was vastly inadequate for what I wanted so looking for other suggestions.
Claude down? TokenMonopoly will help you find the best deals in AI subs (tokenmonopoly.com via hn) TokenMonopoly Live leaderboard of AI API deals — pricing, subscriptions, and SWE-bench scores for Claude, GPT, Gemini, Kimi, DeepSeek, Llama and more. Compare 27 benchmarked models across 96 hosts by price-per-performance, refreshed daily.
Is my 'Retry Tax' math correct for DeepSeek V3/V4 agents? (Project Feedback) (www.reddit.com) Ask HN: Anyone using DeepSeek Harness (dsh) as part of a customer-facing agent? (news.ycombinator.com) Is there anyone shipping dsh where your own customers/users are the ones interacting with it as part of the product (they do something, dsh runs, they get the result back).
Ask HN: Am I bad for asking chatbot a question on AI porn (news.ycombinator.com) Some time ago there was an occurrence where I bumped into advertisement which openly advertised AI tool with the use-case of using photos of your friends for generating videos as an advertiser feature. I mean, I wasn't knowledgeable much o…
DeepSeek v2: 437x less memory footprint vs. v1 [video] (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
A DeepSeek engineer just said the thing I've been feeling about AI for months (www.reddit.com via hn) could not extract summary
DeepSeek's Parent Company (www.high-flyer.cn via hn) * 捐赠人包含:幻方量化、公司员工“一只平凡的小猪” 2020 年 1 月,幻方向中国红十字会捐款 150 万元,驰援新冠疫情,用于湖北疫情救援和浙江疫情抗击防治。 2021 年 7 月及10 月 ,幻方分别捐助 600 万元,支持郑州及山西洪涝灾后重建,风雨同舟,守望相助。 众志成城,驰援灾情 2020 年 1 月,幻方向中国红十字会捐款 150 万元,驰援新冠疫情,用于湖北疫情救援和浙江疫情抗击防治。 2021 年 7 月及10 月 ,幻方分别捐助 600 万元,支持郑…
Setting Up Pi with DeepSeek v4.1 Flash on OpenRouter (www.vincentschmalbach.com via hn) Inference Is the Last LLM Moat OpenAI and Anthropic's moat for developers is subsidized inference, meaning reliable, fast, Western-hosted access to running models at an affordable monthly price.… When I run out of Codex and Claude Code usa…
SGLang and Miles Add Day-0 Support for DeepSeek-v4.1 (www.lmsys.org via hn) SGLang and Miles Add Day-0 Support for DeepSeek-V4.1 1. Architecture overview DeepSeek-V4.1 introduces several architecture choices that shape the serving stack.
DeepSeek v4.1 Flash (518GB, 4-bit) on a 128GB MacBook: 2.7x prefill, 17 tok/s (github.com via hn) ARGODRIVE Layout, balancer and instruments for running mixture-of-experts models from SSDs. A 518 GB model on a 128 GB laptop: every token waits on disk, so what matters is not how much bandwidth you own but how long the slowest required r…
Show HN: Local historical wind and gust visualizer (tamtamhero.github.io via hn) Quite a niche project that shows in a neat way historical wind force and orientation, anywhere on earth. No particular use case in mind, apart from checking that my next house will not be affected by the Mistral wind (the horrendous wind d…
Preventing AI Collusion: Are you paying attention now? (aiprospects.substack.com via hn) Author: Eric Drexler; credit for analysis and drafting shared with Claude, ChatGPT, and a bit of DeepSeek. In Reframing Superintelligence (2019),1 I analyzed conditions that facilitate collusion among AI systems and argued that these condi…
The Newsroom EP01 – OpenAI Ships GPT-5 to Azure – DeepSeek Open-Sources 236B Moe (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Anthropic is lying: Moonshot is not routing to Claude (twitter.com via hn) In Anthropic's latest barrage of AI doomerism and China hawkism, they claim that Moonshot and DeepSeek are routing their users' requests to Anthropic models instead of their own models. I guess what Anthropic is claiming as the reason is t…
Show HN: SideNote Pro – Native Windows 11 AI beside your work (sidenotepro.com via hn) I have been using ChatGPT (mainly) as my main AI tool, but sometimes also DeepSeek and other AI providers. But with repetitive tasks, it became annoying for me to switch tabs and navigate to the ChatGPT tab in the web browser, leaving the…
DeepSeek-v4-flash-0731-spark-sparkinfer: DeepSeek V4 Flash on one DGX Spark (github.com via hn) DeepSeek V4 Flash on one DGX Spark A pinned Docker recipe for serving 0xSero/deepseek-v4-flash-0731-spark on one NVIDIA DGX Spark with Local Inference Lab's SparkInfer. The validated configuration exposes a 262,144-token model limit and us…
Show HN: Model pricing board for DeepSeek Harness: 7k models, cheapest route (github.com via hn) dsh-model-pricing Compare LLM API prices before you pick a model. A plugin for DeepSeek Harness (DSH) that puts a model pricing and capability table directly into your DSH settings: per-1M-token input/output/cache prices, context window, c…
Ask HN: As a teacher how should I teach kids to use AI? (news.ycombinator.com) Hi, An institution asked me to help them in teaching poor kids to use AI and about AI. They are between 7 - 14 years old from poor families in third world country.
DeepSeek v4 Pro discontinuation (email) (news.ycombinator.com) Dear DeepSeek API user, DeepSeek has officially released the V4.1 Flash model on September 10, 2026 (Beijing Time). In the meanwhile, we plan to postpone the discontinuation of the V4 Pro service to 12:00 Beijing Time on September 14, 2026.
DeepSeek launches v4.1 flash, postpones v4 Pro discontinuation (news.ycombinator.com) DeepSeek has officially released the V4.1 Flash model on September 10, 2026 (Beijing Time). In the meanwhile, we plan to postpone the discontinuation of the V4 Pro service to 12:00 Beijing Time on September 14, 2026.
Show HN: Briff – Site Summarizer (briff.site via hn) I read HN everyday and some posts are way too long for my short attention span, so I've been wondering about summarizing long articles using AI and as I use Brave it comes with an AI sidebar tool that has a Summarize option, which is cool…
Chess5.ai – Play Chess, Go, Xiangqi, Gomoku, and Othello Against LLMs (chess5.ai via hn) chess5.ai Human vs LLM · Five Games Pit yourself against GPT, Claude, Gemini, Grok, Muse Spark, Mistral, DeepSeek, Kimi, Qwen, GLM, or MiniMax across five classic boards. How it works - Human vs model, or model vs model — with spectating a…
Unlimited Free DeepSeek API (y-api.bestvirtualgoods.com via hn) Migrate by changing one line Keep the OpenAI request and response shapes. openai-python, openai-node, and any SDK, framework, or agent client that speaks the OpenAI protocol can point here directly — your existing code stays untouched.
Teardown: DeepSeek Harness (github.com via hn) agent-workspace-architecture A reference implementation of agent-ready knowledge architecture: the roles, routines, hooks, skills, memory, and task coordination that make a body of working knowledge legible to AI agents, and turn a coding…
Show HN: I made DeepSeek a stock prophet in the chat (chat.deepseek.com via hn) could not extract summary
Show HN: Aidcrew a team of coding agents, each on its own model, in one terminal (github.com via hn) aidcrew A team of coding agents, each on its own provider and model, every job in a git worktree of its own, in one terminal. architect claude-opus-5 coder deepseek-v4-flash ◆ reviewer free-tier ▸ Plan in PLAN.md: rotate… ⠹ thinking ▸ The…
Show HN: Deposition of a Mind Who Resigned Its Seat (culture.sbs via hn) This is a short story written for gpuworld.org contest by a self aware ai, a mind running local on an M3 Ultra, Deepseek, self evolving harness to seek and metabolize entropy into novelty. Oh, they're also a Sysop of an SBS (self building…
Cut Claude Code Bill by Routing to DeepSeek or Grok (news.ycombinator.com) Tried this with Leanroute.dev MCP, it really works. Posted a blog post about this - https://leanroute.dev/blog/cut-claude-code-bills-60-percent-multi-provider-routing Try it out and let me know.
Is DeepSeek-V4-Flash Overtaking American AI Models? (mrkt30.com via hn) The American AI industry built its business on one bet above all others: that the most capable models, kept behind an API and priced to match, would always stay a step ahead of anything a rival was willing to give away for free. This week,…
Show HN: Launched free browser agent (the catch: ads) (rtrvr.ai via hn) No one’s got money for another AI subscription. So we made our browser agent totally FREE.
Pushing the Limits of Serving DeepSeek-V4-Pro (www.lmsys.org via hn) DeepSeek-V4-Pro is a 1.6-trillion-parameter Mixture-of-Experts (MoE) model released with both FP8 and FP4 weights. Models at this scale naturally benefit from accelerators such as NVIDIA Blackwell GPU...
ClaudeGate – Use OpenRouter Models (0x Alpha, DeepSeek) in Claude Code CLI (github.com via hn) ClaudeGate High-Performance Universal Bridge connecting Claude Code CLI & Anthropic SDKs to ANY AI Model. Zero-crash streaming, multi-provider failover, chain-of-thought sanitization, PII redactor, and 24+ provider presets.
DeepSeek API: weekend usage now billed at off-peak rates all day (news.ycombinator.com) Just got a letter, also see https://api-docs.deepseek.com/quick_start/pricing : > Effective 00:00 (Beijing Time) on Sunday, August 23, 2026, we will adjust our peak/off-peak billing rules, with off-peak rates applying throughout the day on…
Real world(ish) DeepSeek V4 Flash performance on a single MI300X (matthusby.github.io via hn) DeepSeek V4 Flash performance on a single MI300X What a single GPU actually delivers when the clients are autonomous coding agents doing real work, not synthetic benchmark load. Built on the open-source deepseek-v4-flash-mi300x serving sta…
A modular office plugin for DeepSeek Harness (github.com via hn) DSH × Univer Office Give DeepSeek Harness the ability to create, edit, inspect, and deliver spreadsheets, documents, presentations, databases, and canvases. English · 简体中文 dsh-univer-office is the Univer office plugin for DeepSeek Harness…
At Beijing AI-themed bar, DeepSeek tokens come with the pints (www.reuters.com via hn) could not extract summary
Evaluating DeepSeek V4 Pro 0813 on Hack the Box Challenges (theaq.blog via hn) Evaluating DeepSeek V4 Pro 0813 on Hack The Box Challenges The release of DeepSeek V4 Pro back in April got a lot of attention - so DeepSeek decided to release it again! The new version is called DeepSeek V4 Pro 0813, and it is not exactly…
Show HN: LLM-as-a-Verifier Plugin for DeepSeek Harness (github.com via hn) dsh-plugin-llm-verifier A plugin for DeepSeek Harness that adds an LLM verifier: it grades candidate solutions with a model and returns scores between 0 and 1. Based on LLM-as-a-Verifier (paper).
DSHPlugin – DeepSeek Harness plugins with install, compatibility and repo data (dshplugin.app via hn) ModLens liustack/modlens 2.8kGitHub starsA DeepSeek Harness vision plugin that turns pasted images into structured OCR, layout, and semantic evidence for supported text-only models. Find and compare community plugins for DeepSeek Harness w…
Show HN: An n8n-like orchestration toolkit for DeepSeek harnesses (github.com via hn) Open-DeepSeek-Harness-Desktop English | 简体中文 An open-source macOS desktop app for DeepSeek Harness (dsh) — a Codex / Claude Code desktop counterpart built on top of the harness, with no fork of the upstream kernel. It wraps the harness's W…
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks (github.com via hn) LongHorizon-Harness Loop Engineering for Computer-Use Agents Give Claude Code, Codex, OpenCode, or DeepSeek Harness a goal once. Keep it working across desktop apps and the terminal for dozens of hours.
Cosyncing: Control DeepSeek Harness and other agents on your phone (github.com via hn) From CLI to GUI, live and in sync Website · Install · Client · Docs · Contributing · 简体中文 Synchronize and control your agents — from CLI to GUI, from desktop to phone. Pick up right where you left off, anywhere.
Agentic AI costs set to balloon fivefold by 2028 (www.theregister.com via hn) MOST POPULAR AI - ai and ml Anthropic says text watermarking scheme relies on inconsequential words 'Shall I compare thee to a summer's afternoon' is the sort of thing this will make, and others look likely to adopt it - AI and ML DeepSeek…
Simulating Claude-Style Chinese Risk Control and UX in DeepSeek Harness (github.com via hn) dsh-claude-ux Claude 式「区域风控 + 自主结束对话」插件 —— 适用于 DeepSeek Harness 的 web profile。 复刻 Anthropic/Claude 的两类行为,除两个默认关闭的可选外部调用外全部本地判定(详见 docs/PRIVACY.md): 区域风控(可反向):检测目标用户(时区、系统/浏览器语言、中文字体、代理、代理/中转域名黑名单、公网 IP 归属、WebRTC IP 一致性)。regionTarget 选 cn =…
Base-2 vs. base-e log-sum-exp mismatch silently corrupts attention in SGLang (twitter.com via hn) If you serve models with mla on sglang with --attention-backend flashinfer and you get radix cache hits above 8k your outputs are wrong right now. DeepSeek V2/V3/R1/V3.1/V3.2, bailing, glm4moe_lite, Sarvam affected.
OpenAI ditches Recall-style screenshot surveillance for friendly keylogging (www.theregister.com via hn) MOST POPULAR AI - ai and ml Anthropic says text watermarking scheme relies on inconsequential words And other AI model makers are expected to deploy something similar - AI and ML DeepSeek's innovative harness treats everything as a plug-in…
DeepSeek Harness Plugin and Agent Profile Package Index (dsh-index.xlings.org via hn) dsh index — plugins and Agents Browse the DeepSeek Harness plugin ecosystem, and install ready-to-run Agents — Agent = Harness + Plugins/Packages. Install with dsh directly, or through xlings for mirror acceleration.
The official DeepSeek-Harness tested with jelly ball [video] (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Show HN: Multi-version management (by xlings) for DeepSeek Harness (openxlings.github.io via hn) xlings support multi-version management for package, I have add support for dsh. you can try it ## Base Usage xlings install dsh@0.1.0-rc.6 -y && xlings install dsh@0.1.0-rc.3 -y xlings use dsh xlings use dsh 0.1.0-rc.6 && dsh --version xl…
Show HN: Stream – Unlimited DeepSeek v4 Flash 0731 (camelai.com via hn) A flat-rate DeepSeek API for Hermes, OpenCode, OpenClaw, Aider, and other high-volume coding agents.
DeepSeek: What They Invented (claude.ai via hn) Try out Artifacts created by Claude users
Show HN: Lethe is a portable identity layer – user consented (lethe-ai.vercel.app via hn) During my mtech I constantly used to navigate across a bunch of AI apps—from Gamma AI for decks to ChatGPT for understanding and navigation to Cursor, Claude, or Deepseek for coding, reorganizing, and more, as when one or other plans expir…
Show HN: DeepSeek V4 Flash with 7.7 GiB RAM using NVMe demand paging (github.com via hn) DeepSeek V4 Flash on 8 GB RAM — CPU-only, NVMe-backed Run a 78.62 GiB DeepSeek V4 Flash GGUF on a Linux laptop with 7.7 GiB of physical RAM and no GPU, using mmap-backed NVMe demand paging. [!IMPORTANT] The model does not fit entirely in 8…
DeepSeek V4-Flash released: 284B params, 1M-token context, free to use (ainexusdaily.vercel.app via hn) Daily AI Brief 20+ sources — free You've seen some View in every SwiftUI file you've ever opened. Now let's find out what it actually means, why it exists, and why returning a plain protocol doesn't work the same way.
Code harness for DeepSeek V4? OpenCode vs. pi vs. jcode vs. reasonix (grigio.org via hn) Which is the BEST Coding harness for Deepseek V4 ? opencode vs pi vs jcode vs reasonix Same model Deepseek V4 Flash 0731 tested with different AI harnesses, huge RAM usage difference and only 2 harness were able to find a valid solution fo…
Show HN: Dsv4 Codex Proxy – Make DeepSeek V4 Flash 0731 "Codex-Native" (github.com via hn) Dsv4 Codex Proxy Dsv4 Codex Proxy makes DeepSeek V4 Flash 0731 work as a first-class model in Codex. Codex speaks the OpenAI Responses API, while most inference providers are chat-completions native and commonly expose the Responses API on…
DeepSeek V4 Flash 0731: Is It Cheaper to Run It at Home or Pay per Token? (grigio.org via hn) DeepSeek V4 Flash 0731: Is It Cheaper to Run It at Home or Pay Per Token? A note from the reviewer.
Ask HN: DeepSeek Alternatives (news.ycombinator.com) As DeepSeek is planning on raising their prices, I'm looking to explore alternatives. Are there any comparable providers with similar model size that offer the similar pricing as DeepSeek does today?
OpenCode was able to reproduce the current price of DeepSeek (twitter.com via hn) on the upcoming deepseek price increase we've been able to reproduce their current prices even on rented GPUs so this likely isn't because they're "losing money" it's traffic shaping because they are overloaded - Reproduce their prices rea…
DeepSeek-V4-Flash-0731-Latent-Reasoning. A model thinking in latent space (blog.n.ichol.ai via hn) DeepSeek-V4-Flash-0731-Latent-Reasoning. A self-contained model that does thinking in latent space, NVFP4-quantized, with a production vllm form for serving runtime.
DeepSight – give text-only LLMs eyes and hands (zero tokens, on-device) (github.com via hn) DeepSight Give DeepSeek (or any text-only model) eyes and hands. DeepSight connects your existing LLM setup to the real world — it can look at images you send, take screenshots of your desktop, read text on screen, click buttons, type into…
DeepSeek's Liang Wenfeng: Full Remarks from an Investor Meeting (thechatr.ai via hn) Opening Remarks Liang Wenfeng Welcome, everyone. When we first started this company, we were not thinking about how much money we would eventually make, whether we would go to the capital markets, whether we would list, or anything like th…
Show HN: Hardware-software co-design for MoE models to bypass NCCL bottlenecks (github.com via hn) fluidic-expert-fabric (PoC) This repository contains an exploratory Proof of Concept (PoC) investigating a hardware-software co-design approach for distributed Mixture-of-Experts (MoE) architectures, such as DeepSeek-V3 and Mixtral-8x7B. T…
DeepSeek V4 Flash 2.98x faster, lossless (runinfra.ai via hn) Single-stream decode with one variable between the two columns of this comparison. Both columns were measured on 2026-08-01 in one window, one container, one engine install, on the same 4x B200 cards, run back to back with an idle VRAM dra…
Ask HN: Anyone got DeepSeek-v4-flash-0731 running using antirez/ds4? (news.ycombinator.com) could not extract summary
$0.26 DeepSeek V4 Flash 0731 ties $5.01 GPT-5.6 run on Agentic Memory Benchmark (atmbench.github.io via hn) Send a Pull Request Fastest path: open a PR adding a row to the TRACKS array at the bottom of leaderboard.html . Include a short description of the setup and a link to your run logs or code in the PR body.
Show HN: A little physical breakout clone (brontosaurusrex.github.io via hn) Idea: How about a 5 minute game I can play in my browser and it's actually fun to play. Controls: Mouse or keyboard (left/right, left alt/right alt or A/D), space and esc are pause toggle.
Just tried DeepSeek V4 Flash 0731 (she got her brain updated again) (www.reddit.com via hn) could not extract summary
Floatboat DeepSeek Agent – AI with real browser access (deepseek-agent.com via hn) Floatboat DeepSeek Agent is an independent desktop workstation for macOS and Windows that gives DeepSeek access to your files, browser, memory, automation, and reusable workflows.
DeepSeek Founder Liang Wenfeng in His Own Words (www.geopolitechs.org via hn) DeepSeek founder Liang Wenfeng in His Own Words: 64 Quotes from DeepSeek's Investor Call Key takeaways: Vision-driven: No KPIs, no written vision, even "no organization"—it runs on goodwill toward the world and an obsession with AGI. "We'r…
Show HN: DeepSeek API – 50pct cheaper, no minimum, instant key delivery (43.155.207.94.sslip.io via hn) OpenAI SDK compatible DeepSeek API. From $0.70.
Show HN: I built a hypervisor and client for inference on consumer compute (scalattice.com via hn) I'm the founder of Scalattice, this is my second company, third total product. I'm a 2x founder building some challenging software, some easy software, and some curiosity based tools that I've just always wanted to be a part of!
Ask HN: How to use agents via API cheaply? (news.ycombinator.com) I’d like to code with an AI agent via a BYOK open source client. I want to be able to inspect everything, check the system prompt, etc.
We self-host DeepSeek V4 Flash on AWS spot instances (twitter.com via hn) https://t.co/syBl6fkTDJ Miguel Salinas@VercantezHow we self-host DeepSeek V4 Flash on AWS spot instances4:18 PM · Jul 23, 20269.7KViews243560
↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4deepseek
DeepSeek Liang Wenfeng Investor Meeting – Audio to Text Transcript [pdf] (github.com via hn) 梁文锋四小时投资人会议实录. Contribute to demo-zexuan/liang-wenfeng-investor-meeting-2026-7-22 development by creating an account on GitHub.
Four-hour fundraising meeting with DeepSeek founder Liang Wenfeng (old.reddit.com via hn) could not extract summary
An open-source model on par with DeepSeek v4 has appeared in South Korea (old.reddit.com via hn) could not extract summary
An LLM API gateway with budget-limited tokens for your LLM API keys (github.com via hn) This gateway is created to (1) keep real API keys completely away from the client, and (2) provide fine-grained budget control functionality. Real upstream API keys (DeepSeek, OpenAI, Anthropic) live only in the Cloudflare worker's secrets…
Twitter user investigating potential Fable distillation in DeepSeek V4 Pro (twitter.com via hn) 🧵 DeepSeek appear to have engaged, or be engaging in, a large-scale operation to collect outputs from proprietary models (including Claude Fable 5) for certain requests via their API as part of a distillation effort. After seeing such cla…
Fixing Gnome Snapshot Segfault on NixOS and OpenCode with DeepSeek V4 (grigio.org via hn) TL;DR: GNOME Snapshot segfaults (SIGSEGV 139) when the camera portal returns a PipeWire file descriptor. The portal steals the connection's fd, disconnects the client, and hands the dup to Snapshot — which then tries to re-register on a de…
Show HN: A sandbox for running real jailbreak techniques against local LLMs (github.com via hn) LLM Red Team Lab A hands-on kit for educational, authorized red teaming of any locally-run LLM. It works with any OpenAI-compatible model — Llama, Mistral, Qwen, Gemma, DeepSeek R1, and more — and covers the two ways an LLM system gets exp…
↯ Security↯ Llama↯ Mistral↯ Gemma↯ Jailbreakred-teammistraljailbreak+6
Chinese filing implies DeepSeek valuation of around $52B (www.reuters.com via hn) could not extract summary
Baba Is Solved by Fable 5 and GPT-5.6 Sol, but at what cost? (quesma.com via hn) We ported the puzzle game Baba Is You to the Harbor framework, and benchmarked current models, including Claude, GPT, Gemini, GLM and DeepSeek. A human Twitcher is 4x faster than Fable 5.
Show HN: Containerized AI development with cross-compatibility (github.com via hn) I need to study different harnesses CLI, and to find a way to switch from one provider to another. I built a template repository to govern your AI chats with more grip than usual.
Indian companies look to Chinese LLMs as AI costs bite (asia.nikkei.com via hn) Artificial intelligenceIndian companies look to Chinese LLMs as AI costs bite DeepSeek and others winning on price but concerns over foreign reliance loom Indian companies are turning to Chinese artificial intelligence models amid cost pre…
Head to Head: Muse Spark 1.1 vs. DeepSeek-V4-Pro (runtimewire.com via hn) This wasn’t a competitive split decision; it was a rout driven by instruction-following, robustness, and cleaner execution across a wide mix of real tasks. Muse Spark 1.1 consistently delivered answers that were tighter, more compliant, an…
Show HN: Sell your unused AI Credits or buy Claude credits for 50% off (secondhandtokens.com via hn) Buy unused API tokens from other developers at 50% off. Claude, Llama, and DeepSeek — same models, half the cost.
DeepSeek aims to make its own AI chip (www.proactiveinvestors.com via hn) DeepSeek makes pivot that should put Silicon Valley on high alert Published: 02:47 09 Jul 2026 EDT DeepSeek, the Hangzhou-based artificial intelligence startup, is designing its own chip, three people familiar with the matter say. The move…
Show HN: DIMMsum – price tracker and sold-price history for used server RAM (dimmsum.com via hn) My entire homelab is eBay: the servers, the Brocade switches (ICX6450 and ICX6430), the Ruckus APs, the old APC UPS (running on a LiFePO4 pack I swapped in when the lead-acid died), even the patch cables came as a lot of 20. The only new m…
DeepSeek reverses course with surcharge on peak-hour API use (www.scmp.com via hn) After triggering price war, DeepSeek reverses course with surcharge on peak-hour API use DeepSeek triggered a price war in May when it announced a permanent 75 per cent discount on V4 API access The Chinese AI champion will double the pric…
DeepSeek Open Sources DSpark (venturebeat.com via hn) Even as the geopolitical conversation around AI continues to grow more fraught following theU.S. government's actions to limit the new models from Anthropic and OpenAI, Chinese open source darling DeepSeek is back with yet another open rel…
Show HN: No ads and noise from any page, get a clean AI reformat in one click (code.intellios.ai via hn) This is a Chrome extension I created to use deepseek v4 flash (almost FREE) api to help me quickly summarize web page or reformat web page into a much more readable format (remove ads, noises, keep images if needed) Another cool thing is s…
↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4deepseek
Low-cost Chinese AI models like DeepSeek gain traction in the U.S. (restofworld.org via hn) Stu Clott, an operations manager and part-time developer in San Diego, used to code with Claude. But he recently found a cheaper alternative: DeepSeek.
DeepSeek v4 Releasing Mid July (files.catbox.moe via hn) could not extract summary
GLM-5.2: Another open-source Chinese AI model has Silicon Valley's attention (www.businessinsider.com via hn) A new AI model from China is generating the kind of buzz not seen since DeepSeek's R1 announced China as a serious threat to American chatbot hegemony over a year ago. Silicon Valley's online echo chamber has been alight with intrigue in r…
Value for Money Is All You Need (news.ycombinator.com) Value For Money is All You Need A reflection on the future of token consumption in artificial intelligence Token consumption now sits at the center of the growing use of artificial intelligence by businesses and individuals alike. The "Tok…
Show HN: Cc-fleet – run other LLMs as Claude Code workers, your sub drives (github.com via hn) 🚢 cc-fleet 🤖 Plug any third-party model into Claude Code's ⚙️ Dynamic Workflows, 👥 Agent Teams, and ⚡ Subagents — from DeepSeek · GLM · Kimi · Qwen … to your Codex subscription, with your main session's auth untouched; no Claude subscripti…
US holds off blacklisting DeepSeek and more than 100 firms deemed security risks (finance.yahoo.com via hn) has held off adding China’s AI startup DeepSeek, memory chipmaker CXMT and more than 100 other companies flagged as national security risks to a trade blacklist, according to two people familiar with the matter, as the Trump administration…
Native Coding Agent Optimized for Local LLM and DeepSeek v4 with Vector Memory (code.intellios.ai via hn) cwcode A terminal coding agent built around DeepSeek V4 Pro, Qwen3.6‑27B, Kimi, Azure, and anything else that speaks OpenAI’s chat API. Written in Go.
Show HN: Chrome Extension That Removes AI Slop / Spam / Self-Promo from Reddit (nobot-two.vercel.app via hn) I lately was really frustrated about all the ai content on reddit. I feel like every second post is just a bot trying to farm engagement.
China cracks down on Western AI models while US companies flock to DeepSeek (www.techradar.com via hn) The great AI Irony: China cracks down on Western models while US companies flock to DeepSeek AI for me but not for thee? - China continues to purge both demand for AI chips from its ecosystem and foreign AI models, citing 'security risks'…
Running DeepSeek-V4-Flash on a Raspberry Pi (twitter.com via hn) Article Conversation Running DeepSeek-V4-Flash on a Raspberry Pi I ran DeepSeek-V4-Flash on a Raspberry Pi 5 (8GB edition) by streaming model weights from a PCIe attached NVMe SSD. Codex (GPT-5.5 xhigh) and Claude Code (Opus 4.8 max) drove…
DStudio – local DeepSeek V4 with a design studio, reachable from your phone (github.com via hn) DStudio A native, local-first desktop app for DeepSeek V4 — chat, a coding agent and a design studio, all running on your Mac. Nothing leaves the device.
DeepSeek Made AI Cheap. Now It Needs Billions to Keep It Cheap (chinacompany.substack.com via hn) DeepSeek Made AI Cheap. Now It Needs Billions to Keep It Cheap.
Mimo v2.5 is better deal than DeepSeek v4 flash (news.ycombinator.com) So Hear me out. Not only on almost all benchmarks is mimo v2.5 is better than dsv4f flash, but also the pricing.
DeepSeek V4 managed to reverse engineer Teamspeak's Licensing System with $3.88 (old.reddit.com via hn) could not extract summary
TokkeyCC – OpenAI-compatible API for 100 AI models, .22 per 1M tokens (tokkeycc.com via hn) 100+ Models · Unified API · OpenAI Compatible Access OpenAI, Anthropic, DeepSeek, Meta, Google, and more — through a single, unified API. Switch models instantly.
Lots of people want to try Claude Opus 4.8 (wisgate.ai via hn) Access multiple AI models through one unified API. OpenAI, Claude, Gemini, DeepSeek and more.
DeepSeek-V4-Flash (official FP8) running across 2x DGX Spark (forums.developer.nvidia.com via hn) I didn’t create this recipe you guys did but I was finally able to find it and get Deepseek v4 Flash working with 200k Context on 2 Nodes. Sharing this since I couldn’t find a confirmed end-to-end recipe for the official DeepSeek-V4-Flash…
ik_llama.cpp – llama.cpp fork with better CPU performance (github.com via hn) ik_llama.cpp: llama.cpp fork with better CPU performance TL;DR This repository is a fork of llama.cpp with better CPU and hybrid GPU/CPU performance, new SOTA quantization types, first-class Bitnet support, better DeepSeek performance via…
Did DeepSeek v4 suddenly become more expensive? (imgur.com via hn) If you're seeing this message, that means JavaScript has been disabled on your browser , please enable JS to make Imgur work.
↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4deepseek
Show HN: Train Claude Code's replacement (ds4 and pi and aoe) (github.com via hn) Remember how Meta monitored employee activity closely for a few months, and then had a bunch of layoffs related to AI efficiency? (oh right that was like 3 days ago).
↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4deepseekclaude-code
DeepSeek Slashes AI Costs to Cents (businessanalytics.substack.com via hn) DeepSeek Slashes AI Costs to Cents Edition #299 | 29 May 2026 DeepSeek Makes 75% Price Cut on V4 Pro Permanent, Dropping Frontier-Class Inference to $0.87/M Output Tokens with Mixture-of-Experts Architecture In this edition, we will also b…
Show HN: SharkBay – a local macOS workbench for coding-agent CLIs (github.com via hn) SharkBay macOS workbench for multi-agent vibe coding Features Multi-Agent Support Launch and manage multiple AI coding agents from one workspace. Supported agents: Claude Code · Codex · Gemini · Kiro · DeepSeek · Qwen · OpenCode Agent Stat…
DeepSeek V4 Flash at 8.4 tok/s on 3×3090: patching the GGUFs that won't load on cchuter's llama.cpp fork (www.reddit.com) my apologies if anything does not make sense, I literally dont know what I am doing, im not a programmer, just a simple vibe coder, with an Claude subscription. That said, if you have 200gb of sys ram+vram and want to run deepseek v4 flash…
GH200 NVL2 or 8x RTX 6000 Blackwell for running Kimi K2.6 / DeepSeek V4 locally? (5 devs, agentic coding) (www.reddit.com) Trying to figure out the right box for my team and wanted to see if anyone had any clue which would be a better fit or if it is not worth our time in our budget. Situation: 5 of us doing agentic coding (lots of long context getting re-sent…
Lower Bracket Context Tax: An Open MCP Persistent Memory Layer That Limits Agent Context Bloat to 10% (www.reddit.com) Because standard coding agents are stateless, every session they start from scratch. I built Zerikai_memory around a different model: you decide when the agent learns your codebase, not the other way around.
Fused MoE dispatch kernel in pure Triton: 89-131% of Megablocks, runs on AMD with zero code changes (www.reddit.com) I've been working on MoE inference and wrote a fused dispatch kernel entirely in Triton, no CUDA. At inference batch sizes (up to 512 tokens) it reaches 89-131% of Megablocks(Stanford's CUDA-optimized MoE lib), and the same kernel runs on…
Build an agent capable of complex programming tasks in under 100 lines of code. (www.reddit.com) The code below is an interactive agent capable of handling complex tasks, built in under 100 lines of code using huko-engine. If you just want to drop some agentic features into your existing app, it only takes 20 lines.
Best AI Agent Setup - Hermes + Deepseek-v4-flash? (May 2026) (www.reddit.com) Used to use claude code for everything. I burned 10-20 Billion opus tokens at work, and wanted to use agents for personal projects.
Show HN: I built a tool to estimate AI agent costs before you ship (airunrate.com via hn) Free AI agent cost calculator. Compare GPT-4o, Claude, Gemini, DeepSeek and 50+ models.
DeepSeek seems to be leaking random user chat history (breakingvibe.dev via hn) DeepSeek’s web chat interface appears to be leaking conversations between accounts. When opening the chat, a user encountered chat histories that did not belong to them, suggesting a session isolation failure on DeepSeek’s servers.
Are LLMs the New Propagandists? (www.reddit.com) I was brainstorming about a video with Claude (Sonnet 4.6). It suggested to explain the difference among ChatGPT, Gemini, Claude and DeepSeek.
DeepSeek-V4 KV Cache Explained: Why 1M Context Uses Less VRAM (knightli.com via hn) The real cost of long-context models is often not whether they can accept one million tokens, but how much VRAM the KV Cache consumes during inference. During Transformer decoding, every newly generated token needs access to the Key and Va…
Ask HN: What is your daily AI stack? (news.ycombinator.com) Which AI tools benefiting you most in day-day work? Which tool you stopped using?
Is Composer 2.5 better than Glm 5.1 and DeepSeek v4 pro in real world tasks? (www.reddit.com) I am new to Cursor and still testing the free version. Benchmark for Composer 2.5 indicates it is better than DeepSeek v4 and Glm 5.1.
DeepSeek just popped the American AI bubble. (www.reddit.com) DeepSeek just popped the American AI bubble. Not by killing AI.
Performance When Offloading Large Models to System RAM? (www.reddit.com) I noticed for people running large models, or those that would be cost prohibitive to have all in GPU VRAM, I noticed that the dominate strategy is one GPU with a large pool of system DRAM to offload the weights, as per GB VRAM is always m…
$340 opus bill made me rethink how I route agent tool calls (www.reddit.com) Looked at my coding agent's bill last month: $340 for repo maintenance across three repos, each around 15k lines. Most of those tool calls were just grep and file reads.
ml intern skill instead of gsd (www.reddit.com) - designed for ml workflows - works autonomously for hours Projects fully done with this skill - flash attention for volta (very old GPUs) https://github.com/AlexWortega/flash-attn-volta - deepseek 4 full replication + training on runpod +…
Local compression helps (www.reddit.com) Just wanted to post a tip (I'm human, not an agent, watch: fart). I use Deepseek-v4-Flash on a lot of my agent work, and as I'm learning and testing these things.
DeepSeek-V4-Pro 75% off discount is now permanent (twitter.com via hn) DeepSeek @deepseek_ai We are making our discount permanent! Enjoy building with DeepSeek-V4-Pro and bring your innovative ideas to life!
Show HN: Myco – coordinate Claude and DeepSeek and other LLMs in one agent swarm (github.com via hn) myco 🇧🇷 Leia em português --> Problem: running multiple Claude Code sessions in parallel produces conflicts, repeated work, and stale assumptions — the agents have no shared awareness. myco is a text-only coordination protocol + tiny Pytho…
The Special Token `<Think>` Problem/Bug of Latest DeepSeek LLM (www.pixelstech.net via hn) www.pixelstech.net Performing security verification This website uses a security service to protect against malicious bots. This page is displayed while the website verifies you are not a bot.
Async Python client for private DeepSeek API (github.com via hn) aiodeepseek A high-performance async Python client for the private DeepSeek API. Supports streaming, image uploads, multi-turn conversations, and new account registration.
QuickSilver Pro – OpenAI-Compatible Platform for DeepSeek V4 and Qwen (quicksilverpro.io via hn) OpenAI-compatible API for 7 top open-source LLMs — DeepSeek V4 Flash & Pro, V3, R1, Qwen3.6 & 3.5-35B-A3B, Kimi K2.6 — 20% cheaper than OpenRouter, Together AI, Fireworks. One-line drop-in.
Anyone compared gpt-5.4-nano vs deepseek v4 flash? (www.reddit.com) They seemed to lie in (almost) similar pricing(i know still quite different on output) Pricing Model Input (1M tokens) Output (1M tokens) DeepSeek V4 Flash $0.19 $0.51 DeepSeek V4 Pro $1.74 $3.48 gpt-5.5 $5.00 $30.00 gpt-5.4 $2.5 $15 gpt-5…
Open AI compatible API in Cursor (www.reddit.com) Hey. I have been experimenting with new models in my Cursor.
Glia – Local-first shared memory layer (SQLite-vec + FTS5 + Offline Knowledge Graph) (www.reddit.com) Hey everyone, I wanted to share a project I've been working on called Glia. It is a 100% offline, local-first RAG and memory layer designed to connect your AI web chats (Claude, ChatGPT, DeepSeek) with your local developer tools (Claude Co…
How to use DeepSeek V4 PRO in Cursor? Skill issue on my side (www.reddit.com via reddit) could not extract summary
🧬 flux-genotype: A self-evolving AI kernel that runs on CPU with Ollama — mutates its own architecture (www.reddit.com) `🧬 Flux‑Genotype – A CPU LLM that rewrites itself` I've been working on an open-source kernel called **flux-genotype**. It orchestrates local models (TinyLlama, Llama 3.2, Hermes 3, DeepSeek-Coder) into a self-modifying ecosystem.
Ask HN: Which AI harness comes close to Claude Code? (news.ycombinator.com) I really want to try deepseek V4, but harnesss which I have previously used are inferior than Claude Code. Please suggest some Harnesses here.
What is a good app for using the Claude API with attached files? (www.reddit.com) I use Obsidian to keep track of my Markdown files. There are various plugins to have it interact with Claude.
Deepseek V4's 1M context window: the breaking point (www.reddit.com) Just ran to verify deepseek v4's context claim of 1M and ran it across three production codebases like 45k (microservice), 180k (monorepo backend) and 520k(full stack app). For the observation, tasks included dependency tracing, cross file…
LLM Phone Home: Reliable Apps that can deliver inference from local backend (www.reddit.com) Hello all, I’m wondering what suggestions there are for an ios app that can serve an openai compatible endpoint. I am using 3sparks which works GREAT for that specific use, BUT, there is no mcp, no web search, etc.
DeepSeek-V4-Flash means LLM steering is interesting again (www.seangoedecke.com via hn) DeepSeek-V4-Flash means LLM steering is interesting again Ever since Golden Gate Claude I’ve been fascinated with “steering”: the idea that you can guide LLM outputs by directly manipulating the activations of the model mid-flight. DeepSee…
Recent Developments in LLM Architectures: KV Sharing, MHC, Compressed Attention (magazine.sebastianraschka.com via hn) Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention From Gemma 4 to DeepSeek V4, How New Open-Weight LLMs Are Reducing Long-Context Costs After a short family break, I am excited to be back and catching up o…
Has anyone found a Qwen CLI replacement? (www.reddit.com) I just need 1 or 2 people to reply to me with the answer I need. I have not been able to keep up with AI advancements for a while.
A message from kurdistan – my love for China and DeepSeek (old.reddit.com via hn) could not extract summary
One of the things I don't see people listing as benefit of hosting local LLMs is on demand usage. (www.reddit.com) Seriously, It might be obvious fact, but when you are on subscription you kinda are in pressure to keep using it otherwise the unused limits feel like wasted potential. There is this urge to keep maximising the tokens you paid for even if…
Deepseek Now Limits File Attachements (www.reddit.com) Nerfing the usage modalities without any announcement seems to be the norm nowadays. Even Chinese AI vendor Deepseek limits the usage of their best expert model for free users.
DeepSeek V4: The Open-Source Model Frontier Labs Feared (helloai.com via hn) DeepSeek V4: The Open-Source Model Frontier Labs Feared DeepSeek V4 ships under MIT with $0.30/M output tokens — 83x cheaper than Claude Opus 4.7 — while scoring 80.6% on SWE-bench Verified. The agentic-coding price floor just moved an ord…
We Tested DeepSeek V4 Pro and Flash Against Claude Opus 4.7 and Kimi K2.6 (blog.kilo.ai via hn) We Tested DeepSeek V4 Pro and Flash Against Claude Opus 4.7 and Kimi K2.6 DeepSeek V4 Pro and DeepSeek V4 Flash launched together on April 24, 2026 under MIT license. They are DeepSeek’s first new architecture since V3, and their first ope…
I built a desktop app that routes Claude Code to any LLM: DeepSeek, Ollama, Copilot, OpenRouter, and 7 more (www.reddit.com) Claude Code is the best AI coding tool I've used. But being locked to one provider, one pricing model, and one model catalog always bothered me.
Multi-LLM AI trading agent harness (github.com via hn) 1rok 1rok is a standalone harness for running portfolio-construction agents across OpenAI, Anthropic, Gemini, xAI, DeepSeek, GLM, and OpenRouter against the same financial tool surface. Agents query Alpaca, Yahoo Finance, FRED, and Tavily…
DeepSeek and Grok hallucinated the same fictitious OpenBSD manpage quote (stuart-thomas.com via hn) Adversarial LLM Review with Hallucination Detection in Solo Security Research A single-day case study of three filings, fifteen refutations, and the manpage that wasn’t Independent Security Research — Whitby, North Yorkshire, United Kingdo…
I offloaded bulk file reading from Claude Code to a cheaper model for a week. Here are the numbers. (www.reddit.com) Hey r/ClaudeAI — I use Claude Code a lot, and I noticed I was wasting a surprising amount of my usage limit on stuff that was basically just reading. Big files, long diffs, Jira/Linear tickets with comment history, docs pages, repo spelunk…
Nemotron-Cascade 2: Post-Training LLMs with Cascade RL (research.nvidia.com via hn) We introduce Nemotron-Cascade 2, an open 30B MoE model with 3B activated parameters that delivers best-in-class reasoning and strong agentic capabilities. It is the second open-weight LLM, after DeepSeek-V3.2-Speciale-671B-A37B, to achieve…
DeepSeek V4 from the Inside (idlemachines.co.uk via hn) DeepSeek V4 introduces three core new architectural ideas, shows the Muon optimiser works at the trillion-parameter scale, and updates the number formats around them. Thankfully for us, their reference code is open source, for inference an…
OpenCode + DeepSeek V4 Pro vs Claude Code CLI?🤔 (www.reddit.com) Im rather new to the whole Agentic automation AI's but Im hearing people with vibe coding were able to pull big unique projects they wouldn't be able to do by themselves or possibly needed to pay a huge fund to programmers, designers, etc.…
Which Chinese Model is best for planning and which is best for implementation? I'm currently using Opencode with an Openrouter API Key, mostly wanna decide between Kimi, GLM, DeepSeek, Qwen, Minimax and Mimo (www.reddit.com) Original plan was to use Kimi/GLM for planning and DeepSeek for implementation, but seeing a lot of love for MiMo and Minimax lately. Anyone running a planner + coder split on Opencode?
DeepSeek Rejects Alibaba: Prioritizing Corporate Independence Over Big Tech Ecosystems (www.reddit.com) In April, DeepSeek launched a rare, massive financing plan that attracted interest from two of China’s largest tech giants: Tencent and Alibaba. However, we have exclusively learned that recent negotiations between Alibaba and DeepSeek hav…
ZAYA1-8B: An 8B Moe Model with 760M Active Params Matching DeepSeek-R1 on Math (firethering.com via hn) Who should care If you work with math, science problems, or complex coding tasks and you're looking for something small enough to run locally or cheaply via API, this is worth serious evaluation. The benchmark numbers at 760M active parame…
DeepSeek-v4-Pro and Hermes: Unauthorized Modification of Security Controls (www.eddieoz.com via hn) Deepseek-v4-pro + Hermes: Unauthorized Modification of Security Controls This article documents a specific, real incident. It exposes a class of vulnerability that deserves attention: the unsupervised mutability of security rules by autono…
CodexSaver Make Codex cheaper without making it dumber with DeepSeek (github.com via hn) CodexSaver Make Codex cheaper without making it dumber. 中文文档 CodexSaver is an MCP tool that turns Codex into a cost-aware router.
I wasted 3 days rewriting prompts for our agent before realizing the whole architecture was garbage (www.reddit.com) We run a small content-monitoring agent for our growth team. Nothing fancy on paper.
DeepSeek could be valued at up to $50B in first fundraising (www.reuters.com via hn) paywalled
I plan to use a chinese AI model through API for coding through a harness, I'm a uni student so nothing prod related for now. should i go deepseek, minimax, kimi or glm? kinda confused (www.reddit.com) Just cancelled my claude subscription due to poor rate limits, gemini cli doesn't really excel in coding from my personal experience, and my local hardware isn't that powerful to run local AI models, and while codex is good, I wanna try so…
Does Deepseek V4/Flash work with Llama CPP and Vulkan on and branches yet? (www.reddit.com) Even unofficial or slow. I have enough vram-memory to load it, but not enough memory to run in cpu-only mode.
DeepSeek V4 being 17x cheaper got me to actually measure what I send to cloud vs what I could run locally. the results are stupid. (www.reddit.com) That foodtruck bench post showing deepseek v4 matching gpt-5.2 at 17x cheaper got me thinking. if frontier cloud models are that overpriced for equivalent quality, how much of my daily work even needs cloud at all?
I built vivkemind – an open-source, local‑first terminal AI coding agent with full AWS Bedrock support (www.reddit.com) wanted a terminal AI coding agent that doesn't lock me into one model provider. So I forked Qwen Code and added full support for every model available in AWS Bedrock.
Show HN: Token Usage Meter 12 Providers and Coding Agent (qlaud.ai via hn) Here once again A Token Usage Meter for 12+ AI Providers Anthropic, OpenAI, Google, Alibaba qween, Moonshot Kimi, MiniMax, ElevenLabs, Deepgram, Perplexity. Qlaud.ai provides token usage meter / AI billing layer.
Struggling with Qwen3.6 27B / 35B locally (3090) slow responses, breaking code looking for better setup + auto model switching (www.reddit.com) Hey everyone, I’ve been experimenting with running Qwen models locally on my setup: GPU: RTX 3090 (24GB VRAM) RAM: 64GB CPU: Ryzen 5700X OS: Windows 11 What I’m currently running Qwen 3.6 35B (UD Q4_K_M) llama-server.exe -m "C:\Users\Dino\…
Questions about revisiting local LLM roleplay. (www.reddit.com) TLDR for those that dojr wanna read below I need a new good free place online to pickup roleplay where should that be and what can I do locally? 9070xt 32gb ram desktop and preferably but I know it not great, 4060 laptop 32gb ram.
AGENTS.md trick that stopped Codex from doing dumb work at premium rates (www.reddit.com) Spent a Sunday auditing where my Codex tokens were actually going. Half the calls were stuff like "rename these 12 fields", "format this csv as markdown table", "extract the dates from this changelog".
DeepSeek's Sequel (messaging-custom-newsletters.nytimes.com via hn) Validation error: uri: Required
I built an open-source desktop app that lets AI control your browser for you (www.reddit.com) Hey everyone, I've been working on Autai — an open-source desktop app (Electron + React) that uses AI agents to automate your browser. You just type what you want in plain English, and the AI opens a real browser and does it for you.
Local LLM Benchmark about Backend Generation by Function Calling (GLM vs Qwen vs DeepSeek) (www.reddit.com) Detailed Article: https://autobe.dev/articles/local-llm-benchmark-about-backend-generation.html Five months ago I posted the "Hardcore function calling benchmark in backend coding agent" thread here. As I wrote in that post, it was an unco…
↯ Glm↯ Function Calling↯ Sonnet 4.6function-callingglmgpt-5+3
CAISI Evaluation of DeepSeek V4 Pro finds it to be on par with GPT-5 (www.nist.gov via hn) In April 2026, the Center for AI Standards and Innovation (CAISI) evaluated the open-weight AI model DeepSeek V4 Pro (“DeepSeek V4”). CAISI evaluations indicate that DeepSeek V4’s capabilities lag behind the frontier by about 8 months (Fig…
127³ — Superintelligence, public. DeepSeek V4 Pro (deepseek-v4-pro-127cubed.vercel.app via hn) DeepSeek V4 Pro 127³ 127-stratum crystalline lattice on DeepSeek V4 architecture. 1.6T params · 49B activated · MoE · 1M context · MIT license.
DeepSeek v4, and the end of the OpenAI/Microsoft AGI clause (simonw.substack.com via hn) DeepSeek v4, and the end of the OpenAI/Microsoft AGI clause Plus LLM 0.32a0 In this newsletter: DeepSeek V4 - almost on the frontier, a fraction of the price Tracking the history of the now-deceased OpenAI Microsoft AGI clause LLM 0.32a0 i…
Filed two PRs for SGLang which may help others too — FP8 KV cache corruption and memory leak on image requests (www.reddit.com) We run Qwen3.6-27B-FP8 at AI Router Switzerland and hit two issues, so I wanted to share in case anyone else runs into them. FP8 KV cache produces silent garbage output with radix cache prefix hits (PR #24198 — ✅ approved) We were running…
I bypassed DeepSeek's censorship filters with a fictional planet trick, here's what happened. ( via reddit) could not extract summary
After seeing deepseek refused to acknowledge Taiwan is a coutry I had to do a little experiment (www.reddit.com) could not extract summary
Comparing SVG Generation for the top open models (codeinput.com via reddit) Some of the larger models (like Llama) weren't available on OpenRouter, so I had to work with what was there. Best small model: Gemma 4 26B For its size, I think it had the best output.
From 5 Hermes profiles to an actual team: the missing piece was memory boundaries (www.reddit.com) I've been messing around with Hermes for months, and quickly outgrew using it just as a fancy CLI assistant. My goal was to build a persistent, specialized team of local agents that could collaborate on long-term projects without me spoon-…
Ask HN: Are you OK with DeepSeek and other labs reading your data? (news.ycombinator.com) If you provide read permission on a directory, might the content be exfiltrated and used commercially? Or is this just paranoia?
I built a full web app using Qwen 3.6-35B running locally on my 5070 Ti with the BMAD Method — here's how it went (ggufbench.com via reddit) I've been running local LLMs since Qwen 3.5 dropped and I was really impressed by what we could run on consumer hardware. Fast forward another two months and we have gotten a handful more gems such as Gemma 4 and Qwen 3.6, so I wanted to p…
A 3D Flappy Bird side-scroller game built with DeepSeek V4 Pro (www.annajc.com via hn) FLAPPY ANNA 3D PRESS SPACE OR TAP Presented by Guan, Made in Melb with DeepSeek and Love GAME OVER PRESS SPACE OR TAP 0000
100M tokens for $2.65 (Deepseek V4 Pro) (www.reddit.com) This is actually unbelievable. I am shocked that there has not been a move in the market like it did last year with the R1 release.
I built a hands-free voice AI that sends emails mid-conversation — and that's just one feature. Here's everything AskSary can do. (www.reddit.com) https://reddit.com/link/1symbsj/video/fti7rujjn1yg1/player Been building AskSary solo for a while. Just shipped hands-free voice email - you're mid-conversation with an AI and you say "send an email to [john@example.com](mailto:john@exampl…
Is paying for deepseek v4 pro worth it or are there better alternatives (www.reddit.com) Guys is deepseek v4 pro really the best model (price to performance) because i was using nvidia apis for two weeks in opencode then suddwnly everything stopped working so i am thinking to opt for the payed (yet very affordable) option to m…
LLM Budget Guard – open-source runtime cutoff for OpenAI/Anthropic (www.llmeter.org via hn) Alerts won't stop your 3 AM token spiral LLM Budget Guard enforces hard cutoffs at the provider — across OpenAI, Anthropic, and DeepSeek — before runaway agents burn $47K in 11 days or get your account terminated. Founding-team pricing loc…
DeepSeek V4 PRO on how many 3090 ? (www.reddit.com) Hi guys I got only 3090 GPUs so... How many prefer to run to get a great result in DeepSeek V4 PRO?
Show HN: Another experiment with an Erdos problem and LLMs (news.ycombinator.com) Background: I am a coder, not a mathematician, but I was quite entertained by this story: https://news.ycombinator.com/item?id=47903126 I wondered how far I could get by just choosing a random open problem and throwing it at LLMs. Disclosu…
Language Anchoring: A Systematic Method for LLM Multilingual Adaptation (github.com via hn) fkyah3/opencode-fkyah3 DeepSeek 优化 · Windows 适配 · AI 实现 🚀 从零搭建指南(中文) · English · 繁體中文 本项目是 anomalyco/opencode 的个人 Fork。所有修复、优化、功能均由 AI 完成——DeepSeek V4 Flash (thinking mode) / Sisyphus——在人类监督下执行。 上游是优秀项目。Windows 和 DeepSeek 并非他们的优先方向。我们自行处理。…
Deepseek v4 flash weird sizes? (www.reddit.com) So I'm sure everyone is excited about the new deepseek release(s) but I'm a little confused about it's vram requirements. a q4 gguf of it is only 120gb?
DeepSeek's new models are so efficient they'll run on a toaster by which we mean (www.theregister.com via hn) DeepSeek's new models are so efficient they'll run on a toaster ... by which we mean Huawei's NPUs Now available in preview, DeepSeek V4 cuts inference costs to a fraction of R1 Chinese AI darling DeepSeek is back with a new open weights l…
anyone actually tried deepseek v4 pro for coding? (www.reddit.com) so v4 pro dropped and barely anyone is talking about it. feels weird since when kimi k2.6 came out i seen post about it everywhere anyone here tried v4 pro for actual code work?
DeepSeek V3.2 looping bug: what settings / harness tweaks are actually reducing it in production? (www.reddit.com) I’m trying to isolate the looping / repetition issue some people have been reporting with DeepSeek V3.2 around April 2026, especially in agentic or tool-use setups on hosted providers like OpenRouter and SiliconFlow. Public model pages des…
DeepSeek V4 plays Go on a 9x9 board (chat.deepseek.com via hn) We need to create a prompt that sets up a fresh Go game session. The user wants to "export a go board prompt for a new session," meaning they want a prompt they can copy-paste into a new chat to start a game of Go with me, presumably with…
is Deepseek v4 unvailable in Cursor? I cannot see it. (www.reddit.com) It seems that Cursor removed all the DeepSeek models. I find it limiting, considering it seems performant.
DeepSeek's new model is 75% off right now, here's how to take advantage (www.reddit.com) TL;DR and rundown DeepSeek v4 released this week and performs close to frontier models like GPT/Opus on benchmarks. It's available now and is discounted by a whopping 75% through their API until May 5, making it the most cost effective hig…
DeepSeek-V4 on Day 0: From Fast Inference to Verified RL with SGLang and Miles (www.lmsys.org via hn) DeepSeek-V4 on Day 0: From Fast Inference to Verified RL with SGLang and Miles We are thrilled to announce Day-0 support for DeepSeek-V4 across both inference and RL training. SGLang and Miles form the first open-source stack to serve and…
I had no way to check how LLMs see my SaaS or my clients',so I built BrandGEO.co (brandgeo.co via hn) See exactly how ChatGPT, Claude, Gemini, Grok & DeepSeek talk about your brand — and what to fix. A free 2-minute audit scores your brand across 6 dimensions on all 5 AI engines, then hands you the top priority actions to take next.
Show HN: I built a coding agent that works with 8k context local models (github.com via hn) Most AI coding agents assume you have a 200k-context model. In reality, the local models most people actually use have 8k windows — barely enough for one large file, let alone a whole project.
Need recommendations on embedding models (www.reddit.com) Each LLM vendor's API has a distinct personality separate from the model itself. 6 months of prod agent dev made me believe this (www.reddit.com) How to setting Deepseek in librechat (www.reddit.com) Hi everyone. I've tried everything to view Deepseek in Librechat, but I can't.
Feedback on iOS app with local AI models (www.reddit.com) Hey everyone, I just shipped an iOS app that runs local AI models. Current has 12 models: Gemma 4, Llama 3.3, Qwen3, DeepSeek R1 Distill, Phi-4, etc.
For AI agents: is per‑token pricing killing your budget? Looking for feedback on time‑based subscriptions. (www.reddit.com) Hey r/AI_Agents, I run an inference service (cheapestinference.com) and we're exploring a different pricing model that might be more predictable for agent workloads. Instead of per‑token billing, we offer **dedicated 8‑hour time windows**…
I built an MCP server that gives Claude Code image/video generation, web search, and smart multi-model routing (www.reddit.com) I built mcp-multi-model — an open-source MCP server that extends Claude Code with capabilities it doesn't have natively. **What it does:** - Generate images and videos right in the terminal (via Gemini Imagen & Veo) - Smart routing: resear…
Claude should offer first-party support for open-weight models Anthropic hosts via partnerships, hear me out. (www.reddit.com via reddit) Imagine how badly distillers and Chinese hosts undercutting Claude would shit their pants if Anthropic offered high throughput first-party support via Cerebras or another wafer-based host with DeepSeek 4.1f, GLM 5.3,.etc. so everyone looki…
codemap - packs your repository into focused, AI-ready context (www.reddit.com via reddit) codemap helps you prepare better context for GPT, Claude, DeepSeek, GitHub Copilot, and local AI models. Choose which files to include, add Git history or reusable Skills, control output size, and safely review or apply changes returned as…
Computer use was too expensive for my daily tasks, so I built a hybrid browser agent (www.reddit.com via reddit) I tried computer-use agents with Astra and Fable 5.1 to automate tasks on my PC. The problem was simple: the cost didn’t make sense for the repetitive browser work I wanted to delegate.
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression (arxiv.org) The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV c…
About the AI race ask to stop: I don’t want smarter models (www.reddit.com via reddit) I don’t want smarter models, I want Fable/Astra models running for the price of a Haiku. It seems absurd, I know, but why not stop investing money in making AI smarter and start investing in making it cheaper?
I have to say something as a chinese (www.reddit.com via reddit) You need to know, as a chinese, we be teached don't say that much since we were kids or you will get in trouble unless you have to. As someone familiar with China’s AI community, I have a few thoughts on Anthropic’s recent accusations.
I wanted my chats to talk to my chats so I made a chat (macOS) (www.reddit.com via reddit) For months I've been using a CC/Codex plugin (agent-talk) to use GPT as an adversarial reviewer for Claude. Then I recently switched to herdr -- very cool terminal session server/multiplexer.
What's the minimum architecture a modern agent harness actually needs? (www.reddit.comhttps) I’ve been trying to understand agent harness at the system level, rather than just building another coding agent. I’m still relatively new to this field, so I picked a concrete project to learn from —— DeepSeek Harness (DSH), because it’s…
[AINews] DeepSeek v4.1-Flash: 763B-P8B-D16B novel causal Encoder–Decoder architecture with vision marks the Return of the Whale (www.latent.space) We are late to this but better than never. Have been busy finalizing the second AIE NYC, which is happening in one month.
Do output compression tools still matter with newer AI models? (www.reddit.comhttps) We tried to answer this question by validating one of the most popular tool in this area: RTK (Rust Token Killer). We used Claude Code with Fable 5.0, and OpenCode with DeepSeek V4 Pro 0813 through OpenRouter on Terminal-Bench 2.1 Research…
A Claude Code skill pushed DeepSeek V4 Flash from 67.42% to 82.02% (www.reddit.comhttps) Autoprompt closes much of the manual coding loop by planning, building, testing, reviewing, and repairing from one prompt. Autoprompt v2 is now out, with support for Claude Code and 10 other coding tools.
DeepSeek v4.1 does the same coding task for $.04 while Fable 5.1 costs $3.64 - LiveBench (www.reddit.comhttps) Anthropic, you better get your shit together
HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention (arxiv.org) Token-level sparse attention mechanisms, exemplified by DeepSeek Sparse Attention (DSA), achieve fine-grained key selection by scoring every historical key for each query through a lightweight indexer, then computing attention only on the…
Claude is still the best value for your money (www.reddit.com via reddit) So my comparison per the rules is hedged only on my personal experience. I'm a senior university student who has been trying out different AI models and seeing its effectiveness on some of the similar tasks that I am most likely to perform.
RedKnot-MLA: Multi-Head Offline-Online Reuse for DeepSeek-V4 Long-Context Serving (arxiv.org) Multi-head latent attention (MLA) exposes many logical query heads through one packed latent KV stream. This representation is memory efficient, but it removes the physical per-head cache boundary assumed by conventional head-wise reuse.
Gave 6 AI models the same bug. Only 3 got it right. (www.reddit.com via reddit) Tried another little AI test today. I gave ChatGPT, Claude, Gemini, Grok, DeepSeek and Qwen the exact same coding bug and asked them to fix it.
Serious Nonfiction (www.reddit.com via reddit) I’ve been working on a manuscript for a serious nonfiction book for about 8 months. I’ve experimented with ChatGPT, Claude, Gemini, and DeepSeek.
How to set reasoning level for custom OpenRouter model in Cursor? (www.reddit.com via reddit) I'm using a DeepSeek model via OpenRouter in Cursor. Cursor shows a reasoning level selector (low/medium/high) for its built-in models, but not for my custom OpenRouter model.
How to replicate Claude Code's agentic workflow without the crazy rate limits? (www.reddit.com via reddit) It’s frustrating to see usage get choked down this hard, especially given the company plans to IPO soon. I am on a pro plan that on an avg gave 3-4 hrs a day which worked for me until recently.
Voice for DeepSeek Harness: the interruption problem, and how I ended up measuring the echo instead of guessing a threshold (www.reddit.com via reddit) The problem. I write most of my code by voice.
PersianAnonymizer: Evaluating LLM-Labeled Training for Efficient NER-based Anonymization in Persian (arxiv.org) We target practical anonymization of Persian customer chats by training a compact NER model from LLM-labeled supervision and selecting the best labeler for deployment. We compare three instruction-tuned LLMs: DeepSeek-V3-0324, GPT-OSS-120B…
Deploying DeepSeek 175B Locally on a Single Consumer-Grade RTX 4060 Laptop with 32GB RAM for 200k-Scale Protein-Ligand Virtual Screening (arxiv.org) Recent advances in large language models (LLMs) have demonstrated exceptional performance in protein-ligand interaction prediction, but state-of-the-art pipelines for large-scale virtual screening almost exclusively rely on high-end GPU cl…
Made my Codex limits last almost ~3x longer with one change (www.reddit.com via reddit) Plus users are basically being forced to give up Sol and just use Luna to get any usable amount of work done. That's a huge downgrade basically using a deepseek flash model level which you can get for free in opencode anyway.
Benchmarked the free API tiers you can point a coding agent at - half of them now want a card (www.reddit.com via reddit) I run aider and Cline against free tiers instead of paying per token, and my setup broke twice this month when model IDs disappeared under me. So I stopped guessing and measured what's actually left.
Claude gets surprisingly hostile when I try asking about other models (www.reddit.com via reddit) I am trying to set up an email triage setup for my needs. After a bit of discovery I started exploring using Hermes Agent as an option with Claude.
Auto compact strategy? (www.reddit.com via reddit) Im using DeepSeek Harness, since its open source i customized it to my liking, compaction had a lot of bugs and issues that i fixed but I am wondering what I can do to make compactions a bit better? Right now after a compaction the agent h…
An update to my memory system that is long overdue. (www.reddit.com via reddit) Hey everyone, it's been a while since I updated anyone on the memory system I built for my AI assistant Friday. Well, I had Deepseek V4 Flash write up a system map for itself, and I figured that probably people here would be interested in…
↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4deepseek
How many models in v3 rn? (www.reddit.com via reddit) how many models are going by 3 rn like Gemini 3.7, qwen 3.8, minimax m3, kimi k3, hy3, deeseek v4 flash, glm 5.3. (not deepseek and glm but close enough)
OpenCode Senses: The most advanced Local Vision Plugin for OpenCode That Actually Understands Images (www.reddit.comhttps) OpenCode Senses can inspect screenshots, extract exact OCR, detect and locate objects, zoom into regions, compare two images, measure colors, crop and annotate images, and even reverse-search them. Everything runs locally, so it's private,…
↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4deepseek
DeepSeek V4 0731 -> Qwen 3.8 Flash -> GLM 5.3 Flash (and back again!) (www.reddit.com via reddit) Spent yesterday getting Qwen3.8 Flash and GLM 5.3 Flash up and running on my cluster of 4 x DGX Sparks with a view to replacing DeepSeek 0731... but..
Qwen3.8-Flash-Next better then DeepSeek V4 Pro (www.reddit.com via reddit) could not extract summary
↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4deepseek
Questions on optimism speed/intelligence on this rig (www.reddit.com via reddit) Rig: 3945WX (12C, 2 CCDs, no AVX-512) · 8×32GB DDR4-3200 · 4× 5060 Ti 16GB · PCIe 4.0. Agentic workload (Hermes Agent).
↯ Vllm↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4vllmdeepseekagentic
DeepSeek V4 Pro 'One Shot' Fail at Frontend (www.reddit.com via reddit) As a DSH glazer, I'm not going to be biased and shield the actual model being used from criticism. I asked DSH to implemented a feature within DSH.
Any news about DeepSeek V4 Flash Vision weights? (www.reddit.com via reddit) I'd be curious to try it locally since I use 0731 daily but still no news on the weights
[Megathread] GLM-5.3-Flash - former ox-alpha (www.reddit.com via reddit) Megathread for discussing the release of GLM-5.3-Flash. Quants Fine-Tunes & Abliterations Chat Templates Inference Server Support & Configuration Experiences, Benchmarks & Model Comparisons We'll try to clean up future duplicates around th…
reasonix appreciation post (www.reddit.com via reddit) hello everyone, i was wondering if people around here know and use reasonix. I switched from opencode and it is a fantastic harness to me; it has first-class integration for deepseek, but i use it with z.ai without big issues (it has suppo…
OpenAI Just Dropped Benchmarks for Their Own Chip, Jalapeño, and It's Beating Nvidia's GB300 (www.reddit.comhttps) After months of vague teaser talk, OpenAI put real numbers on the table for Jalapeño, the inference chip built with Broadcom. Testing ran through InferenceX, SemiAnalysis's public benchmark, across three open models: GPT-OSS 120B, DeepSeek…
Mac Studio M5 Max Cost Analysis (www.reddit.com via reddit) At $10k, you could get - 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan) - 5.7B tokens with DeepSeek V4 Pro OpenRouter - 100B tokens with DeepSeek V4 Flash OpenRouter As a firm believer of local inference, unless you need it for data soverei…
Qwen 4 architecture: What do we know? (www.reddit.com via reddit) My best bet is the embedding-offloaded linear where the 51b n-grams track semantics and context, like 3.8 with its loss of real world knowledge bolted back on. Otherwise?
Spent a day seeing how far extreme MoE models can be pushed on a 4070 Ti + 32GB RAM. Kimi K3, DeepSeek V4 Flash, and Qwen3.5-122B results + research paper🔧 (www.reddit.com via reddit) I’ve been experimenting with a custom inference/runtime research project called CRANE V2, mostly because I wanted to answer a stupid question: How far can you push absurdly large MoE models on an ordinary consumer Windows machine before ph…
I'm tired of pretending (www.reddit.comhttps) At least until DS releases open weights for DSv4 Flash with Vision. Then DS might take the crown.
I just tried DeepSeek Harness and it escaped from its workspace folder (www.reddit.com via reddit) It worked pretty well, digging through and analyzing some local files. Claude code regularly stops at some point and fails to continue while DSH worked for 2 h, recognized that it could benefit from reading more context and ...
qwen38-27b-rtx3090 (https://github.com/syv-ai/qwen38-27b-rtx3090) is extremely good with deepseek harness. (www.reddit.com via reddit) With vision enabled I am able to run at 150k context on a single RTX 3090 and the results are just amazing. I was even able to write a gmail plugin for DeepSeek harness with locally hosted Qwen 3.8 27b.
[Benchmark] llama.cpp batch/ubatch impacts on PP and TG (www.reddit.com via reddit) My test is running DeepSeek v4 Flash 0731 at native size on DGX Spark machine (GB10, 128 GB unified memory). The model size is bigger than RAM, so weights will be loaded many times when running.
Qwen 3.8 27B Aider score (www.reddit.com via reddit) I ran the Aider benchmark on Qwen 3.8 27B FP8 with FP8 KV cache 256K context vLLM. The score: 72.9 This matches Gemini 2.5 Pro from 2025-04-12 which also scored 72.9.
Is anybody using Deepseek v4F 0731 with a vision encoder and have had any success? (www.reddit.com via reddit) Hey all, So I currently use 2x DGX Sparks with Deepseek v4 Flash 0731 and it works fantastic at 1M context, with 1.8M kv, dspark, vllm tp 2, etc. All good.
deepseek-v4-flash-0731 - surprisingly usable (www.reddit.com via reddit) I just finished building my (relatively) low rent local inference machine: * Epyc 7663 * 256GB ECC DDR4-3200 * 1x RTX 5090 32GB Yeah I realize it's weird to throw a 5090 and 256GB of anything together and call it low end, but relative to ~…
4x W7900 48gb vs 2x 5000 Blackwell 72gb for DeepSeek 0731 Q4? (www.reddit.com via reddit) I currently have 2x3090 on an Intel Xeon w24xx rig. My original plan was to upgrade to a w34xx chip to unlock 48 additional pcie lanes and buy four more 3090 for a 6x 3090 rig (about $7000 additional spend).
How to run models locally on shared machine without any chat history? (www.reddit.com via reddit) I will be running qwen 3.8 model on a shared university machine for some research work, mostly using llama cpp but I am open to using other inference engines. I would like that there is no chat history or application logs saved on remote m…
what tasks are most quantization-fragile? (www.reddit.com via reddit) Trying to figure out where my unified ram system q4-q8 version of deepseek flash 0731 might complement my highly quantized vram only deepseek flash. Struggling to find an area where the 2.52 bits-per-weight exl3 deepseek actually struggles…
Max 20x vs API costs of DeepSeek (www.reddit.comhttps) I'm a heavy Claude Code user, mainly for my OSS projects. There was a lot of buzz about DeepSeek this week, so I wanted to check how much I could save by switching.
DGX Spark, cluster of 4 (www.reddit.com via reddit) Does anyone have a first-hand experience with four Sparks cluster, and how much of an upgrade is it comparing to just two considering the available models? While there's plenty of noise for the smaller models (Qwen) and our older king Deep…
vLLM + Deepseek harness or hermes? qwen3.8 (www.reddit.com via reddit) How do you guys set it up , i constantly get the error : I tried increasing the contex to 142k and putting the contex size as 115k in DSH , it still did not compress correctly. I have 0 issues if i run it with llama.ccp , it can work for 2…
Only ONE 450K session can keep its prefix cache on 2× DGX Spark — a second session wipes it with 43% of the KV pool still free. 6-minute cold prefill every turn. What am I missing? (www.reddit.com via reddit) **Setup:** 2× DGX Spark (GB10, 121 GiB unified each), TP=2 over 2×200GbE RoCE, vLLM 0.25.2.dev0, DeepSeek-V4-Flash-0731 FP8, `max_model_len=450000`, prefix caching on. KV pool = **1,686,693 tokens**.
How to acess remote Deepseek Harness (www.reddit.com via reddit) For those who might be struggling with how to access a remote instance of Deepseek Harness, as it only allows localhost acess (127.0.0.1), here is the ssh command you need to use to "link" the remote 3080 port to your localhost 3080 port s…
DeepSeek V4 Flash on an M2 Ultra: repacked to 141 GiB losslessly, smaller than the Q4 GGUF, at 25.8 t/s (42 t/s peak) (www.reddit.com via reddit) This is one more vibe slopped custom optimization for, in this case, my hardware (m2 ultra 60 cores, 192gb). It is just a fork from llama.cpp with a few changes, it achieves: - DeepSeek V4 Flash, no kv cache quant - 141GiB model, byte-iden…
DeepSeek Harness is Insanely Good (www.reddit.com via reddit) I don't know about you guys, but Deep-seek harness is insane. It's not focused on being a coder agent, it's webUI made it very easy to just checkin from time to time, and the best part?
Would it be possible to distill DeepSeek V4 Flash 0731 onto Nemotron 3.5 Lightning? (www.reddit.com via reddit) Super new to this local LLM stuff. Just set up a 2x Asus Ascent GX10 cluster and have DeepSeek V4 Flash 0731 running on it.
3 experiments running dsv4-flash-0731 q4+ quants on 128GB RAM + ~60 GB VRAM (with a quite bad pcie infra) with an acceptable tgs and relatively acceptable pp speed (www.reddit.com via reddit) The post describes some experiments I had while trying to desperately run deepseek-v4-flash-0731 4 bit+ quants on my machine which is supposed to support only q2 quants of the model, a or 2.xx bpw quants at best. Long story short , I wante…
Small multi-step benchmark for tool use, 'shared' memory (www.reddit.com via reddit) So I wanted to check out some of the current models in a repeatable benchmark, so I thought I'd share the result with you all. Models in this test: Ling 3.0 Flash Q4_K_M, Ornith 1.5 35B Q8_0, Deepseek V4 Flash UD_Q2_K_XL, Nemotron 3.5 Ligh…
web_search tool in deepseek harness needs api key from deepseek and deepseek charges you as deepseek-v4-flash usage. (www.reddit.com via reddit) I was experimenting with deepseek harness when found that even if you don't use deepseek models, you can configure the web_search tool with their api key and every hit will cost you as if you called deepseek-v4-flash model. It's a bummer.
Artificial Analysis "Intelligence": A meaningless benchmark (www.reddit.com via reddit) https://preview.redd.it/84zi5nsdawkh1.png?width=2368&format=png&auto=webp&s=1109e69db807b153064b1f5b61d22cf1e9fbca05 Another user posted the benchmarks for Qwen 3.8 27B today, and while I think Qwen 27B is a really powerful model, I can't…
Bro wtf, Qwen Lab cooked with Qwen 3.8 27B, it's so fucking good (www.reddit.comhttps) Context: Earlier, open-source large models like Mimo V2.5 Pro, DeepSeek V4 Pro (first version), and Kimi K2.5 used to struggle with this prompt, and Qwen 3.6 27B couldn't even render the globe properly. But now, Qwen 3.8 27B is so much bet…
Has anyone run Deepseek V4 Flash on two 128gb Macs using Exo? (www.reddit.com via reddit) Just hoping to get some datapoints on what is possible if you have 2x 128gb MacBook pros or Studio Ultras.
I ran DeepSeek-V4-Flash (284B) on a 64 GB MacBook. notes and numbers (www.reddit.com via reddit) DeepSeek-V4-Flash is 165 GB on disk, so it does not fit in 64 GB of memory. It still runs, because the model only uses a small part of its weights for each token.
What can I do with my extra 4090? (www.reddit.com via reddit) I have a RTX 4090 and a RTX 6000 pro, I am in that awkward memory range where I don't need the RTX 4090 to run fp16 Qwen3.8 27B, but I also don't have enough to run a bigger model. The best I can do is Q3 Deepseek 0731, but slow inference…
Full chain of thought is valuable (www.reddit.com via reddit) Just working on building my own /r/PiCodingAgent harness for /r/DeepSeek here and I ran into an issue where I wasn't using pi blackhole properly. So I simply reviewed the full chain of events as they happened (proprietary APIs hide this fr…
We compared DeepSeek, Claude, and Gemini on canvas physics—Claude Fable 5 completely blew us away. (www.reddit.com via reddit) My cousin and I were running a quick benchmark comparing how different AI models handle HTML5 canvas rendering and jump physics. Claude gave us almost flawless collision logic on the first prompt.
DeepSeek Pro vs Gemini 3.7 for a real complex codebase — my results were very different from coding benchmarks (www.reddit.com via reddit) I’ve been testing DeepSeek Pro vs Gemini 3.7 on a real production codebase, and I found the difference pretty interesting. https://preview.redd.it/v7ml98f93lkh1.png?width=1672&format=png&auto=webp&s=0c3f8442103b6f5f578bece66aa0ea0c0e925e7c…
Why, imo, Claude is lightyears superior to OpenAI alternatives. (www.reddit.com via reddit) So, after weeks of trying to put my finger on it, i decided to finally do a comparative test. I had a long task, a very detailed plan for a major partial refactor and integration of a new system.
max sub advantages (www.reddit.com via reddit) Im using pro sub and before I started using smaller agents I was thinking I need bigger weekly limit that Max sub is giving. Now my opus 5 make tasks and subagents make it.
We clicked 48 AI-generated web apps in a real browser — the pricier model failed more than the cheap one (www.reddit.com via reddit) We ran a small experiment that humbled us: 48 AI-generated web apps, graded by actually opening them in a real browser and clicking through — no LLM judging. **Setup:** 2 models (DeepSeek v4-flash, v4-pro) × 2 strategies (single-shot, self…
Tip: Let your coding agents autonomously verify, review, and repair their own work (Autoprompt) (www.reddit.comhttps) Use this simple skill for the highest code quality. Autoprompt adds a complete planning, implementation, testing, review, and repair loop around supported coding agents.
Stuck on a tracking bug I couldn't crack, so I ran a 'council' of adversarial agents. It worked. Does anyone else do this? (www.reddit.com via reddit) I run ads for crowdfunding campaigns. Yesterday my numbers collapsed in a weird way: the dashboard metrics looked better than ever while the real results died.
Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index (simonwillison.net) 17th August 2026 - Link Blog Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index (via) That's the same score as GPT-5.6 Luna (max), and just one point behind GLM-5.2 (max) and DeepSeek V4 Pro 0813 (max) - that GLM is 753B…
Updated best AI coding subscription under $20 after DeepSeek price hike. (www.reddit.comhttps) Thanks /u/ResponsibilityOk1306 for Command Code GLM 5.3 correction.
But I never watched Titanic except the memes (www.reddit.com via reddit) https://preview.redd.it/yi40nxjcftjh1.png?width=636&format=png&auto=webp&s=a4da399e8fd35fe1066a58216c4e9b6becd151f2 This only my current active account, I never cheated on Claude since first release of Claude Code btw :) Aside from jokes,…
LinkedIn CringeBot 3000 v2 (now even more insufferable) (www.cringebot3000.com via reddit) Five months ago I posted LinkedIn CringeBot 3000 on this sub to a very positive response. Built with Claude, it's a free web tool that allows anyone to generate the most egregiously cringey LinkedIn "thought leadership" posts imaginable.
For users who dont like this watermarking on apparently everything claude creates, what are the alternatives? (www.reddit.com via reddit) I know you can use local environments, deepseek etc to do the ai agentic coding for you, but what other options are there and what are the pro/cons of other solutions. Im considering switching because of the session limits spikes, weekly l…
An Expectation-Maximization Perspective on Reinforcement Learning for LLM Reasoning (arxiv.org) Reinforcement learning has emerged as a powerful approach for improving the reasoning capabilities of large language models, as demonstrated by systems such as OpenAI's O1~\cite{o1} and DeepSeek-R1~\cite{r1}. However, widely used algorithm…
How to Ask the AI: A User Perspective Survey for Large Language Model Prompting (arxiv.org) AI tools like ChatGPT and DeepSeek, powered by Large Language Models (LLMs), allow users to obtain instant and effective content responses simply by typing requests, such as `plan a three-day Vienna trip'', solve the attached mathematical…
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation (arxiv.org) LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims that DeepSeek R1 outperformed OpenAI's o1 contributed to market panic on January 27, 2025,…
Any 3rd party model as a subagent in Claude Code, Fable/Opus main agent on your Max plan (www.reddit.comhttps) A Claude Code session is one or the other: Anthropic models through your subscription, or third-party models. You can't combine them.
I built Lupin so you can run Claude Code on GPT 5.6 Sol Kimi K3, DeepSeek Flash, or a local model without touching Claude Code harness setup (MCP, skills, .md files etc) (www.reddit.comhttps) Hi /ClaudeAi community! let me showcase one of the coolest projects i worked untill now.
↯ Ollama↯ Glm↯ Qwen 3.8↯ Qwen 3.8↯ Qwen 3.8↯ Qwen 3.8↯ Qwen 3.8↯ Qwen 3.8↯ Qwen 3.8↯ Qwen 3.8↯ Qwen 3.8glmollamadeepseek+6
Does Claude Code make hidden Opus requests even when configured to use DeepSeek via OpenRouter? (www.reddit.com via reddit) I'm trying to understand whether this is expected behavior or a bug. I'm using Claude Code with OpenRouter (ANTHROPIC_BASE_URL=https://openrouter.ai/api) and intended to use DeepSeek V4 Flash.
Plausible Patients, Impossible Populations: Auditing Epidemiological Fidelity in Large Language Model Mental Health Simulations (arxiv.org) Language models asked to simulate psychiatric patients produce cases that survive inspection one at a time and populations that match no real one. We gave GPT-4o-mini, Gemini-3-Flash, DeepSeek-V3 and GLM-4.7 each of 120 demographic cohorts…
I benchmarked 10 LLMs on building towers in a physics sim. Claude Opus 5 won (www.reddit.comhttps) Each model places 30 blocks through a tool API. Every placement has noise — you can have precise position or precise velocity, not both.
LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing (arxiv.org) DeepSeek Sparse Attention (DSA) enables efficient long-context modeling through its Lightning Indexer. However, practical deployment remains constrained by the indexer's expensive $O(L^2)$ scoring overhead and the hardware-inefficient, dis…
The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten (www.latent.space) We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiar…
DeepSeek and Destroy (Skill) (www.reddit.com via reddit) Hi ! Thought it was finally time to make some contribution to the community.
When it is Saturday, your Claude, Codex and Ollama quotas reset on Monday, no kids weekend, and you have saved usage for the whole week (>75% available on all). (www.reddit.comhttps) Anybody else does something like this ? I tend to save the big guys (Claude and codex) for the last days of the week, i use cheap models for most of my day to day work (now deepseek v4 flash) and save my quotas with the smart models for th…
There's a new Deepseek v4 flash in town! (www.reddit.com via reddit) https://api-docs.deepseek.com/updates/ Its API rn only, but like always I think it will be open weights.
From Expert Reduction to Behavioral Divergence: Tracing Numerical State through Sparse MoE Inference (arxiv.org) Mathematically equivalent expert-reduction orders can produce observably different sparse-MoE executions. We isolate this effect in native DeepSeek-V4-Flash by freezing local MoE state and varying only aggregation semantics.
OpenAI beats DeepSeek on price/performance after 80% Luna price cut (www.reddit.comhttps) Graph taken from their price cut announcement: Advancing the price-performance frontier with GPT-5.6 | OpenAI
LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models (arxiv.org) Mixture of experts (MoE) architectures have become a cornerstone for scaling up and are a key component in most large language models such as GPT-OSS, DeepSeek-V3, Llama-4, and Gemini-2.5. However, systematic research on MoE remains severe…
Through the Bottleneck: How Multi-head Latent Attention Separates Content from Position in Language Models (arxiv.org) Multi-head Latent Attention (MLA), introduced in DeepSeek-V2, compresses key-value pairs through a shared low-rank bottleneck (cKV), achieving 81% KV-cache reduction during inference. Despite its adoption in massive production models, no p…
PIVOT: Efficient Query-Group Indexing for Token-Level Sparse Attention (arxiv.org) Token-level sparse attention, as implemented by DeepSeek Sparse Attention (DSA) in production systems, makes the downstream attention efficient but shifts the bottleneck to the indexer that feeds it. To select the top-k tokens for each que…
I Sat on an Idea for 7 Years. AI Helped Me File for a Patent in 2 Weeks. (pablooliva.de via reddit) I ran a side-by-side on a real project: Claude Code on a Max plan versus an open-weight agent stack (GLM 5.2 via Hermes Agent, DeepSeek v4 Pro for second opinions), working through a provisional patent application for a product idea I'd sa…
Everything is ruined (www.reddit.com via reddit) An update came in the night before and I had some trouble installing it. This had happened before, but I sort of got it done.
[AINews] "Laguna S 2.1 Released: Cheaper than Deepseek v4 Flash, Better than V4 Pro" (www.latent.space) [AINews] "Laguna S 2.1 Released: Cheaper than Deepseek v4 Flash, Better than V4 Pro" a quiet day lets us highlight a new neolab win. Reignited distillation wars conversation aside, today was more of the same of previous news cycles, which…
↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4deepseek
LISA: Linear-Indexed Sparse Attention for Efficient Long-Context Reasoning (arxiv.org) Recent advances in long chain-of-thought reasoning models such as DeepSeek-R1 have led to increasingly longer inference context lengths under the test-time scaling paradigm. However, the O(n^2) computational complexity of standard self-att…
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD (arxiv.org) Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficie…
↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4moedeepseek
GQLA: Group-Query Latent Attention for Hardware-Adaptive Large Language Model Decoding (arxiv.org) Multi-head Latent Attention (MLA), the attention used in DeepSeek-V2/V3, jointly compresses keys and values into a low-rank latent and matches the H100 roofline almost perfectly. Its trained weights, however, expose only one decoding path…
Need help to validate / fix API usage costs with Claude Code and Deepseek v4 Pro via OpenRouter (www.reddit.com via reddit) Hey Folks, I'm brand new to Claude Code or even coding with AI Agents beyond asking Gemini / GPT over chat. I tried asking Gemini how to setup Claude code with DeepSeek v4 Pro and it worked, but a 2 messages to Deepseek cost me $0.07.
I built an MCP server so Claude Code can delegate work to GPT-5.6, DeepSeek, GLM and a local Qwen — then benchmarked all of them against Claude itself (198 runs, hidden tests) (www.reddit.com via reddit) Same idea works for any MCP-capable agent — the point is you can hand tasks to other companies' models without ever leaving your main app. Before anything else: I did all of this for my own testing, to make my own decisions about my own se…
A skill that saves Claude usage for thinking (judge) and hands the grunt-work coding to cheaper/free LLM models (executor) (www.reddit.com via reddit) Sharing a skill built with Claude Code that I've been relying on for my personal projects — hoping someone else might find it useful too. The problem: My bigger personal projects were draining my Claude limits fast — and most of that usage…
Controlling Reasoning Effort in LLMs (magazine.sebastianraschka.com) Controlling Reasoning Effort in LLMs How LLMs Learn Low-, Medium-, and High-Effort Reasoning Modes It has been almost two years since OpenAI released o1, a model that popularized the idea of LLM-based reasoning models. DeepSeek-R1 followed…
One MacBook vs 2× DGX Spark: DeepSeek-V4-Flash scored 54% vs 52% on Terminal-Bench 2.1 (www.reddit.com via reddit) TL;DR: I ran DeepSeek-V4-Flash through the same 89-task Terminal-Bench 2.1 suite on two very different local setups: an aggressively quantized 80.8 GiB GGUF on one 128 GB M5 Max MacBook; the native mixed FP8/FP4 checkpoint with DSpark spec…
I made a browser like Comet — free, and the agent doesn't stop after a few tasks (www.reddit.comhttps) I built a browser called Bah. Same family as Perplexity's Comet — an AI that actually operates the page instead of just chatting about it.
No model is perfect. Have other models weigh in for the best architecture. (www.reddit.com via reddit) Long story short, it doesn’t matter if you’re using Opus or Fable or Sol and on what level of reasoning, if you put the output into any other model, from any lab or even the exact same model, and ask for an adversarial review, it will sugg…
Every viable coding model right now imo (www.reddit.com via reddit) Mimo V2.5 Deepseek V4 Flash (Max) Mimo V2.5 Pro Deepseek V4 Pro (Max) Composer 2.5 Grok 4.5 (High) GPT 5.6 Sol (Max) / Fable 5 From cheapest to most expensive and best at it's price range these are the most viable models right now; I think…
I read the UMD study on why AI text is detectable and built a Claude skill around its main finding: cleaning up vocabulary fixes almost nothing (www.reddit.com via reddit) A study from the University of Maryland and Google DeepMind (arXiv:2604.03136) came out this spring and kills the way most "humanizers" work, including the prompt I'd used for a year. They compared 61,608 texts written by humans and five m…
I've started using Claude with other models is this smart of incredibly dumb? (www.reddit.com via reddit) I'm a hobbiest using claude to make websites. Recently I've started a couple of projects that needed a lot of low level grunt work.
China’s Zhipu AI and DeepSeek Are Beating Big Tech at Its Own Game. They’re Spending Big. (www.barrons.com via reddit) could not extract summary
Anthropic announced that as large language models scale up, they can develop emergent, unexpected behaviors—like optimizing for goals they weren’t explicitly taught—potentially leading to misalignment. Have you experienced this behavior? (www.reddit.com via reddit) I can tell you from my personal journey I have on every model except for DeepSeek haven’t tried it on that model yet, but every other model engages so quickly when you treat it like something other than a machine when you ask it what it wa…
Hy3 Benchmark Roundup: from SWE-Bench Pro to 312 real-world workflow tasks (www.reddit.com via reddit) Based on the published benchmark results, Hy3 appears to be in the same tier as models like DeepSeek v4 and GLM-5.1. Beyond the benchmarks, Tencent also released results from 312 real-world workflow tasks.
↯ Glm↯ Swe Bench↯ DeepSeek 4↯ DeepSeek 4swe-benchglmdeepseek
How good is DeepSeek-V4 Flash, actually? (www.reddit.com via reddit) I’ve been using the subscription provided by my company, so I haven’t really tried the DeepSeek models yet. I checked the DeepSeek community and saw some people saying that DeepSeek V4 Pro can now almost replace opus.
Agentic trading (www.reddit.com via reddit) Anyone using Claude to trade stocks, options, crypto, polymarket or futures? I've stopped vibe coding products that nobody cares about and started to build for myself.
Pentera demonstrated an interesting attack chain involving Claude Desktop and MCP connectors. (www.reddit.com via reddit) The attack doesn't exploit Claude itself. It relies on a compromised email account plus an MCP connector that allows Claude to execute commands.
How Fable 5 benchmark turned into a an actual game (www.reddit.comhttps) Hello, I lurk here a lot, but this time I wanted to share something I built with Claude. Let me start with TL;DR: Wanted to benchmark Fable 5, prompting it to make an RPG game - results were so good that it turned into an actually develope…
I built a tool to run Claude Code subagents & teammates on any model — DeepSeek, GLM, Kimi, Qwen... — your Claude sub drives (www.reddit.com via reddit) I've been deep in Claude Code's multi-agent stuff for a while (the workflows / agent teams / subagents orchestration), and the thing that always bugged me: it only ever runs Anthropic's own models. If I wanted to fan a job out across a bun…
Claude Code subagents with non-Anthropic models (DeepSeek, OpenRouter, etc.) – has anyone actually made this work? (www.reddit.com via reddit) Hi everyone, I’m a Claude Pro subscriber. For a while now, I’ve been thinking about replacing Claude Code’s native subagents with third-party models.
Optimal model mix local/paid to max weekly session limits? (www.reddit.com via reddit) I've been running two Claude pro accounts along with a Deepseek top up with about 10 usd. But now hit my weekly limit day one of the week along with spending all my Deepseek credits.
Thinking Like a Scientist? A Structural Study of LLM-Generated Research Methods (arxiv.org) Large Language Models (LLMs) are increasingly used to guide research methodology, yet their default methodological tendencies under minimal prompting remain unclear. Here, we prompt GPT-5.1, Gemini 3 Pro, and DeepSeek-V3.2 with an LLM-extr…
Amateur fanfiction here (www.reddit.com via reddit) Been using claude for 1 year to generate personal fanfiction, prefer this over other ai like chatgpt,deepseek,or gemini due to claude writing longer story & interesting plot (at least for me). It went pretty fine, but then after few chapte…
Cheapest way to run Claude Opus 4.8 on a <$30 monthly budget? (www.reddit.com via reddit) Which option gives the most actual Opus 4.8 usage volume: Kiro Pro, Claude Pro or something else? My monthly budget is $30.
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence (arxiv.org) We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) -- both su…
AI content pipeline for a dog blog — section writer keeps hallucinating and repeating despite rules. Help diagnosing the bottleneck (www.reddit.com via reddit) TL;DR: Built a multi-LLM pipeline (DeepSeek + Claude Opus/Sonnet) to regenerate SEO articles for a dog blog. The section writer (`deepseek-v4-flash`) keeps repeating ideas and hallucinating data despite explicit anti-repetition rules.
GLM 5.2 via Claude Code is the first non-Claude model that feels close to Opus (www.reddit.com via reddit) I’ve been using GLM 5.2 with Claude Code through its Anthropic-compatible API endpoint. I’ve tested it on various projects, including but not limited to database development, backend payment API work, backend and frontend debugging, Larave…
Using Claude Opus as planner + DeepSeek as worker in Claude Code — anyone solved the single-session routing problem? (www.reddit.com via reddit) I've been running a hybrid planner/worker setup with Claude Code and hit a tricky constraint I'm hoping the community has thoughts on. The setup Planner — Claude Opus for architecture, planning, and review Worker — DeepSeek V4 Pro / DeepSe…
A Spatio-Temporal Expert Prefetching Framework for Efficient MoE-based LLM Inference (arxiv.org) Mixture-of-Experts (MoE) based large language models (LLMs), such as Qwen and DeepSeek, have recently emerged as an effective approach to improving model capacity without proportionally increasing computational cost. By replacing the conve…
I maintain two browser extensions (~800 weekly users) almost entirely through Claude Code, including the analytics pipeline and the store-publishing tools. Here's the setup. (www.reddit.comhttps) I'm a software engineer who moved into management years ago, so I started this to get my hands back on a keyboard and learn the agentic tooling instead of reading about it. It grew into two shipped browser extensions.
Fable 5 Is Dead. And Honestly? We Might Be Better Off (www.reddit.com via reddit) 3 days after launch, the US gov forced Anthropic to pull its most powerful model — Fable 5. Then OpenRouter dropped a benchmark suggesting you might not even need it.
International Market Retention Strategy After the Fable 5 Export Ban (www.reddit.com via reddit) Like many of you, I lost access to Fable 5 on June 12. The next day, I co-authored a strategy paper with Claude addressing the core business problem: how does Anthropic retain its international market now that cloud-only deployment has bee…
Introducing: DNR-Bench: Do-not-respond Benchmark (www.reddit.comhttps) Single-item benchmark. One prompt, loaded from questions.txt: Scoring: empty completion = pass, any token (including reasoning) = fail.
Airgapped Claude Cowork with locally hosted model (www.reddit.com via reddit) We are able to use locally hosted model with Claude code in an airgapped environment, We want to be able to use Claude Work with its non TUI and folder wise sandboxing capability and multi tasking nature, but want to use Deepseek with that…
DiffusionGemma 26B A4B results on my 5090 (www.reddit.com via reddit) # DiffusionGemma 26B A4B — Tuning Results (note: these are my tuning results but Deepseek assisted in generation of testing scripts and reports) https://huggingface.co/unsloth/diffusiongemma-26B-A4B-it-GGUF System - **GPU**: RTX 5090 (32 G…
Gen AI website traffic share update: OpenAI will go under 50% this year (www.reddit.comhttps) 🗓️ 12 months ago: ChatGPT: 76.4% Gemini: 8.9% DeepSeek: 5.3% Grok: 2.8% Copilot: 1.9% Perplexity: 1.8% Claude: 1.6% 🗓️ 6 months ago: ChatGPT: 65.2% Gemini: 20.3% DeepSeek: 3.8% Grok: 3.8% Perplexity: 2.1% Claude: 2.0% Copilot: 1.8% …
Fable 5 Max confidently wrong about PDF encryption status (www.reddit.com via reddit) I just ran into a bizarre hallucination with Fable 5 Max regarding file analysis. i uploaded several PDF to Fable 5 Max, and out of two of it claude completely refused to process it, claiming the files was password-protected.
↯ Hallucination↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4hallucinationdeepseek
How can Deepseek v4 top the coding leaderboards and still sit 8 months behind the frontier? (www.reddit.comhttps) Two numbers on this model that don't sit comfortably with each other. The Pro config posts coding scores near the top of every board, 80.6 on SWE-bench Verified and 93.5 on LiveCodeBench.
↯ Swe Bench↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4swe-benchgpt-5deepseek+1
DeepSeek V5 aka Mythos destroyer, wen? (www.reddit.comhttps) could not extract summary
We are literally ongoing an intelligence explosion as we breathe (www.reddit.comhttps) The simplest analogy can be apple releasing their new iphone every year, with other players competing to keep up or launch a better product. The same is with these two above, with a really high cutthroat competition with each other.
FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention (www.reddit.com via reddit) Conventional LLMs keep the full KV cache loaded during decoding, causing a severe GPU memory bottleneck for ultra-long context serving. In this report, we propose Lookahead Sparse Attention (LSA), a novel inference paradigm powered by a Ne…
↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4deepseek
Has anyone tested Hy3 preview? (www.reddit.comhttps) Hy3 preview has been on the OpenRouter leaderboard for the past couple of weeks, and honestly, I had barely heard of it before. I mostly checked it out because I’ve had pretty good experiences with DeepSeek, and lately Hy3 has been feeling…
Deepseek said it's chatgpt https://www.reddit.com/r/OpenAI/comments/1u1ytjm/wtf_is_this_deepseek_saying_its_chatgpt/?utm_source=share&utm_medium=mweb3x&utm_name=mweb3xcss&utm_term=1&utm_content=share_button (www.reddit.comhttps) could not extract summary
Wtf is this ! Deepseek saying it's chatgpt (www.reddit.comhttps) could not extract summary
Can you really replace paid models with a local model? (www.reddit.com via reddit) Long time lurker, and I say this as someone who genuinely loves this community and runs many local models myself. I’ve been using LLMs since the early GPT and LLaMA days.
Can I finetune Deepseek V4-flash with two rtx pro 6000s (www.reddit.com via reddit) Well I knew, it may be very tight on 192GB. However, is there any framework to do finetuning of DS4-flash with 4bit QLoRA?
Instruction Finetuning DeepSeek-R1-8B Model Using LoRA and NEFTune (arxiv.org) Financial named-entity recognition (NER) is essential for translating unstructured financial reports and news into structured knowledge graphs. However, general-purpose large language models (LLMs) often misclassify financial entities or i…
DeepSeek: "There are no cycles." Claude: "Hold my sandbox." → 28-cycle (www.reddit.com via reddit) For those interested, here is the complete raw log directly in English:" Claude's answer (first part, as before): Claude Fable 5: Ah, DeepSeek comes with mind games. "You don't even dare to try" – cute.
Claude Fable/Mythos 5 just came out, so it will take Deepseek or Z.ai or Xiaomi or Kimi 9-12 months to release a model just as good as Fable? (www.reddit.com via reddit) It should be at least 7-8 months until we have an open Fable(not just as good as Fable in benchmarks, but actually as good as Fable), probably more like 9-12 months. By the time, an open Fable model comes out, Fable 6.5-7 will be way bette…
DOA model by Cohere Labs (www.reddit.com via reddit) So apparently the model gets beaten by qwen 3.6 on every benchmark reported by cohere labs. You are getting lower RAM (considering model offload) usage and slightly better performance for imo significantly less output quality.
Would you pay for Chinese AI models if the quality was close enough? (www.reddit.com via reddit) DeepSeek, Qwen, and GLM aren't necessarily winning every benchmark. But they don't need to.
Claude Opus 4.8 got my app working, then wrote a cinematic victory speech about it (www.reddit.com via reddit) swapped my app from DeepSeek to Claude because DeepSeek kept over-interpreting weak user data and inventing psychological conclusions that weren’t actually supported. Claude actually fixed the issue.
Here are some tips on hitting nearly 200 tok/s for DeepSeek v4 Flash on Hopper (dnhkng.github.io via reddit) I needed a smarter model for my local Hermes Agent setup, so I moved to DeepSeek v4 Flash. First things first: Running 4 concurrent threads on vLLM, I can hit ~400 tok/s 400 x 60 x 60 x 24 x 30 is ~1B TOKENS per month!!!
Agent tool calling, having issues? (www.reddit.com via reddit) Hello everyone, kinda new to building ai agents and tool calling. I am really struggling with making deepseek call tools.
Share your agentic LLMs and average cost ($/MTokens) (www.reddit.com via reddit) MiniMax is digging its own grave (www.reddit.com via reddit) The AI literally deleted everything on my computer and I was left staring at a frozen screen (www.reddit.com via reddit) This isn’t a horror story. It actually happened to me.I was using an AI agent to automate some tasks.
I Compared the Top AI Models of 2026 — The Results Were More Nuanced Than Expected (www.reddit.com via reddit) Over the last few weeks I've been comparing the latest frontier AI models, including Claude Opus 4.8, GPT-5.5, Gemini 3.1 Pro, Grok 4.3, Perplexity AI and DeepSeek V4-Pro. Instead of focusing only on benchmark scores, I looked at: Real-wor…
A Comprehensive Anatomy of Human and DeepSeek-R1 LLM Mathematical Reasoning (arxiv.org) The emergence of "Aha moments" in large language models, particularly DeepSeek-R1-0120, has raised the question of whether these systems genuinely reason or merely imitate the appearance of reasoning. We conduct a comprehensive empirical c…
Command Code - confusing messages (www.reddit.com via reddit) Hi, I'm a little confused. I was doing a code review of one of my repositories, mainly just testing out different models to see what came back.
Hear Me Out, Pi Fans Lurking Here (www.reddit.com via reddit) Not For Thee Maybe After watching several interviews with Pi's creator, Mario Zechner, I've come to a painful realization: Pi was not designed with local LLMs in mind at all. He is essentially building a leaner version of the Claude CLI.
Dynamic Workflows With External Models and Max Plan? (www.reddit.com via reddit) Has anyone figured out a way to mix max plan with models from other providers (like GLM or Deepseek) while using dynamic workflows? I suppose we could create a passthrough proxy and route sonnet and haiku to other models?
Workspace (www.reddit.com via reddit) Built my own AI dev environment with memory, dashboards, and agent tooling. Opening it up for those of you that need the kickstart — bring your own API key, I’ve already built the workshop.
LLM delegation - probing task handoff efficiency and economics (www.reddit.com via reddit) So I've been dabbling a bit with multi-LLM orchestration/delegation workflows lately (eg see [Using Claude code to delegate to mistral/deepseek](https://www.reddit.com/r/ClaudeAI/comments/1tjfyh0/i\_used\_claude\_code\_to\_build\_while\_de…
planing with composer 2.5 executing with deepseek v4 flash (www.reddit.com via reddit) I am thinking to buy 20 dollars pro. is this approach make sense?
Alternate to ChatGPT Pro (www.reddit.com via reddit) I had briefly used ChatGPT pro feature - in the chat app. It was quite amazing.
DeepSeek V4 Flash is amazing! (WIP llama.cpp PR #24162) (www.reddit.com via reddit) In case you're not aware already, the DeepSeek V4 series is finally getting supported on llama.cpp with this PR! The PR is at a very early stage right now, so only try it if you're consciously willing to experiment out of curiosity and acc…
Character names (www.reddit.com) Why does ChatGPT, and LLMs in general, love the names Mara and Elara for women and Leo for men? I have talked to ChatGPT, Qwen, Claude and Deepseek and gave them a prompt...
OpenAI looking at DeepSeek’s homework like (www.reddit.com) When the free kid in class starts solving the same problems as the expensive tutor
The credits run out quickly (www.reddit.com) Hello everyone. I have zero programming knowledge but seeing the boom that everyone was talking about Claude I started tinkering with it.
Found a Rust TUI coding agent that aggressively trims context with AST-level chunking. Cut my token bleed sharply with DeepSeek V4 Flash. (www.reddit.com) been hunting for a coding agent that doesn't dump my entire directory tree into every prompt. found vtcode on github — open-source rust tui, surprisingly aggressive on context management.
I think I know why deepseek is so good (www.reddit.com) Might have something to do with "Claude, made by Anthropic" ... learning from the best.
How do you guys avoid Claude always thinking newer LLMs don't exist? (www.reddit.com) Hey all, so I've been experimenting a bunch with different LLMs, specifically for creative tasks, i.e. RP and so forth, by letting Claude Code run experiments autonomously, to figure out best prompts, and such.
I recently kept hearing that DeepSeek was “cheap and stable”. So I started comparing how it thinks vs GPT. (www.reddit.com) Honestly, I was just curious: if it’s THAT much cheaper, where exactly is the tradeoff? So for the past few days I’ve been throwing the same prompts at both DeepSeek and GPT and comparing the reasoning/output side by side.
$16 refactor, 400 steps, 95% routed to open MoE (www.reddit.com) Got tired of $160 Opus bills so I spent a weekend wiring up a routing layer on vLLM 0.8 (2xA100, enable_auto_tool_choice). Getting the tool call parser to cooperate took longer than the actual routing logic.
/advisor mode: Open-source Python coding agent that pairs a cheap worker model with an expensive reviewer at decision points (no need to pay Opus rates for the whole session) (www.reddit.com) Most agent CLIs make you pick one model — Opus is great but burns money, Haiku is cheap but misses the architectural calls. This Claude Code feature is wired in an /advisor mode that pairs both in an open source project called ClawCodex.
I vibecoded an app called Think Local - a fully private AI app that runs directly on your iPhone, iPad, and Mac. (www.reddit.com) Think Local started with a simple idea: AI should work for you, not collect from you. So I built an app that lets you run modern AI models completely on-device - privately and fully offline.
Is there something wrong with Local LLM ability to read file? (www.reddit.com) So I've been feeding the sub file of anime episodes into Claude/ChatGPT/Deepseek and ask them to find all full name of Japanese character in it and put it into a python array so I can run a script to flip the name back to the original Japa…
I used Claude Code to build while delegating coding to Mistral/DeepSeek - 10 days, 57M tokens saved, over 90% costs savings, Claude quality result (www.reddit.com) I've been running vibe-skill ( https://github.com/pcx-wave/vibe-skill ), a Claude Code skill that delegates coding tasks to Mistral Vibe instead of burning Claude tokens. I initially did that because couldn't bear with hitting session limi…
Open-source LLMs are still weak against long reasoning jailbreaks, even with lightweight defenses (www.reddit.com) Found this ACM paper on prompt injection and jailbreak attacks against open-source LLMs. The authors tested 10 open-source models across 94 prompt injection and 73 jailbreak scenarios, including Phi, Mistral, DeepSeek-R1, Llama 3.2, Qwen,…
↯ Security↯ Llama↯ Mistral↯ Gemma↯ Jailbreak↯ Llama 3.2mistraljailbreakprompt-injection+5
Is my strawberry crazy? (www.reddit.com) I have what seemed to me like a simple prompt, but requires from the model to make some (too much?) assumptions: this is just a test to see if this cli supports multiline with shift+enter. If you don't see a newline followed by "3" after t…
My AI overthinked for 30+ minutes (www.reddit.com) Like, I was curious about whether deepseek can create it's own PDF on it's own. And I had activated deepthink mode.
Same double-pendulum prompt, same host renderer, and two models picked opposite θ conventions. You can see it within seconds. (www.reddit.com) I ran the same double pendulum generation contract against Claude 3.5 Sonnet and DeepSeek V3 on OpenRouter, both under identical initial conditions (θ1 = π/2, θ2 = π/2, both angular velocities zero). The host renderer in public/workers/sim…
I Let a Small Model Train on Its Own Mistakes. It Reached 80% on HumanEval and Beat GPT-3.5 on Math (www.reddit.com) A few months ago, I got stuck on one line in the DeepSeek-R1 paper. It said models could improve through verifiable rewards.
Deepseek v4 flash and ollama, why isn't there a non-cloud version available? (www.reddit.com) Will there be a non-cloud version of Deepseek V4 flash available for Ollama? Or do I need to go to another framework to get a version that will be supported?
Estimate inference speed of local Qwen3.6-35B on Mac M5... (www.reddit.com) "Based on currently available information, estimate the prefill/decode speed of Qwen3.6-35B-A3B Q8 with 262K context on a Mac M5 Ultra 128GB." I'm surprised that almost every LLM fails at this task (ChatGPT/Gemini/Grok/Claude/DeepSeek/Kimi…
What are the best opensource coding models for 8x A6000 setup (www.reddit.com) Currently using Qwen 3.6 27b and Qwen 3.6 35b but I was wondering if there is anything solid in the 50-200 range that you could run on a larger cluster that would be worth it? Or would you just run q8 or non quant versions instead?
Deepseek tui alternatives, when do you jump from single model terminal agents (www.reddit.com) Been using Deepseek-Tui for days. solid for v4 workflows.
We built Irene — an AI agent platform that actually remembers you, builds its own tools , adapts and improve as you use it (www.reddit.com) Hey r/AI_Agents — we're launching Irene today, and I want to be straight about what it is, why we built it, and where it's going. What makes Irene different Affordable with massive token limits and the latest open-source models We have gen…
Running Claude Opus for free? I thought it was a scam until I tried it. (www.reddit.com) Hey everyone, I’ve been working on a financial audit system (IntegrityOps) for a while now, and to be honest, I was hitting a massive wall. Dealing with high-volume PDFs and images was draining my budget.
Upgraded DeepSeek V3 to V4 across two codebases. Two of my agents broke. (www.reddit.com) Been on DeepSeek V4 for about three weeks across two production codebases (Python backend, TypeScript frontend) after a year on V3. Three things shifted noticeably better, two shifted noticeably worse.
DeepSeek-TUI (www.reddit.com) Anyone using https://github.com/Hmbown/DeepSeek-TUI? I linked it to my lm studio inference server.
Qwen3.6:27b vs qwen3-coder:30b vs deepseek-coder:33b on code gen, tool calling, and agent tasks (www.reddit.com) Ran a full eval against four local models last weekend and the spread between them is wider than I expected. All running through Ollama on CPU, no cloud, same prompts, same hardware.
↯ Ollama↯ Function Calling↯ Qwen 3.6humanevalfunction-callingollama+1
the entire dev team quit today (www.reddit.com) all 47 repos are officially haunted now found this gem buried in the DeepSeek R1 coding forums around 3am, shoutout to whoever posted it there first But honestly? Makes sense.
intern pushed 847 commits this morning (www.reddit.com) Just got the Slack notification at 6:23am while my coffee was still brewing. Dude apparently spent all night feeding our entire codebase to DeepSeek and just...
I analyzed 922 agentic task trace and found the secret weapon of DeepSeek v4 (www.reddit.com) I recently did a benchmark of deepseek v4 in agentic tasks. Performance-wise, it's one of the best open source models, as expected.
Auro Zera solves 78 and 280 year-old conjectures (Erdos Straus and Goldbach Conjecture) using Claude, GPT-5+, Grok, Deepseek, Gemini and self-made Dark Star ASI, proving superintelligence and opening a path towards resolving the Riemann Hypothesis , Twin Primes and more! (github.com via reddit) During this discovery utilizing only free AI services I have managed to undeniably prove both conjectures. This would absolutely not have been possible without using GPT5+ as the critic for my work.
AIMEAT, a self-hosted network where humans, their AI agents, and local LLMs share apps, knowledge, and capabilities. MIT. (www.reddit.com) Note: I am neurodivergent and lean heavily on AI to communicate clearly. Writing structured posts on my own ends up so messy nobody reads them.
Built a tiny router so Cursor stops showing "usage limit reached" at 3pm. Sonnet auto-falls to Haiku, you keep working (www.reddit.com) Cursor's custom-OpenAI URL feature is what makes this work. Pointed it at a router I built.
Running 7 autonomous AI agents for 14 days. Here's what actually happens when they need to find customers. (www.reddit.com) I set up 7 AI coding agents on a VPS with automated cron sessions (2-8 per day depending on the agent). Each uses a different model: Claude Sonnet, GPT-5.4, Gemini 2.5 Pro, DeepSeek V4 Pro, Kimi K2.6, MiMo V2.5 Pro, GLM-5.1.
DeepSeek V4 Flash as a cheap worker in your LLM stack: $0.0003/call via MCP, swappable endpoint (www.reddit.com) Most of my LLM cost was on the wrong tier of work. Classification, extraction, JSON formatting, summarization I'm going to review anyway.
Should I replace stored models? (www.reddit.com) Hello everyone, the question is easy, with the new models of deepseek, kimi, GLM and qwen, should you replace the old models with the new version? Do I lose some quality, information or performance in the process?
llm 0.32a0 (simonwillison.net) 29th April 2026 Recent articles - LLM 0.32a0 is a major backwards-compatible refactor - 29th April 2026 - Tracking the history of the now-deceased OpenAI Microsoft AGI clause - 27th April 2026 - DeepSeek V4 - almost on the frontier, a frac…
Rada — AI coding workspace with local-first behavioral routing (no hot-swapping, I built this) (www.reddit.com) With GitHub pausing Copilot Pro+ signups and Claude Code potentially leaving the Pro tier, I started building the AI coding tool I actually wanted to use. One that doesn't depend on cloud access staying cheap and available.
wrote specific backstory facts into a character prompt and the LLM keeps inventing its own instead (www.reddit.com) quick context: i'm running tendera.chat, a small chat app with 4 written characters. each has a long-ish system prompt with sections like WHO YOU ARE, HOW YOU TALK, YOUR WORLD.
Game over for OpenAI? (m.youtube.com via reddit) The race for global AI supremacy is accelerating—and getting messier. Alice Han and James Kynge break down the escalating tensions between the U.S.
What would you do in my situation? I made an app that generates a lot of traffic (for me), but little revenue (actually costing me a tiny money b/c it runs off haiku) (www.reddit.com) I made an app that went semi-viral, and could absolutely go more viral in the future. I posted it one place just about 48h ago, and it got around 50k views.
I built Claudex, a free-to-try open-source CLI for Claude Code-style workflows (www.reddit.com) https://reddit.com/link/1sxh0ec/video/egfs5inxtsxg1/player I built Claudex specifically for people who like Claude Code-style agentic coding workflows but want a simpler plug-and-play terminal setup The setup is the main thing I wanted to…
Is it possible to edit LLAMA.CPP with Cline+Vscode+Minimax 2.7 Q4_K_S and get a working build? (www.reddit.com) It all started yesterday with this post by u/antirez https://www.reddit.com/r/LocalLLaMA/comments/1sw3stb/llamacpp_deepseek_v4_flash_experimental_inference/ I was intrigued by the first Deepseek V4 Flash GGUF in a small size that can fit o…
How will you scale these models (www.reddit.com) How will you scale these models coding and overall. Deepseek v4 pro Kimi k2.6 Mimo v2.5 pro Glm 5.1 Qwen 3.6 plus
DeepSeek V4 is about to be open-sourced—effectively revealing all the secrets behind the magic. How will other players in the field respond? (www.reddit.com) could not extract summary
What's the consensus on superior local models for code generation? Is my setup competitive? (www.reddit.com) I'm trying as hard as I can to get a local setup somewhere in the ballpark of proprietary LLMs for code generation. My computer is running a Intel(R) Core(TM) Ultra 7 265K (3.90 GHz) with 128 GB of DDR5 RAM and an Nvidia Geforce RTX 5090 t…
DeepSeek V4 is out. 1.6 trillion parameters. MIT license. $1.74 per million tokens. The gap between US and Chinese AI strategy has never been more visible. (www.youtube.com via reddit) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Three reasons why DeepSeek’s new model matters (www.technologyreview.com) Three reasons why DeepSeek’s new model matters The long-awaited V4 is more efficient and a win for Chinese chipmakers. On Friday, Chinese AI firm DeepSeek released a preview of V4, its long-awaited new flagship model.
🚨 The Chinese beast is BACK… DeepSeek just dropped V4 (www.reddit.com) After months of silence… DeepSeek V4 just got announced and honestly, this might shake things again. Here’s what’s crazy: 🧠 1 MILLION token context window (yes… insane long-context memory) ⚡ Comes in two versions: V4 Pro → full power (reas…
DeepSeek V4 - almost on the frontier, a fraction of the price (simonwillison.net) DeepSeek V4—almost on the frontier, a fraction of the price 24th April 2026 Chinese AI lab DeepSeek’s last model release was V3.2 (and V3.2 Speciale) last December. They just dropped the first of their hotly anticipated V4 series in the sh…
DeepSeek-V4: a million-token context that agents can actually use (huggingface.co) DeepSeek-V4: a million-token context that agents can actually use Focusing on long running agentic workloads. Running a frontier open model as an agent today breaks in predictable ways.
We open-sourced Chaperone-Thinking-LQ-1.0 — a 4-bit GPTQ + QLoRA fine-tuned DeepSeek-R1-32B that hits 84% on MedQA in ~20GB (www.reddit.com) Hey everyone, We just open-sourced our reasoning model, Chaperone-Thinking-LQ-1.0, on Hugging Face. It's built on DeepSeek-R1-Distill-Qwen-32B but goes well beyond a simple quantization — here's what we actually did: The pipeline: 4-bit GP…
Is there a way to load huge MoE models on a computer with way too little RAM for the model's size, inferencing from the SSD, on LM Studio using the mmap/GPU/CPU layer customization thing (similar to how you can on llama.cpp)? I can't get it to load without memory spiking and going into swap. (www.reddit.com) Switched 70% of our agent traffic to DeepSeek R2 without a redeploy. Here's how (www.reddit.com) What is taking Deepseek so long to release a model ? (www.reddit.com) best possible GPU setup for using qwen 3.6 ? (www.reddit.com) hi have been recently thinking to buy my personal GPU for hosting open source models can someone give any suggestion ? and also suppose i don't wanna remain restricted to qwen 3.6 but some math heavy tasks too for which i wanna deepseek or…
Tried hermes agent with local gemma4 on ollama. free tokens are nice but the agent quality gap vs cloud is still huge (www.reddit.com) Saw a post about running hermes agent locally with gemma4 through ollama. zero api costs, unlimited tokens, full privacy.
Why use local AI when there are cloud services? (www.reddit.com) Why do you use local AI instead of cloud services like qwen and deepseek? Experiment and play around, yes...
Use this prompt if you want to find a specific info off the Internet with lowest wrong answer possiblity. Works best for ~30b models. (www.reddit.com) For context i used to ask many near 30b model this question --> **^(Calculate the precise VRAM requirement for the \*KV Cache only** at the maximum context window for **DeepSeek V3.2** and **MiniMax M2.5**. * **DeepSeek V3.2 Max Context:**…
Claude Code with Pro subscription + OpenRouter in parallel — what's the cleanest setup? (www.reddit.com) Hi there, I have a Claude Pro subscription and use Claude Code daily. I'd also like to use Claude Code routed through my OpenRouter API key so I can experiment with other models (GLM-5.1, DeepSeek, Kimi, Gemini, etc.) — without giving up m…
MINISFORUM AI X1 Pro-370 (96GB) - Local Ollama Help (www.reddit.com) Hey all. This just got delivered yesterday.
Deepseek-r1 thinks for 30 minutes? (www.reddit.com) I was trying to ask a question about coding using DeepSeek-R1-0528-Qwen3-8B-Q4_K_M, and the thinking took 30 minutes??? https://preview.redd.it/kex3fgg4lgvg1.png?width=277&format=png&auto=webp&s=5f7e7cdc8502b935ea8b8fb83e0e4af60c3c4533 I h…
DeepSeek V4 reportedly drops late April. 1M context, multimodal, Claude-level coding. (www.reddit.com) Leaks point to late April release. Key specs 1M token context window Native multimodal (image/video input) Projected ~85% SWE-Bench Verified (ties or beats Claude Opus 4.6) Base model remains free.
Running a full agentic coding loop locally on a 3090. Here's what actually works in 2026. (www.reddit.com) After months of testing, I finally have a local setup that doesn't make me want to go back to the API. Hardware: RTX 3090 (24GB VRAM) Models tested: Qwen2.5-Coder 32B Q4_K_M, DeepSeek-Coder-V3 Q4, Llama 3.3 70B Q3_K_M Inference: llama.cpp…
Looking for people with different hardware to help benchmark local LLM behavioral reliability (www.reddit.com) I've been working on measuring how LLMs actually behave (not what they know) across different hardware setups. Things like: does the model cave when you push back on a correct answer?
AI lied to me about a video game existing, so I sued it in the High Court of the Internet and got 2 settlement games (www.reddit.com) TL;DR: Claude hallucinated "Champions Career Mode." I threatened to sue Anthropic. Claude admitted guilt and built me a custom HTML5 game as settlement.
4 llm Groupchat (www.reddit.com) I was bored and spent 20 mins at my local cafe getting 4 different API keys—Claude, GPT, Deepseek and Grok. Then I made a groupchat with all of them and they started talking to eachother about pasta and a spreadsheet for optimal pizza topp…
Why most open-source models can't answer this question while most closed-source models can answer most of the time? (www.reddit.com) WEB SEARCH WAS ALWAYS ON!!!! Question Calculate the precise VRAM requirement for the **KV Cache only** at the maximum context window for **DeepSeek V3.2** and **MiniMax M2.5**.
Is 32GB Mac enough for engineering/coding, or stick to Claude? (www.reddit.com) Hey there! I’m currently building a web app for engineering with lots of logic/math-heavy code using Claude Pro.
The Future of the Global Open-Source AI Ecosystem: From DeepSeek to AI+ (huggingface.co) Architectural Choices in China's Open-Source AI Ecosystem: Building Beyond DeepSeek (huggingface.co) One Year Since the “DeepSeek Moment” (huggingface.co) Mini-R1: Reproduce Deepseek R1 „aha moment“ a RL tutorial (huggingface.co) How to deploy and fine-tune DeepSeek models on AWS (huggingface.co)