could not extract summary
#minimax
246 items
Ryan Lee from MiniMax posts article on the license stating it's mostly for API providers that did a poor job serving M2.1/M2.5 and may update the license for regular users! (www.reddit.com) I'm glad we have deepseek (www.reddit.com) other companies are slowly going away from open weight, not releasing base models, delaying open weight distribution, not releasing top models (this one I think is fair, but still), and I also noticed they stopped publishing research (old…
Open Models - April 2026 - One of the best months of all time for Local LLMs? (www.reddit.com) Any underrated or overlooked models? FYI MiniMax-M2.7 switched their license(from MIT to Non-Commercial) so it's not in graph.
MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video (blog.comfy.org via hn) MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video An open-weights omni-modal video model with real stereo sound and 2K output — this powerful model is greatly optimized in ComfyUI and can run locally on a 3060.…
Minimax M2.5 vs. GLM-5 vs. Kimi k2.5: How do they compare to Codex and Claude for coding? (www.reddit.com) Guys we have to change the pelican test (www.reddit.com) So i have been seeing more of those pelican on a bike svg tests and while they work i feel like (and maybe you guys do too) they are getting kinda benchmaxxed so we should switch things up soon and this is my idea generate me a html svg of…
MiniMax m2.7 under 64gb for Macs - 91% MMLU (www.reddit.com) https://huggingface.co/JANGQ-AI/MiniMax-M2.7-JANGTQ Used TQ as quantization method where it matters. Finally mac users under 64 gb - esp base m5 users can get a real cloud SOTA-like level LLM running from home.
Update LICENSE · MiniMaxAI/MiniMax-M2.7 at edf8030 (huggingface.co via reddit) RyanLee's(MiniMax) recent tweets for same. I just updated our license.
Antirez/h3.c: MiniMax H3 inference engine for Mac computers (github.com via hn) h3-metal Native MiniMax-H3 inference for Apple Silicon. The project is being built as a sequence of working vertical slices: deterministic host/model metadata first, then portable Metal block parity, prompt encoding, prompt-to-video/audio,…
Dual dgx spark (Asus GX10) MiniMax M2.7 results (www.reddit.com) MiniMax released MMX-CLI: one CLI for text, image, video, speech, music, vision, and web search — no MCP server needed. Works natively in Claude Code, Cursor, OpenClaw. (www.reddit.com) MiniMax just open-sourced MMX-CLI, a command-line tool built specifically for AI agents. Seven command groups: mmx text, mmx image, mmx video, mmx speech, mmx music, mmx vision, mmx search.
MiniMax M2.7 GGUF Investigation, Fixes, Benchmarks (www.reddit.com) Hey r/LocalLLaMA, we did an investigation into MiniMax-M2.7 GGUF causing NaNs on perplexity. Our findings show the issue affects 21%-38% of all GGUFs on Hugging Face (not just ours).
My first impressions of Minimax M2.7 (Q5_K_M) vs Qwen 3.5 27b (Q8_0) (www.reddit.com) I'm not sure if the AesSedai's Q5_K_M version of Minimax M2.7 is too much lobotomized or if the model itself is kind of weak. I did a simple experiment with both models running with the recommended parameters.
Those of you running minimax 2.7 locally, how are you feeling about it? (www.reddit.com) Im running the raw version straight from the minimax release on hugging face (https://huggingface.co/MiniMaxAI/MiniMax-M2.7) on 3 rtx pro 6000's on vllm. So no quantization.
Running Minimax 2.7 at 100k context on strix halo (www.reddit.com) Just wanted to share because it took me a lot of tweaking to get here: llama-server -hf unsloth/MiniMax-M2.7-GGUF:UD-IQ3_XXS --temp 1.0 --top-k 40 --top-p 0.95 --host 0.0.0.0 --port 8080 -c 100000 -fa on -ngl 999 --no-context-shift -fit of…
2x Asus Ascent GX10 - MiniMax M2.7 AWQ - cloud providers are dead to me (www.reddit.com) Hello, I've been on a quest to get something "close enough" of Opus 4.5 running locally, for agentic coding, as SWE with 15 years of experience. I tried with one spark (yeah I'm calling my Asus Ascent GX10 sparks - they're the same), with…
A 1-bit quant of MiniMax 2.7 that runs from a CD at 1500 tk/s would be nice. (www.reddit.com) Badda Boom.
Single question llm comparison (www.reddit.com) MiniMax M2.7 AWQ-4bit on 2x Spark vs 2x RTX 6000 96GB - performance and energy efficiency (www.reddit.com) Hello, This model/quant is my daily driver and I wanted to have some reference benchs for comparing my setup with a 3x more expensive and 4x time power hungry setup. Results first, methodology after, link at the end with all results Model:…
Kimi K2.6-Code-Preview, Opus 4.7, GLM 5.1, Minimax M2.7 and more tested in coding (www.reddit.com) Hi everyone. It's been a while since I posted (was a lil burned out), but some of you may have seen my older SanityHarness posts.
Bench 8xMI50 MiniMax M2.7 AWQ @ 64 tok/s peak (vllm-gfx906-mobydick) (www.reddit.com) Inference engine used (vllm fork): https://github.com/ai-infos/vllm-gfx906-mobydick/tree/main Huggingface Quants used: cyankiwi/MiniMax-M2.7-AWQ-4bit Relevant commands to run: docker run -it --name vllm-gfx906-mobydick-mixa3607 -v ~/llm/mo…
I Made LLMs Play Texas Hold’em. The Smallest Model Beat a ~1T Model by Being Too Dumb to Fold (www.reddit.com) Made LLMs play Texas Hold’em against each other. 6 models at the table: a tiny 1.2B running locally on my 16GB MacBook, a couple mid-size ones, and cloud models going up to about 1 trillion parameters.
MiniMax M2.7 ultra uncensored heretic is Out Now with 4/100 Refusals, Available in Safetensors and GGUFs Formats! (www.reddit.com) llmfan46/MiniMax-M2.7-BF16-ultra-uncensored-heretic: https://huggingface.co/llmfan46/MiniMax-M2.7-BF16-ultra-uncensored-heretic llmfan46/MiniMax-M2.7-ultra-uncensored-heretic-GGUF: https://huggingface.co/llmfan46/MiniMax-M2.7-ultra-uncenso…
Updated Minimax m2.7 still doesn't allow coding a product. But before the next riot starts, Ryan Lee has already confirmed that they are still working on the license, and sale of products built by m2.7 is permitted. (www.reddit.com) could not extract summary
Show HN: A free CLI coding agent, powered by ads (freebuff.com via hn) We subsidize Deepseek 4.0, MiniMax M3, and more!
Pushing the limit: minimax m2.7 q8_0 128k on 2x3090, 256GB DDR4 (www.reddit.com) CPU is just a secondhand 10900x. Using 128k context, unquantized kv cache.
Your local LLM predictions and hopes for May 2026 (www.reddit.com) Which of these do you think we'll get in May? Also, feel free to pick/rank which ones you'd want the most badly: more Gemma4 models (124b?) (other sizes?) more Qwen3.6 models (9b?
Comparing GPT-5.4, Opus 4.6, GLM-5.1, Kimi K2.5, MiMo V2 Pro and MiniMax M2.7 (www.codejam.info via hn) Free 100M AI tokens for Kimi and MiniMax models (inference.dahl.global via hn) cheap and free from corporate oversight Low prices Open models Zero data retention First 100 million tokens free DEMO CHAT — try it Affordable AI inference Get access to powerful open models through a decentralized GPU network built to red…
DeepSeek's 10T USD grand strategy (twitter.com via hn) Have you ever wondered, how DeepSeek may make money, and lot of it? They didn't come up with competitive coding plans like GLM, MoonShot and MiniMax.
Testing MiMo-V2.5-IQ3_S with 1'048'576 context (www.reddit.com) llama-server.exe --model "H:\gptmodel\AesSedai\MiMo-V2.5-GGUF\MiMo-V2.5-IQ3_S-00001-of-00004.gguf" --ctx-size 1048576 --threads 16 --host 127.0.0.1 --no-mmap --jinja --fit on --flash-attn on -sm layer --n-cpu-moe 0 --threads 16 --parallel…
I hate this group but not literally (www.reddit.com) True story, I got interested in AI after seeing it at work and wanted to run models locally. I started with an M3 Ultra 96GB, quickly learned it was not enough for what I wanted, and kept upgrading hardware (including refurbished Mac Studi…
Cuda + ROCm simultaneously with -DGGML_BACKEND_DL=ON ! (www.reddit.com) I invested quite a bit of time and it wasn't easy but finally I can run models like Minimax 2.7 Q4 using Cuda+ROCm at the same time bypassing Vulkan. load_tensors: offloaded 63/63 layers to GPU load_tensors: CUDA0 model buffer size = 83650…
Tenstorrent TT-QuietBox 2 Specifications (Blackhole) (www.reddit.com) Source: https://docs.tenstorrent.com/systems/quietbox/quietbox-bh-2/specifications.html Currently supported models: https://tenstorrent.com/developers From the specification docs above: CPU: Ryzen 7 9700X 65W Granite Ridge 3.8GHz Memory: 2…
Nvidia Super Acceleration for MiniMax H3 Video (nvlabs.github.io via hn) MiniMax H3 Super Acceleration fast draft generation and high-resolution refinement, powered by Sol Engine 6.85 s for a 5-second 768p video · 14.93 s for a 10-second video H3 Super Acceleration first uses H3 with a LoRA to generate a four-s…
MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities (twitter.com via hn) MiniMax (official) @MiniMax_AI Introducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities - Coding & Agentic Frontier: 59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 34.8% SWE-fficiency, 28.8% KernelBench H…
Minimax M3 on Open Router (openrouter.ai via hn) MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding, and tool use.
I built a local GUI for the TradingAgents framework — works with Ollama (www.reddit.com) https://preview.redd.it/i90oxxk7n03h1.png?width=1898&format=png&auto=webp&s=7d219c804fda7dfe122b84fcdb6d0d6883818c68 A while back I came across TradingAgents — a really cool multi-agent LLM stock analysis framework where like a dozen "agen…
Best AI coding plan alternative to Claude and ChatGPT (news.ycombinator.com) With the lowering usage limit in Claude, I am thinking of jumping ship to Chinese AI, since the benchmark is already very near compared to Sonnet or Haiku 4.5 , but for a fraction of the price. I am not worried about where is my data endin…
Considering two Sparks for local coding (www.reddit.com) I'm currently running a 4x RTX 3090 system (96GB VRAM, DDR4 2133 RAM) and have tested opencode and pi.dev using Qwen3.5-122B-A10B (AWQ) up to 200k context for web app coding (html/js/python). I'm now seriously considering picking up two Sp…
Agile as a cat (www.reddit.com) https://preview.redd.it/kgkv6knv2dyg1.png?width=1026&format=png&auto=webp&s=d2e37f1914136ad672bcecf98741eee5e8cd69da MiniMax M2.7 AWQ 4bit hallucinated a URL and instantly pivoted to treating its own error as a joke. That made me laugh (do…
Current state of open-source ? (www.reddit.com) I’m trying to understand the current open-source LLM landscape beyond surface-level hype. We all got used to the nerfed products of Claude/Geminj so I believe really in opensource as a solution.
VC-Attention: Faster Low-Bit Attention Without Retraining (www.nunchux.ai via hn) VC-Attention: Faster Low-Bit Attention Without Retraining Attention speedup over BF16 FlashAttention-4 [1] on B200 and B300. We benchmark the attention workload in MiniMax-H3 when generating 243 frames at 1344×768.
Show HN: Fake Zoom, hang out with AI coworkers and feel the synergy (fake-zoom.pages.dev via hn) This is so stupid. I made a fake Zoom call with AI coworkers.
Show HN: MiniMax H3 on a 16GB Mac, 5 days after open weights (github.com via hn) VPIPE Lightweight local multimodal AI pipelines and custom Metal inference for Apple Silicon Macs. Multimodal graph: video, audio, images, text, and tool actions in one pipeline.
MiniMax H3: Open weights Omni-modal model (news.ycombinator.com) could not extract summary
What Is MiniMax H3? Everything You Need to Know About the Hailuo 3.0 Video Model (minimaxh3.art via hn) On July 31, 2026, Chinese AI company MiniMax officially launched MiniMax H3 — the third-generation model in its Hailuo video family, also known as Hailuo 3.0. First previewed at WAIC 2026 just two weeks earlier, H3 arrives with a clear amb…
China's MiniMax releases H3 video model (www.reuters.com via hn) could not extract summary
MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities (twitter.com via hn) MiniMax H3: Omni-Reference, Commercial-Grade Generation, Unbeatable Cost Efficiency, Open Weights - As I said, MiniMax-H3 is Open! #1 Video Editing (With Audio) #2 Text to Video (With Audio) #2 Image to Video (No Audio) artificialanalysis.…
Optimizing MiniMax M3 Sparse Attention on Nvidia Blackwell (fireworks.ai via hn) Fireworks built a KV-stationary sparse-attention kernel for MiniMax M3 on NVIDIA Blackwell (SM100), reaching ~980 TFLOP/s: 1.9–2.4× a query-stationary baseline and ~1.6× open-source MSA. The post walks through the Q-outer vs KV-outer desig…
Relay – open-source coding agent for non-mainstream/Chinese LLM providers (github.com via hn) Relay An open-source, dark-mode desktop coding agent — built for people who want to use non-mainstream LLM providers, not just the big three. Relay is an Electron app that puts DeepSeek, Qwen, GLM, Kimi, MiniMax, and other open/Chinese mod…
MiniMax-M3: A native multimodal model with 1M context (huggingface.co via hn) MiniMax-M3 is a native multimodal model with 1M context. It has ~428B parameters and ~23B activated parameters.
MiniMax debuts AI model built for long and complex coding tasks (www.scmp.com via hn) MiniMax debuts AI model built for long and complex coding tasks Shanghai-based company says M3 can process data five times faster than its predecessor, while also slashing inference costs Chinese artificial intelligence start-up MiniMax ha…
MiniMax teased M3 Sparse Attention: 9.7x prefilling, 15.6x decoding at 1M (twitter.com via hn) Don’t miss what’s happening People on X are the first to know. Log in Sign up Post Conversation Skyler Miao @SkylerMiao7 Something BIG is coming 2:49 PM · May 26, 2026 307.3K Views New to X?
JANGQ-AI/MiniMax-M2.7-JANGTQ_K : mixed-bit quant of MiniMax M2.7 - 74 GB on disk (huggingface.co via reddit) MiniMax-M2.7-JANGTQ_K MiniMax M2.7 — 74 GB on disk (down from ~230 GB FP8 source) — mixed-bit JANGTQ_K quantization in JANGTQ-PRESTACK layout. Source: MiniMaxAI/MiniMax-M2.7 (62 layers, 256 routed experts top-8, 196K context) Quantization:…
Strix Halo 128GB on Proxmox - Vulkan vs ROCm benchmark matrix (www.reddit.com) Ryzen AI MAX+ 395, Bosgame M5, 128GB LPDDR5x. Proxmox VE 9.1 LXC containers with GPU passthrough.
I got better results when I made each AI tool do one job (www.reddit.com) I spent too much time trying to find one AI dev tool that could do everything. Planning, coding, fixing, reviewing, maybe filing my taxes too It never really worked.
Accelerating MiniMax-H3 768p Video Generation on a Nvidia DGX Spark in 1 Minute (nvlabs.github.io via hn) Sol-H3- Spark Accelerating MiniMax-H3 768p Video Generation on a Single NVIDIA DGX Spark in 1 Minute NVIDIA Research, Efficient AI Team & Singapore Lab. A two-stage pipeline specialized for a single NVIDIA DGX Spark generates a 384p draft…
Max H3 – MiniMax H3 Max AI Video Generator (www.maxh3.com via hn) A performance with spatial sound Human movement · prop contact · working-room ambience Max H3 creator workspace MiniMax H3 Max delivers faster, more stable AI video generation with stronger prompt adherence. Turn text, key frames, or a ref…
SlopTV: A livestream of AI slop generated from comments, Minimax H3 on 2x5090 (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Endless sitcom using Minimax H3 and a turbo LoRA (www.twitch.tv via hn) Channel One | AI Interdimensional Cable (Original)
Minimax H3 Max Live – An infinite broadcast is generated by the chat (www.twitch.tv via hn) Chat Directs the Show with Minimax H3 Max | Live AI Broadcast Experiment | Streaming just chatting for 113 viewers.
MiniMax H3 Max sets the new Pareto Frontier for video generation (twitter.com via hn) BREAKING: MiniMax H3 Max sets the new Pareto Frontier for video generation, nearly 50x faster than the base model. This model is post-trained by @fal on @MiniMax_AI H3, and it's in a league of its own: no other Image to Video model on the…
MiniMax-Music3 (huggingface.co via hn) MiniMax Music 3 MiniMax Music 3 is a high-performance music generation model for creating complete songs up to five minutes long. Conditioned on lyrics and a detailed music description, it generates structurally coherent songs with express…
MiniMax M3: How Sparse Attention Makes Long-Horizon Agents Practical (twitter.com via hn) https://t.co/v9huIornsf elvis@omarsar0ArticleMiniMax M3: How Sparse Attention Makes Long-Horizon Agents Practical GLM 5.2 has taken over much of the AI timeline lately, and most of the conversation has centered on how it stacks up against…
Flash-MSA: Accelerating Million-Token Training with Sparse Attention Kernels (nanduruganesh.github.io via hn) [Github] [MiniMax Paper] [Trainer] Several frontier models [1, 2, 3, 4, 5] use sparse attention to greatly speedup their inference, though no one has posted code to train it efficiently. Today I introduce the world's first performant open-…
Zhipu AI, MiniMax shares for Hong Kong investors as lock-ups end (www.scmp.com via hn) Zhipu AI, MiniMax shares to provide gut check for Hong Kong investors as lock-ups end Shares worth US$11.5 billion to hit market as record wave of lock-up expirations starts and some firms eye share placements Hong Kong’s stock market coul…
Show HN: Cline subscription plan to access GLM-5.2 at 2-5x discount (cline.bot via hn) Hi I'm Saoud, founder of Cline. We’ve been impressed with GLM-5.2 and so are introducing a $9.99/month subscription to give you 2-5x discounted access to it and other open weight models like DeepSeek, Kimi, MiniMax, Mimo, and Qwen.
GLM 5.2 ranks #2 in Code Arena: Frontend (twitter.com via hn) Exciting news: GLM-5.2 (Max) ranks #2 in Code Arena: Frontend, with +29pt over Claude Opus 4.7 (Thinking) and only behind Fable 5! GLM-5.2 is the best open model vs Kimi-K2.6 and Minimax-M3 by a large margin.
Boundaries of Stationary Feature Learning: A Minimax Barrier for Scaling Laws (zenodo.org via hn) I do not derive the Chinchilla scaling law; I map the boundaries of the regime in which such a derivation could even be attempted. Working in the μP feature-learning setting on a Sobolev-on-manifold data model, I establish what the station…
MiniMax M3 Benchmarks One Pager (filecdn.minimax.chat via hn) could not extract summary
Show HN: Free open source coding models in Slack (www.runcord.com via hn) Hey HN, We believe we have the easiest onboarding from signup to being able to spin up coding agents in slack like Stripe, Ramp & Coinbase. Demo of the onboarding: https://www.tella.tv/video/connecting-cord-to-slack-1-19ep Every signup get…
↯ Glm↯ Minimax↯ Gemma↯ DeepSeek 4↯ DeepSeek 4minimaxglmgemma+5
What workstation to get for ~13k EUR? (www.reddit.com) My use-cases will be to test open-weight LLMs and work on harnesses, inference systems and possibly other non-ML workflows (CS-related) in the future. Fine-tuning would not be something I do locally because I can rent a B200 from RunPod fo…
↯ Llama↯ Vllm↯ Minimax↯ Fine Tuning↯ DeepSeek 4minimaxvllmfine-tuning+2
Chinese AI Coding Plan (www.reddit.com) With the lowering usage limit in Claude, I am thinking of jumping ship to Chinese AI, since the benchmark is already very near compared to Sonnet or Haiku 4.5 , but for a fraction of the price. I am not worried about where is my data endin…
Is Qwen3-coder the best kept secret out there? (www.reddit.com) So I'm brand new to this scene but I'm using Claude to help me fine tune a model for a startup idea I have in the Healthcare space. I have been working with the 27-35B parameter mdoels (Qwen3.6, Gemma 4) and the couple of 120B+ models (Qwe…
Mesh LLM to build private personal AI, using open models (www.anarchai.org via hn) Model Catalog Filter Connected Peers | ID | Role | Version | Status | Model | Latency | VRAM | Share | | --- | --- | --- | --- | --- | --- | --- | --- | | | Host | 0.65.1 | Serving | MiniMax-M2.5 | <1 ms | 256.0 GB | 42% | | | Host | 0.65.…
M3 Ultra + DGX Spark = M5 Ultra-lite? (www.reddit.com) So I saw an article recently about exo disaggregated prefill with DGX Spark and M3 Ultra - prefill on one machine and decode on another. DGX Spark apparently has 4x matmul performance over an M3 Ultra - same as the M5 Ultra should have.
What's the best suscription under 20$? (www.reddit.com) I’m pretty overwhelmed. I feel like there are so many options that I don’t know which one to choose, and trying things until I find a decent one isn’t really my thing—even though I enjoy it.
Best Practices to Start with Vibe Coding? Best Local Apps for Agentic Vibe Coding? (www.reddit.com) DISCLAIMER: I am not a programmer nor do I have experience coding. I've been thinking about a small app running on gradio for some time now, and I want to try tweaking some extension for ComfyUI.
eGPU vs system RAM (www.reddit.com) OpenCode + Self host Minimax-2.7 via SGLang? (www.reddit.com) anyone knows how to setup opencode to work with self hosted minimax-2.7 properly? It has <think> and </think> in the message and OpenCode failed to parse the answer correctly.
Use Claude, ChatGPT, or MiniMax Subscriptions in Cursor (open-vsx.org via hn) Ungate A Cursor-first extension for using Claude, ChatGPT, and MiniMax subscriptions in Cursor instead of paying for API tokens. How it works Ungate lets you use Claude, ChatGPT, and MiniMax in Cursor through account subscriptions instead…
Minimax M2.7 on Q3_K_S or Smaller Model with greater precision? (www.reddit.com) I currently am looking for models to fit into my single DGX Spark for use. I have an RTX Pro 6000 and also a 5090 as well that I'm considering using in combination if the DGX Spark is too slow, but the intent here is to play around with Op…
Ollama Cloud - Pro (www.reddit.com) Hi. I've been looking at ollama cloud's Pro offering ($20), which says "Run 3 cloud models at a time".
Ask HN: Former grok-code-fast-1 users, what coding model are you using now? (news.ycombinator.com) I get good, cheap, fast feature coding success with grok-4.1-fast for planning and grok-code-fast-1 for execution. But according to the Openrouter usage stats, grok-code-fast-1 is now old hat - usage dropped off a cliff in mid-Feb.
HyperFlow – one LoRA for all MiniMax-H3 tasks (ref2va/t2va/fl2va) (github.com via hn) HyperFlow for MiniMax-H3 HyperFlow is Video Rebirth's 8-step LoRA for MiniMax-H3, obtained by data-free flow self-distillation and running on the official diffusers Modular Pipeline. Diffusers' default 50-point sigma schedule performs 49 m…
MiniMax H3 Max Prompts (github.com via hn) Awesome MiniMax H3 Max Prompts English · 简体中文 Learn MiniMax H3 Max through real public examples: watch a shot, read the breakdown, then copy a prompt to make your own version. Explore action, performance, animation, short stories, referenc…
MiniMax H3: measured cost per clip on seven rented GPUs at four providers (qrun.cloud via hn) Measured GPU runs What rented GPUs actually cost per unit of work. Measured runs on rented GPUs: what each provider actually charged, and the cost per finished unit of work — a video clip, an image, a thousand training steps, a million tex…
Measuring Malicious Intermediary Attacks on the LLM Supply Chain (twitter.com via hn) I bought a Fable dataset from one of the top Chinese LLM routers yesterday. With just 6TB data, I can take over 7 Chinese/CIS gov entities & 19 top Chinese firms like Xiaomi, Huawei, NIO, Minimax using SSH keys, VPN configs, Aliyun key…
Chess5.ai – Play Chess, Go, Xiangqi, Gomoku, and Othello Against LLMs (chess5.ai via hn) chess5.ai Human vs LLM · Five Games Pit yourself against GPT, Claude, Gemini, Grok, Muse Spark, Mistral, DeepSeek, Kimi, Qwen, GLM, or MiniMax across five classic boards. How it works - Human vs model, or model vs model — with spectating a…
Show HN: MusicMaxxer – a songwriter's MiniMax Music 3 UI (github.com via hn) Tired of paying for Suno, I made this local desktop app for Windows and macOS to generate songs through MiniMax's hosted Music API (free for up to 3 songs/minute!). I wanted something tailor-made for a songwriter, so I implemented some QoL…
Show HN: Highlander – realtime MiniMax FastH3 with audio at $0.02 per second (www.highlander.sh via hn) Video generated faster than it plays. We serve MiniMax Fast H3 — the FastH3 VSA checkpoint distilled from MiniMax H3 — on eight H100s.
China's MiniMax sees revenue nearly quadruple in first half as AI demand surges (www.reuters.com via hn) could not extract summary
Show HN: MiniMax H3 – Turn text and images into AI video clips (minimax3.com via hn) MiniMax now lists H3 as available, with confirmed 4–15 second output, 768P and 2K options, and multimodal references. Generate on this site, or read our API guide for the workflow, supported formats, and site pricing.
MiniMax M3 Medium hits 73.17% F1 on DeepSearchQA, near GPT-5 High (huggingface.co via hn) MiniMax M3 DeepSearchQA Skill Eval Evaluates minimax/minimax-m3 on google/deepsearchqa using a Pi agent, You.com MCP tools, and a research skill optimized for this harness, model, and tool surface. MiniMax M3 Medium Reasoning with the You.…
AMA: MiniMax H3 Team – Ask anything about open video generation model (old.reddit.com via hn) could not extract summary
Run MiniMax-H3 Locally with SGLang Diffusion on 2× RTX 5090s or 1× RTX Pro 6000 (twitter.com via hn) @MiniMax_AI H3 is live in SGLang Diffusion, with day-0 serving support 🎬 This open model matches Seedance 2.0 at 1/3 the cost, or $0 if you run it locally on 2x 5090 or 1 RTX 6000. With SGLang Diffusion, you can build visual concepts, m…
MiniMax-H3 weights are up (huggingface.co via hn) MiniMax H3 System Overview MiniMax H3 is a general-purpose, omni-modal generative system. It supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio…
Minimax H3 on Anime (www.reddit.com via hn) could not extract summary
If U.S. labs slow down AGI development, this could be 2028 (news.ycombinator.com) If U.S. labs slow down AGI development, this could be 2028: Moonshot Kimi 6, Alibaba Qwen 5, Z.ai, and MiniMax are all claiming AGI-level capabilities.
Show HN: Same castle prompt, 8 LLMs, 24 procedural Three.js worlds (castle-bakeoff.pages.dev via hn) Fable 5 · GPT 5.6 Sol · Kimi K3 · Grok 4.5 · Gemini 3.5 Flash · MiMo V2.5 Pro · MiniMax M3 · GLM 5.2 — low-poly, semi-realistic, very realistic.
Harry Partridge on X: "GLM 5.2 With Vision" / X (twitter.com via hn) https://t.co/iJsDrlGy45 Harry Partridge@part_harry_ArticleGLM 5.2 With VisionGLM 5.2 is one of the best currently available open source language models. However, unlike other flagship models like Qwen, Kimi and Minimax, GLM 5.2 does not su…
MiniMax M3 vs. GLM 5.2: Codegen comparison across autonomous coding tasks (thinkwright.ai via hn) Thinkbench, our custom evaluation harness, was used to drive both models through the same autonomous coding loop: read files, write files, run shell commands, and stop when the task was complete. The scored suite covered greenfield builds,…
Testing MiniMax M3 on refactoring, screenshot debugging, music recommendations (andlukyane.com via hn) A hands-on look at MiniMax M3 through Claude Code — what its new MiniMax Sparse Attention (MSA) is and how it differs from the lightning-attention and full-attention designs of earlier MiniMax models, plus three real tasks: auditing and re…
Inference Optimization for MiniMax Sparse Attention (www.together.ai via hn) - Together AI is the preferred cloud partner for MiniMax M3. Together AI will host the open-weights model as a developer endpoint upon its public release.
MiniMax M3 Review: Matching GPT-5.5 and Opus? (thomas-wiegold.com via hn) I ran my usual coding tests — two websites, a poker sim, and a code audit. Here's how MiniMax M3 actually stacks up against GPT-5.5 and Opus 4.8.
MiniMax M3 on Qubrid AI (news.ycombinator.com) Coding & Agentic Frontier. 1M-context MSA.
been pairing M2.7 with Hermes Agent for a few weeks. holds up surprisingly well. anyone else running this combo? (www.reddit.com) been self-hosting hermes agent locally for a few months and rotating through different model backends for it. tried claude sonnet 4.5, gpt-5.5, qwen 3.6 coder, and most recently minimax m2.7.
↯ Minimax↯ Sonnet 4.5↯ Sonnet 4.5↯ Sonnet 4.5↯ Sonnet 4.5↯ Sonnet 4.5↯ Sonnet 4.5minimaxgpt-5qwen+1
Finally tested an AI video tool that works directly in Claude without setup (www.reddit.com) Been using Claude for everything creative lately and got tired of switching to Runway every time I needed video. Found out Higgsfield supports MCP, connected it once, and now Claude generates video directly in chat.
Help me choose an LLM Provider which doesn't take my life savings (www.reddit.com) Hi everyone 👋 I’m trying to choose an LLM provider for my personal projects and side experiments, but I also don’t want my API bill to quietly consume my entire salary 😅 My primary use cases are: Coding assistance Agentic workflows Browser…
Testing MiniMax M2.7 via API on three real ML and coding workflows (andlukyane.com via hn) Testing MiniMax M2.7 via API on three real ML and coding workflows I recently got access to some MiniMax M2.7 API credits, so I decided to plug this model directly into Claude Code and run it on three workflows I do regularly. The same tas…
Full Hermes Agent tutorial (Spanish with English auto-translation). Computer Use, MCP Blender, Hindsight memory and multi-agent setup (www.reddit.com) Spent weeks running Hermes Agent in production on my Mac Mini M4 before recording this. Wanted to show things nobody else was covering.
Has anybody been able to achieve reliable agentic performance with cheap/open source models? (www.reddit.com) Basically the title. Recently I've been trying various open source and comparatively cheaper models like minimax m2.7, qwen models and glm5.1 in Pi agent from openrouter, and the performance on coding tasks have be moderately adequate at b…
Regex Chess: A 2-ply minimax chess engine in 84,688 regular expressions (nicholas.carlini.com via hn) by Nicholas Carlini 2025-01-05 Over the holidays I decided it's been too long since I did something with entirely no purpose. So without further ado, I present to you ...
Show HN: Ungate – use Claude and GPT subscriptions in Cursor without API costs (github.com via hn) Ungate A Cursor-first extension for using Claude, ChatGPT, and MiniMax subscriptions in Cursor instead of paying for API tokens. How it works Ungate lets you use Claude, ChatGPT, and MiniMax in Cursor through account subscriptions instead…
Which Chinese Model is best for planning and which is best for implementation? I'm currently using Opencode with an Openrouter API Key, mostly wanna decide between Kimi, GLM, DeepSeek, Qwen, Minimax and Mimo (www.reddit.com) Original plan was to use Kimi/GLM for planning and DeepSeek for implementation, but seeing a lot of love for MiMo and Minimax lately. Anyone running a planner + coder split on Opencode?
I plan to use a chinese AI model through API for coding through a harness, I'm a uni student so nothing prod related for now. should i go deepseek, minimax, kimi or glm? kinda confused (www.reddit.com) Just cancelled my claude subscription due to poor rate limits, gemini cli doesn't really excel in coding from my personal experience, and my local hardware isn't that powerful to run local AI models, and while codex is good, I wanna try so…
I built vivkemind – an open-source, local‑first terminal AI coding agent with full AWS Bedrock support (www.reddit.com) wanted a terminal AI coding agent that doesn't lock me into one model provider. So I forked Qwen Code and added full support for every model available in AWS Bedrock.
Show HN: Token Usage Meter 12 Providers and Coding Agent (qlaud.ai via hn) Here once again A Token Usage Meter for 12+ AI Providers Anthropic, OpenAI, Google, Alibaba qween, Moonshot Kimi, MiniMax, ElevenLabs, Deepgram, Perplexity. Qlaud.ai provides token usage meter / AI billing layer.
Comparing SVG Generation for the top open models (codeinput.com via reddit) Some of the larger models (like Llama) weren't available on OpenRouter, so I had to work with what was there. Best small model: Gemma 4 26B For its size, I think it had the best output.
Free llm APIs from Nvidia (www.reddit.com) So build[.]nvidia[.]com[/]models give access to free APIs for llms ranging from SLMs to frontier models. I tried building with it and let's say the APIs are so slow to respond.
Minimax vs Qwen vs Kimi vs Mimo(Omni) vs Glm ( via reddit) could not extract summary
My frustrating experience with MiniMax models! (www.reddit.com) I keep on hearing from community here that Minimax models are pretty solid, their benchmark are also always respectable but I am never able to get decent result from them. I have tried local setup (multiple harness) I have even tried their…
How does a self correcting loop for AI agents work? (www.reddit.com) Hey guys, just checked out minimax 2.7, where they used AI to train itself, and ran over a hundred loops, and it improved it's performance by 30%, how does that work, can I also run a script that makes AI store it's memory in a loop on a m…
Model API Performance (news.ycombinator.com) We’ve been benchmarking a few models on our API platform and got some interesting performance numbers: - MiniMax M2.5 → 0.118s time-to-first-token, 103 tokens/sec - GLM 5.1 → 120 tokens/sec throughput - Kimi K2.5 → 0.643s TTFT, 69 tokens/s…
What Am I Doing Wrong? Models Won't Listen, At All (GLM 5.1, MiniMax M2.7, Kimi K2.5) (www.reddit.com) What am I doing wrong here? I can't get models to follow my instructions, pretty much at all.
Near-Optimal Pure Single-Loop Extragradient Method for Strongly Convex--Strongly Concave Minimax Optimization (arxiv.org) We study smooth strongly convex--strongly concave minimax optimization with general nonlinear coupling in the deterministic unconstrained setting. We propose a pure single-loop damped extragradient method with fixed parameters and two new…
Minimax-Optimal Online Contract Design with Unrestricted Bounded Contracts (arxiv.org) We study repeated contract design when a principal observes outcomes but not the actions that generate them. The principal may use any bounded outcome-contingent payment vector, and the agent's best response can make expected profit discon…
Matching Multi-Loop Complexities with a Single Loop: Optimal Optimization Stationarity and Best-Known Game Stationarity in Nonconvex--Concave Minimax Optimization (arxiv.org) We introduce a new single-loop algorithmic framework for smooth nonconvex--concave minimax optimization. The resulting projected damped extragradient method combines projected extragradient updates, dual momentum, and a moving proximal cen…
Colla-Q: Toward Collaborative Experts in MoE Quantization via Minimax Precision Balancing (arxiv.org) In this paper, we present a Mixture-of-Experts (MoE) quantization method based on activation entropy. Although quantization reduces memory and computational costs, it can substantially degrade performance.
Stochastic Gradient Descent for Operator Learning in Hilbert Spaces: Convergence Rates and Minimax Lower Bounds (arxiv.org) This study investigates the use of stochastic gradient descent (SGD) to learn operators between general Hilbert spaces. We study weak and strong regularity conditions for the target operator that characterize its structure and complexity.
Riemannian ascent--descent for nonconvex nonconcave minimax landscapes: convergence to basin saddle points and applications to distributionally robust optimization (arxiv.org) We study a class of distributionally robust optimization (DRO) problems for the statistical risk problem, formulated as minimax problems over the product of a Euclidean space and a Riemannian manifold. Because the resulting minimax landsca…
Nearly Minimax-Optimal Regret for Linear Contextual Bandits with Arbitrary Adaptive Action Sets (arxiv.org) We study stochastic linear contextual bandits with arbitrary action menus that may depend on the fixed parameter and the interaction history. We establish matching upper and lower bounds, up to logarithmic factors.
Bias-Corrected Subspace Intersection: Minimax-Optimal Shared Subspace Estimation in Multi-View Data (arxiv.org) Estimating a low-dimensional subspace shared across noisy data matrices is a fundamental problem in multi-view matrix estimation. We study this problem under the two-view JIVE model, where each data matrix contains shared and view-specific…
Non-Adaptive 1-Bit Mean Estimation: Minimax Rates and the Sample-Interval Tradeoff (arxiv.org) We study distributed one-dimensional mean estimation under a 1-bit communication constraint. Each agent observes one sample, drawn independently from an unknown distribution, and returns a single bit in response to a query $Q: \mathbb{R}\t…
Sharp Structure-Agnostic Minimax Risk for Partial Linear Models (arxiv.org) We characterize the sharp structure-agnostic minimax risk for coefficient estimation in the partial linear model when the outcome and treatment nuisances are learned by two distinct black-box learners, which resolves the open problem in do…
Minimax Lower Bound for Estimating Diffusion-based Local Intrinsic Dimension (arxiv.org) While diffusion-based methods have recently emerged as effective tools for probing the intrinsic geometry of high-dimensional data, their statistical difficulty remains largely unexplored. We study estimation of the finite-scale population…
Erm, what? (www.reddit.com via reddit) https://preview.redd.it/v6s6tqnkrlnh1.png?width=2487&format=png&auto=webp&s=cdfdf416dc46186aa6389cc0d9d4b1eddf049cb2 My Claude did this when I asked it to write me a prompt for Minimax H3.
Minimax bounds for watermarked and masked recursive discrete distribution estimation (arxiv.org) Watermarking has been proposed as a way to identify synthetic samples in estimation settings where no metadata is available to distinguish them from real samples, but its precise effects remain unexplored. In the absence of a distinguishin…
Continuity-Free Near-Minimax Leading-Order Regret for CVaR-UCBVI (arxiv.org) For finite-horizon tabular CVaR reinforcement learning, prior work proves a $\widetilde{O}(\tau^{-1}\sqrt{SAK})$ leading regret bound for arbitrary normalized return laws and the sharper $\widetilde{O}(\sqrt{SAK/\tau})$ rate under a densit…
I asked Claude to plan a cheap 30-second AI video. It split the job across 3 models ($4.93) (www.reddit.com via reddit) Small MCP experiment: I asked Claude to plan a low-cost 30-second vertical video without sending every shot to the most expensive model. It came back with five shots: - 20s total on LTX 2.3 Fast for the establishing, transition and cutaway…
Sharp Minimax Regret for Infinite-Memory Logistic Prediction (arxiv.org) We study online prediction for a specific finite-alphabet, exogenously driven source with infinite input memory. Independent Rademacher inputs $(Ut)$ are observed sequentially, and the next binary mark has logit $\sum{j=1}^{t}\thetajU{t+1-…
Cone Extended Rayleigh Quotients for Directed Graph Learning: Minimax Spectral Certificates, Sensitivity, and Adaptive Control (arxiv.org) Directed graph learning naturally leads to trainable nonsymmetric propagation operators with distinct right and left spectral structures. Building on the two-sided cone Rayleigh framework for generalized pencils \[ B_\theta-\lambda G, \] w…
[audio.cpp] Release 0.7: 62 audio model families (85+ variants), Arena UI for model comparison, MiniMax Music 3, FireRed TTS3/Audio, ControlFoley, Personaplex, and more (www.reddit.comhttps) audio.cpp 0.7 is out :) This release adds a lot of new audio models and a new way to compare them locally. Audio.cpp is now at 62 model families and 85+ model variants.
How many models in v3 rn? (www.reddit.com via reddit) how many models are going by 3 rn like Gemini 3.7, qwen 3.8, minimax m3, kimi k3, hy3, deeseek v4 flash, glm 5.3. (not deepseek and glm but close enough)
Functional linear regression from sparse to dense designs: a pooling-ridge method and minimax optimality (arxiv.org) Functional data analysis is an important statistical field that treats data as random functions. In practice, the random functions are often not fully observed but instead measured at discrete times.
Minimax Alternating Regret for the Experts Problem and Online Convex Optimization (arxiv.org) In this paper, we study alternating regret in online convex optimization (OCO), motivated by the success of alternating learning dynamics in two-player games. Although previous works have shown that $o(\sqrt{T})$ alternating regret is achi…
When three models all claim SOTA, how do I pick for a local agent stack (www.reddit.com via reddit) I personally stopped reading the launch table once GLM-5, MiniMax M2.5, and Gemini 3 Deep Think dropped in two days and all claimed the same coding, reasoning, and agent wins. They optimize different constraints.
Two-Sided Nearest Neighbors: An adaptive and minimax optimal procedure for matrix completion (arxiv.org) Nearest neighbor (NN) algorithms have been extensively used for missing data problems in recommender systems and sequential decision-making systems. Prior theoretical analysis has established favorable guarantees for NN when the underlying…
Would anyone running a dGPU/eGPU with a Strix Halo care to share tuning tips? (www.reddit.com via reddit) I installed an R9700 in my Strix Halo machine over the weekend, via Oculink, and so far it hasn't been life-changing. First I tried running the Unsloth Q4_K_XL quant of DSv4 Flash 0731, and that failed.
Today I merged the first feature branch written entirely by my 4060Ti 16GB! (www.reddit.comhttps) Howdy folks, You might (or likely not) know me around here with shilling of Pi harness, and the use of Pi as productivity assistant and KB manager. Lately, I have also been telling anyone who listens to try Qwen 3.8 27B IQ3_K_XXS by Unslot…
Minimax is increasing their token and token plan costs (www.reddit.com via reddit) For those that don't know, Minimax is increasing their rates and token plan costs by ~60%-65% on the 25th. This is pretty much the last chance to lock in at the current rates if you use minimax and haven't yet.
Minimax Optimality of Score-Entropy Discrete Diffusion (arxiv.org) Discrete diffusion models have demonstrated strong performance across a range of datasets, including natural language data and graph-structured data. Among many variants, score-entropy discrete diffusion (SEDD) has achieved particularly st…
DGX Spark, cluster of 4 (www.reddit.com via reddit) Does anyone have a first-hand experience with four Sparks cluster, and how much of an upgrade is it comparing to just two considering the available models? While there's plenty of noise for the smaller models (Qwen) and our older king Deep…
I trained a game music generator (www.reddit.com via reddit) I trained a instrumental game music generator. The 1.2B DiT was trained on 1 cloud H100 from scratch in 8 days; I used the VAE from Stable Audio 3.
Minimax Optimal Estimator and Improved Error Rate for the MLE in Logistic Regression with Gaussian Design (arxiv.org) We study finite-sample parameter estimation in logistic regression with Gaussian design, where the goal is to estimate $\mathbf{\theta}^\in \mathbb{R}^d$ with $R=\|\mathbf{\theta}^\|2\ge 1$ from i.i.d. samples $\{(\mathbf{x}i,yi)\}{i=1}^n,…
How Many Samples Are Needed to Determine Causal Direction? Sharp Minimax Bounds for Bivariate LiNGAM (arxiv.org) We study how many observations are needed to determine the causal direction between two linearly related variables. Classical LiNGAM theory shows that independent non-Gaussian disturbances identify the direction, but does not quantify the…
Consistent Model Chasing Is Minimax Optimal: The Exact Value of Scalar Adversarial Adaptive Control under Large Parametric Uncertainty (arxiv.org) We solve exactly a fundamental problem of adaptive control against adversarial disturbances: regulate the scalar system $x{t+1} = axt + ut + wt$, $x0=0$, $\|w\|\infty \le 1$, where the constant pole $a \in [-\Delta, \Delta]$ is unknown in…
Minimax and Adaptive Covariance Matrix Estimation under Differential Privacy (arxiv.org) Estimating covariance matrices is fundamental to a wide range of statistical applications. This paper studies minimax and adaptive estimation of high-dimensional covariance matrices under $\rho$-zero-concentrated differential privacy ($\rh…
Finite-Time Minimax Bounds and an Optimal Lyapunov Policy in Queueing Control (arxiv.org) We introduce an original minimax framework for finite-time performance analysis in queueing control and propose a surprisingly simple Lyapunov-based scheduling policy with superior finite-time performance. The framework quantitatively char…
Learning with Bilevel-Minimax Optimization for Efficient and Reliable Transfer Attacks (arxiv.org) Transfer-based adversarial attacks craft adversarial examples using surrogate models to mislead black-box victim models. Beyond perturbation generation, transferability is fundamentally governed by the coupling of initialization, surrogate…
Has anyone compared MiniMax-M3 for coding-agent workflows? (www.reddit.com via reddit) I am comparing a few model options for coding-agent work and MiniMax-M3 caught my attention because it is described as supporting coding, tool use, and long-context tasks. The questions I cannot answer from the documentation are fairly pra…
Luna high weekly token experience (www.reddit.com via reddit) I am planning to use GPT-5.6 luna high as my main autonomous coding agent. Before this, I was using MiniMax M3, which gives around 1.7B tokens monthly.
Deciding When to Switch: E-Processes for Adaptive Minimax Training for Generative Adversarial Nets (arxiv.org) Modern data science increasingly gives rise to hypothesis-testing problems that are not naturally formulated in terms of parameters within prespecified statistical models. One important example is the dynamic evaluation of optimization alg…
Hierarchical Empirical-Bayes Naive Bayes: Minimax Smoothing and Calibration with AODE Extension (arxiv.org) The Naive Bayes (NB) classifier remains a standard choice for categorical data, yet its widely used smoothing rules, such as Laplace, Lidstone, Krichevsky-Trofimov, and the $m$-estimate, all prescribe a fixed smoothing strength that ignore…
Minimax-Optimal Policy Regret in Partially Observable Markov Games (arxiv.org) We study sequential decision-making in partially observable environments against strategic, adaptive opponents, modeled as partially observable Markov games (POMGs). The central challenge is to learn latent dynamics from partial observatio…
Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions (arxiv.org) Distributionally robust Markov decision processes provide a principled framework for sequential decision making under model uncertainty. We study how many samples are necessary and sufficient to learn an $\varepsilon$-optimal robust policy…
Minimax Optimal Early-Stopped Gradient Descent for Gaussian Mixture Classification (arxiv.org) In overparameterised classification, training data can be linearly separable even when the underlying distribution is not. In this setting, gradient descent (GD) on the logistic loss diverges in norm while converging in direction to a max-…
Minimax-Optimal Semiparametric Contextual Dynamic Pricing with Multimodal Revenue (arxiv.org) We study contextual dynamic pricing with arbitrary covariate sequences and bounded, possibly nonbinary purchase quantities. Demand follows a semiparametric surplus-index model with an unknown linear valuation parameter and an unknown Hölde…
PipeNetwork/minimax-h3-mlx (simonwillison.net) 4th August 2026 - Link Blog PipeNetwork/minimax-h3-mlx. MiniMax released MiniMax-H3 two days ago - they describe it as a "a general-purpose, omni-modal generative system", which in practice means it accepts text, images, audio and video an…
Online Algorithms via Minimax and Posterior Matching (arxiv.org) Competitive analysis is central to the study of online algorithms, but upper bounds are often highly problem-specific. We develop a more unifying methodology via the minimax viewpoint.
Optimizing Minimax Regret in Uncertain MDPs with Small Sets of Policies (arxiv.org) Sequential decision-making in real-world applications often involves uncertainty about the environment's model. Uncertain Markov decision processes (UMDPs) represent the possible environments as a set of MDPs with shared states and actions…
Simple-regret rates and minimax optimality of fixed-prior expected improvement in Mat\'ern and squared-exponential RKHSs (arxiv.org) We study the expected improvement (EI) policy for minimizing a deterministic objective function $f$ on a nonempty compact set $\mathcal X \subset\mathbb R^d$. We assume that $f$ belongs to the RKHS $\mathcal H_k$ of a continuous positive-s…
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence (arxiv.org) We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The flagship M2 contains 229.9B total parameters with only 9.8…
When Kernel Ridge Regression Meets the H\"older-Zygmund Class: Minimax Optimality and Failure of Properness (arxiv.org) We study kernel ridge regression for nonparametric regression over the Hölder-Zygmund class. Using an RKHS equivalent to a Sobolev space of smoothness s+d/2, we prove that misspecified KRR attains the minimax L2 rate n^{-2s/(2s+d)}.
Transfer Learning in High-Dimensional Clustering: Minimax Thresholds and Applications in Single-Cell Data (arxiv.org) Clustering is a fundamental problem in statistics, with applications across many scientific disciplines. In many modern applications involving clustering, the primary dataset (the target data) is accompanied by related datasets (the source…
Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets (arxiv.org) The tremendous success of Transformer models in fields such as large language models and computer vision necessitates a rigorous theoretical investigation. To the best of our knowledge, this paper is the first work proving that standard Tr…
Identification and Learning of Semantic Observation Kernels: Partial Observation, Uniform Recovery, & Minimax Limits (arxiv.org) Probabilistic text generators supply conditional distributions over tokens and complete verbal continuations, whereas scientific use often requires a posterior over a finite state. Large language models are the leading example: phrase prob…
Minimax Lower Bounds of Kernel Discrepancy Estimation: MMD, HSIC, KSD (arxiv.org) Over the past 20 years, kernel discrepancies have been leveraged as a highly powerful tool for quantifying the disagreement of distributions, with numerous successful applications in two-sample, goodness-of-fit, and independence testing, a…
A free LLM passed every quality check in my production benchmark, then silently returned empty output 33% of the time (www.reddit.com via reddit) I benchmarked Claude Haiku 4.5 against 8 free/alt models (NVIDIA NIM plus a second free-tier provider) on a real task: writing outreach proposals from actual job postings, not a synthetic prompt. Same production prompts, same 3 real jobs,…
↯ Haiku↯ Minimax↯ Haiku 4.5↯ Haiku 4.5↯ Haiku 4.5↯ Haiku 4.5↯ Haiku 4.5↯ Haiku 4.5↯ Haiku 4.5↯ Haiku 4.5minimaxhaiku
I kept losing my agents' work in the chat scroll. So I gave them a board - each agent gets a card, and finished work gets handed to a human to review. Built with Claude in a couple hours. (www.reddit.comhttps) I run a few agents for research and drafting. In one long chat the good output was always buried 200 messages up, I couldn't tell done vs running, and my team couldn't see any of it.
Honest Physical-Support Inference after Latent Dictionary Learning: Collision Singularities and Minimax Resolution (arxiv.org) Sparse-support uncertainty is usually quantified by treating the dictionary as known, an assumption that can produce overconfident, label-dependent conclusions when the dictionary is learned from latent sparse mixtures. Near collisions of…
Minimax and Bayes Optimal Best-Arm Identification (arxiv.org) This study investigates minimax and Bayes optimal strategies for fixed-budget best-arm identification. We consider an adaptive procedure consisting of a sampling phase followed by a recommendation phase, and we design an adaptive experimen…
Kimi moment. I think the writing is on the wall for Anthropic and OpenAi (www.reddit.com via reddit) New day new model....... can't wait to see in a few weeks how Minimax 3 Pro (2.7T parameters) and GLM 5.3 reinforces the narrative.
Fable 5 is currently ranked #10 at document generation. Every model above it is cheaper. (www.reddit.comhttps) I expected Anthropic's flagship model to be expensive but sit near the quality ceiling. The current results are considerably worse than that.
No model is perfect. Have other models weigh in for the best architecture. (www.reddit.com via reddit) Long story short, it doesn’t matter if you’re using Opus or Fable or Sol and on what level of reasoning, if you put the output into any other model, from any lab or even the exact same model, and ask for an adversarial review, it will sugg…
Minimax Theory of Likelihood-Based Deep Learning for Speckle Regression (arxiv.org) Speckle noise is a multiplicative noise commonly encountered in coherent imaging modalities such as synthetic aperture radar, optical coherence tomography, and digital holography. Although deep learning methods, in practice, have achieved…
Price of Fairness in Bandits: A Tight Minimax Characterization (arxiv.org) In bandit problems, standard regret-minimizing algorithms treat exploration as an amortized cost, which can expose early participants to unfair ex-ante losses in settings such as clinical trials. Recent work addresses this by evaluating th…
started splitting my cursor sessions between opus and m3. one thinks, the other cranks (www.reddit.com via reddit) been using cursor full time for about six months. was on opus the whole time and never questioned it because the output quality was there.
Bilateral Trade Under Heavy-Tailed Valuations: Minimax Regret with Infinite Variance (arxiv.org) We study contextual bilateral trade under full feedback when, conditionally on the context, trader valuations have bounded density but infinite variance. We first extend the self-bounding property of Bachoc et al.
First-Order Softmax Weighted Switching Gradient Method for Distributed Stochastic Minimax Optimization with Stochastic Constraints (arxiv.org) This paper addresses the distributed stochastic minimax optimization problem subject to stochastic constraints. We propose a novel first-order Softmax-Weighted Switching Gradient method tailored for federated learning.
Bandit PCA with Minimax Optimal Regret (arxiv.org) We study the bandit-feedback version of online principal component analysis (Bandit PCA): in each round $t = 1,\dots,T$, the adversary selects a $d \times d$ symmetric gain matrix $Gt$ with spectrum in $[0,1]$ and rank at most $r$; the lea…
Beyond Bayesian Nash: Learning Minimax-Regret Equilibria for Adversarial Team Games under Asymmetric Information (arxiv.org) Adversarial team games (ATGs) with asymmetric information, such as adversarial path-finding, goal search, and reachability games on graphs, require strategies that are robust to hidden opponent types, such as a hidden goal flag, and to dec…
Accelerated Fully First-Order Methods for Bilevel and Minimax Optimization (arxiv.org) We present in this paper novel accelerated fully first-order methods in \emph{Bilevel Optimization} (BLO). Firstly, for BLO under the assumption that the lower-level functions admit the typical strong convexity assumption, the \emph{(Pertu…
NL-PAC: Specification Ambiguity and Certified Minimax Risk Floors in LLM-Mediated Supervision (arxiv.org) Large language models increasingly provide labels, evaluations, and feedback for tasks specified in natural language. When a specification admits multiple readings but the supervision channel does not reveal which is operative, additional…
Fixed-Gaussian Spectral Algorithms: Minimax Optimal Rates for Misspecified Learning and Transfer (arxiv.org) The principal objective of this work is twofold within nonparametric regression settings: (1) to establish the minimax optimal convergence rates for fixed-bandwidth Gaussian kernel spectral algorithms when the true regression function resi…
Minimax Estimation of Kernel Stein Discrepancy: Trace versus Hilbert-Schmidt Scales (arxiv.org) Kernel Stein Discrepancy (KSD) compares a sample to a fixed target distribution known only through its score, and is widely used for goodness-of-fit testing, sample quality assessment, and approximate inference. We study the estimation of…
Adversarial Contamination Meets Hard Thresholding: An Iterative Algorithm with Signal Adaptivity and Minimax Optimality (arxiv.org) Pervasive data contamination -- stemming from measurement errors, outliers, or adversarial corruption -- has motivated the development of robust statistical methods. In this context, we propose a two-stage Adversarial Contamination-resista…
Be honest guys, how reliable are your agents for building softwares? (www.reddit.comhttps) I'm an SDE and I feel like I'm not getting much productivity out of coding agents. Yeah, they can generate code and build features, but most of what I get isn't really deployment-ready or easy to maintain.
What models are you all using lately? I find Minimax M3 useful for most coding and price quality to make the most sense. I’m using it via openrouter in cursor along with composer 2.5 for the heavier stuff. I’m building a Shopify app with TypeScript, Prisma, Polaris. Are there any better combos? ( via reddit) could not extract summary
Are third-party memory systems actually better than the built-in memory_wiki in Openclaw? (www.reddit.com via reddit) When I first installed openclaw, I immediately set up an obsidian vault. When they added the memory_wiki plugin, I migrated everything to that and deleted obsidian..
Minimax PAC Bounds for Learning in Exogenous Contextual MDPs (arxiv.org) We study PAC learning in tabular discounted Markov decision processes with exogenous i.i.d. contexts, with discount factor $\gamma$, finite state space $\mathcal X$, action space $\mathcal A$, and context space $\mathcal Z$.
Black-Box Assisted Regression: Phase Transitions and Minimax Optimality (arxiv.org) Foundation models are often used as fixed black-box predictors for downstream tasks with limited labeled data, but their predictions may be biased and unsafe to trust blindly. We study this setting through black-box assisted nonparametric…
Minimax Limits of k-Fold Cross-Validation via Majority (arxiv.org) We study the mean-squared error of $k$-fold cross-validation as a risk estimator, with particular emphasis on how its accuracy depends on the number of folds $k$. Despite the widespread use of cross-validation, principled guidance for choo…
Minimax Quantile Lower Bounds for Interactive Statistical Decision Making with Privacy (arxiv.org) Minimax risk and regret are expectation-based criteria and do not capture rare but consequential failures. To address this concern, we develop a $\delta$-explicit minimax-quantile theory for interactive statistical decision making (ISDM).
Claude Cowork 3P Gateway returned no usable models. Add entries under Models to test inference without discovery error (www.reddit.com via reddit) Hello, I'm having this problem with Cowork 3P, I cannot use Minimax model with Cowork 3P. This API key works when I use it with Hermes Agent Desktop, but it does not work with Cowork whatever I do.
GLM 5.2 and MiniMax M3 are a lot closer/better to Sonnet 4.6 than I expected on coding-agent workloads (www.reddit.comhttps) We benchmarked GLM 5.2, MiniMax M3, Kimi K2.7-code, Qwen 3.7-Plus and Sonnet 4.6 across nearly 1,000 coding-agent scenarios. The scenarios were run twice.
Quantile of Means: A Bonus-Free Ensemble Method for Minimax Optimal Reinforcement Learning (arxiv.org) Optimal Reinforcement Learning (RL) algorithms typically rely on carefully constructed count-based uncertainty estimates to drive exploration. Although theoretically sound, such estimates are hard to compute in practical settings and there…
PM tried M3's 1M context on a real Q3 brief: where it held, where it broke (www.reddit.com via reddit) I'm a PM, not a researcher. My job is pulling 12-18 sources into one strategy doc and not losing the caveats.
Learning from Biased and Costly Data Sources: Minimax-optimal Data Collection under a Budget (arxiv.org) Data collection is a critical component of modern statistical and machine learning pipelines, particularly when data must be gathered from multiple heterogeneous sources to study a target population of interest. In many use cases, such as…
Enhancing LLM Safety Through a Theoretical Minimax Game Lens (arxiv.org) The rapid advancement of large language models (LLMs) necessitates effective mechanisms to ensure their responsible deployment by accurately distinguishing unsafe content from benign content. While substantial safety datasets are available…
Reviewing speed optimizations on llamacpp for large MoE models on multiGPU rigs? (fitparams vs -ngl/-ncmoe vs other flags, P2P, overclocking) (www.reddit.com via reddit) In anticipation of MiniMax reported upcoming open-weight release of M3, wanted to do comprehensive review of what I’m aware of regarding speed optimizations. Hopefully it can be helpful reference for some people too.
As we know Minimax M3 is just going to be open sourced in few days and because of that I was surfing on internet searching for its scores and I found out pretty interesting results. Is Minimax M3 really that good in agentic stuff and in coding? Is it better than older gpt models? (www.reddit.com via reddit) Has anyone personally compared the Minimax M3 model against other proprietary models to determine its relative performance tier? I am trying to understand where it currently ranks in the broader Al landscape.
Minimax M3: Are they capping about open weight? I can't find the download link anywhere (www.reddit.com via reddit) https://www.minimax.io/blog/minimax-m3 They advertise it as open weight and have these words everywhere in their advertisements, but they have not released it.
Can you really replace paid models with a local model? (www.reddit.com via reddit) Long time lurker, and I say this as someone who genuinely loves this community and runs many local models myself. I’ve been using LLMs since the early GPT and LLaMA days.
Agentic Setup: Minimax 2.7 vs qwen 3.6 (www.reddit.com via reddit) I'm currently using Minimax 2.7-AWQ-4bit for an specific coding agentic workflow. I see many of you are currently using Qwen3.6 and wanted to know how does it compare with Minimax2.7 .
Algorithmic and Minimax Complexities in Kernel Bandits (arxiv.org) Claude Fable/Mythos 5 just came out, so it will take Deepseek or Z.ai or Xiaomi or Kimi 9-12 months to release a model just as good as Fable? (www.reddit.com via reddit) It should be at least 7-8 months until we have an open Fable(not just as good as Fable in benchmarks, but actually as good as Fable), probably more like 9-12 months. By the time, an open Fable model comes out, Fable 6.5-7 will be way bette…
MiniMax is digging its own grave (www.reddit.com via reddit) A Temporal Spatial Minimax Rate for Smoothly-Varying Distributions in Wasserstein Space (arxiv.org) We study the minimax rate of estimating a future value $\mu{tn+h}$ of a curve $t\mapsto\mut$ in the $2$-Wasserstein space $\mathcal{P}2(\mathbb{R}^d)$ from finitely many noisy snapshots of its past, under an adiabatic bound $\|\nablat^k v\…
Generalization in Deep Neural Networks: Minimax Rates for Gradient Methods (arxiv.org) Understanding the generalization performance of over-parameterized neural networks has become a central topic in deep learning theory. While recent advances, particularly works under the Neural Tangent Kernel (NTK) regime, have shed light…
Running a 24/7 AI agent dev team: I route each role to a different LLM (Claude/Kimi/MiniMax/GPT) to dodge a ~$2k/mo API bill. Setup + what actually breaks. (www.reddit.com via reddit) Context: I run an autonomous engineering "org" of AI agents on my own product. Once it grew past ~5 agents and started running around the clock, it maxed my Claude Max weekly limit by mid-week.
AA comparison of the latest local models (www.reddit.com via reddit) I picked models I consider local (usable on 3×3090), so there are no 300B models, and you should probably skip 200B models too (but MiniMax and Step are pretty fast in Q3) Gemma-4 12B is still missing
Anyone has experience between Mimo flash v2.5 pro vs Composer 2.5 (cursor pro+) (www.reddit.com via reddit) I have Mimo subscription alongside Claude Code Max. You won’t believe how suck Claude Opus can be at certain task but it does get more job done than any other model I have tried.
Minimax optimal differentially private synthetic data for smooth queries (arxiv.org) Literature-Guided Minimax Optimization of Virtual Epilepsy Neurostimulation (arxiv.org) I went from 1 to 10 apps on the App Store in 4 months - vibe coding as a senior iOS dev (www.reddit.com) I code for 20 years and make mobile apps for 15+. This February I decided to try vibe coding, but at scale.
dual spark with llama.cpp (www.reddit.com) I'm daily driving dual Asus GX10 (spark) with vllm and it's fantastic. But I want to try model that is GGUF only and won't fit into single spark.
My 1.2B model won 2 out of 5 poker tournaments against models up to 1T params. (www.reddit.com) I made 6 LLMs play Texas Hold’em against each other. Ran 5 tournaments on my 16GB MacBook.
We built Irene — an AI agent platform that actually remembers you, builds its own tools , adapts and improve as you use it (www.reddit.com) Hey r/AI_Agents — we're launching Irene today, and I want to be straight about what it is, why we built it, and where it's going. What makes Irene different Affordable with massive token limits and the latest open-source models We have gen…
Spec decoding for minimax m2.7? (www.reddit.com) MTP was not released for m2.7, so would anyone have experience with setting up speculative decoding for minimax m2.7 and its results? Whether via EAGLE3 or a distilled variant
Opus 4.6 is Vicious (www.reddit.com) This is the hardest I've ever seen it riff. Full shared link at the bottom, but here are some highlights.
[Research use case] MiniMax-M2.7 with small context, CPU+GPU (5090) setup on Llama.cpp (www.reddit.com) I was experimenting yesterday with running oversized models with smaller context size, hoping that leaving them overnight could compensate for the slow token generation and periodic pauses for compaction or task chunking. Summary: For rese…
Is it possible to edit LLAMA.CPP with Cline+Vscode+Minimax 2.7 Q4_K_S and get a working build? (www.reddit.com) It all started yesterday with this post by u/antirez https://www.reddit.com/r/LocalLLaMA/comments/1sw3stb/llamacpp_deepseek_v4_flash_experimental_inference/ I was intrigued by the first Deepseek V4 Flash GGUF in a small size that can fit o…
What's the smallest reasonable quant for coding? (www.reddit.com) What’s your LLM routing strategy for personal agents? (www.reddit.com) TL;DR I try to keep most traffic on very cheap models (Nano / GLM‑Flash / Qwen / MiniMax) and only escalate to stronger models for genuinely complex or reasoning‑heavy queries. I’m still actively testing this and tweaking it several times…
Use this prompt if you want to find a specific info off the Internet with lowest wrong answer possiblity. Works best for ~30b models. (www.reddit.com) For context i used to ask many near 30b model this question --> **^(Calculate the precise VRAM requirement for the \*KV Cache only** at the maximum context window for **DeepSeek V3.2** and **MiniMax M2.5**. * **DeepSeek V3.2 Max Context:**…
Need suggestions for local AI Machine (www.reddit.com) I’ve been running various AI harnesses like OpenClaw, ForgeCode, ClaudeCode, etc. Most of these are running via OpenRouter or Minimax (credits/subscription model).
But why Local LLM? How does this make economic sense vs API? (www.reddit.com) Hey guys, come fight me: how do you justify local LLMs from a value perspective? It doesn't seem economical?
I made a simple proxy to let Claude use MiniMax models as subagents (www.reddit.com) I made this due to the usage problem. Enjoy and tell me what you guys think!
Optimizing MiniMax 2.7 - Experts vs Layers for best VRAM/RAM utilization (www.reddit.com) I'm curious if there is a rule of thumb regarding how to best load Minimax given varying amounts of VRAM/RAM configurations. Is there a way to estimate how many experts versus layers to offload for individuals running either 16GB/24GB/32GB…
Best setup for MiniMax-M2.7 (230B) | 3x RTX 5090 | Threadripper 9975 | 512GB RAM (www.reddit.com) I have the following hardware and want to run MiniMax-M2.7 (230B) locally. What is the best software stack and configuration to maximize performance?
Mac Studio Performance Suggestion For minimax (www.reddit.com) I need help. I want to self-contain my MiniMax 2.7 and Qwen 3.5 (122 billion parameter) models.
Why most open-source models can't answer this question while most closed-source models can answer most of the time? (www.reddit.com) WEB SEARCH WAS ALWAYS ON!!!! Question Calculate the precise VRAM requirement for the **KV Cache only** at the maximum context window for **DeepSeek V3.2** and **MiniMax M2.5**.
Stop donating your salary to OpenAI: Why Minimax M2.5 is making GPT-5.2 Thinking look like an overpriced dinosaur for coding plans. (www.reddit.com) ↯ Hallucination↯ Glm↯ Minimax↯ Swe Benchswe-benchminimaxaltman+5
Aligning to What? Rethinking Agent Generalization in MiniMax M2 (huggingface.co)