Something keeps nagging at me about the Chinese AI space lately. Every few months a new Chinese model drops that closes the gap with US frontier models a little more(not by throwing more compute at it, just genuinely clever engineering at…
#glm
378 items
Chinese AI companies are shipping faster and cheaper than anyone expected and I'm not sure the west has a good answer for it (www.reddit.com) Major drop in intelligence across most major models. (www.reddit.com) As of mid Apr 2026, I have noticed every model has had a major intelligence drop. And no I'm not talking about just ChatGPT.
2x 512gb ram M3 Ultra mac studios (www.reddit.com) Zai replaced the network architecture running GLM-5.1 inference and the gains are pretty wild (www.reddit.com) Been following the infrastructure side of AI more lately and stumbled on this from Zai. They upgraded the network architecture on a thousand-GPU cluster running GLM-5.1 coding inference from the standard ROFT setup to something they built…
GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost (reinvently.co.uk via hn) Seventeen models on the Featherbench leaderboard: glm-5.3 leads at 100%, eight tie at 96%, and checker errors flip the safety ranking. See who tops the board.
GLM-5.3: Frontier Coding with Emergent Cyber Capabilities (z.ai via hn) could not extract summary
GLM-5.3 is now open-weight (twitter.com via hn) Z.ai on X: "GLM-5.3 is now open-weight. Our most capable model for agentic coding and cyber defense is now available to download, run, and customize.
ZCode: Claude Code from the Makers of GLM (zcode.z.ai via hn) GLM Coding Lite 适合轻量开发任务 包含基础使用额度 - 适合轻量迭代与小型仓库 - 持续获得最新旗舰模型与功能 - 支持 20+ 编程工具,包括 ZCode 深度适配 GLM Coding Plan GLM 深度适配 ZCode,让 Agent 编程更稳定、更高效。 GLM Coding 适合轻量开发任务 包含基础使用额度 GLM Coding 适合专业开发工作流 包含 Lite 全部权益与 5 倍 Lite 额度 GLM Coding 适合高频与大规模任务…
I'm glad we have deepseek (www.reddit.com) other companies are slowly going away from open weight, not releasing base models, delaying open weight distribution, not releasing top models (this one I think is fair, but still), and I also noticed they stopped publishing research (old…
Recent Open models from last 6 Months - Nov 2025 - Apr 2026 (www.reddit.com) I created this chart with recent open models from last 6 months. Few might be older than that possibly.
Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models (news.ycombinator.com) I’ve been building Echo (https://echo.tracerml.ai/), an experiment in making one AI system out of a pool of open-weight models rather than choosing a single model and using it for every task. It started with a simple experiment.
GLM-5.2 is the new leading open weights model on Artificial Analysis (artificialanalysis.ai via hn) June 17, 2026 GLM-5.2 is the new leading open weights model on the Artificial Analysis Intelligence Index Z ai’s GLM-5.2 is the new leading open weights model on the Artificial Analysis Intelligence Index scoring 51 and it sits on the Pare…
Do you guys think there’s a high chance of Singularity being open source? (www.reddit.com) GLM 5.1 is dominant in almost every aspect in Design arena, surpassing Opus 4.6 in many tasks. Although user experiences vary dependent on subscription plans for both of those one of them is open source.
(Interactive)OpenCode Racing Game Comparison Qwen3.6 35B vs Qwen3.5 122B vs Qwen3.5 27B vs Qwen3.5 4B vs Gemma 4 31B vs Gemma 4 26B vs Qwen3 Coder Next vs GLM 4.7 Flash (www.reddit.com) Minimax M2.5 vs. GLM-5 vs. Kimi k2.5: How do they compare to Codex and Claude for coding? (www.reddit.com) Guys we have to change the pelican test (www.reddit.com) So i have been seeing more of those pelican on a bike svg tests and while they work i feel like (and maybe you guys do too) they are getting kinda benchmaxxed so we should switch things up soon and this is my idea generate me a html svg of…
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents (arxiv.org via hn) We present GLM-5V-Turbo, a step toward native foundation models for multimodal agents. As foundation models are increasingly deployed in real environments, agentic capability depends not only on language reasoning, but also on the ability…
GLM-5.2: The Most Powerful Open Model yet and the Brutal Reality of Running It (vettedconsumer.com via hn) Every few weeks the "best open model" crown changes hands. This week it's GLM-5.2, from the Chinese lab Z.ai — and unusually, the claim has teeth: it sits at #1 on the independent Artificial Analysis Intelligence Index.
ZAI might stop open-weighting their models? (www.reddit.com) Ever since the company went public, they’ve been making a lot of changes that clearly seem to be prioritizing profit without regard to their customers. For example, with their coding plans: - They promised/advertised that the Lite coding p…
All major LLMs are lib-left. Even Grok, half the time (unslop.run via hn) I ran the 62-item politicalcompass.org test 30 times each on sixteen models: OpenAI's GPT-5.x and GPT-4o, Claude, Gemini, Grok, Llama, Mistral, and China's DeepSeek, Qwen, Kimi and GLM. Fifteen land in the libertarian-left quadrant.
Ox-Alpha Is GLM (dejan.ai via hn) Prompt injection and gzip-NCD compression analysis reveal that OX Alpha, a mysterious LLM on OpenRouter, is GLM developed by Z.ai. A stealthy new model called OX Alpha has popped up on https://openrouter.ai/ and is climbing up the leaderbo…
Running gpt and glm-5.1 side by side. Honestly can’t tell the difference (www.reddit.com) So I have been running gpt and glm-5.1 side by side lately and tbh the gap is way smaller than what im paying for On SWE-Bench Pro glm-5.1 actually took the top spot globally, beat gpt-5.4 and opus 4.6. overall coding score is like 55 vs g…
Abliterlitics: Benchmarks and Tensor Comparison for Heretic, Abliterlix, Huiui, HauhauCS for GLM 4.7 Flash (www.reddit.com) This is a follow up to the previous benchmark and tensor analysis of abliteration techniques across the Qwen model family. Same approach, same toolkit, new model family.
The pacman benchmark: finally a viable local agentic coding agent with Qwen 3.6 27b (www.reddit.com) One way I like to test new models, is by one-shoting (with a good prompt) a single webpage clone of the classic arcade game pacman. I usually do 3 attempts and keep the best one.
Tested how OpenCode Works with SelfHosted LLMS: Qwen 3.5, 3.6, Gemma 4, Nemotron 3, GLM-4.7 Flash - v2 (www.reddit.com) I have run two tests on each LLM with OpenCode to check their basic readiness and convenience: - Create IndexNow CLI in Golang (Easy Task) and - Create Migration Map for a website following SiteStructure Strategy. (Complex Task) Tested Qwe…
Lossless model compression experiment: GLM-5.2 in 25% less memory (brianbell-x.github.io via hn) A full GLM-5.2 scan found 30.168% K15 charged-format accounting. A separate byte-split representation was decoded bit-for-bit across all 59,509 BF16 tensors at 24.967% reduction.
Ask HN: What was the last task where only a frontier model could do it? (news.ycombinator.com) ive been seeing a recurring claim that open (weight) models 6 months behind the frontier are good enough for the majority of ‘work’. if you've had a concrete task in the last month where GLM/DeepSeek/Kimi/Qwen failed and Opus/Fable/GPT suc…
GLM-5.2: Frontier Intelligence, Open Weights (twitter.com via hn) Introducing GLM-5.2: Frontier Intelligence, Open Weights - Significant improvements in coding and agentic tasks - Strong long-horizon capabilities with a 1M context window - Two levels of reasoning effort: GLM-5.2 (max) pushes the limits,…
Smaller, faster, safer: running Kimi and GLM at scale (blog.cloudflare.com via hn) Smaller, faster, safer: running Kimi and GLM at scale Workers AI runs inference for some of the best open models in the world on GPUs in Cloudflare data centers close to your users. Two of the most capable, and most demanding, are Moonshot…
Is Qwen3.6 current king for local agentic use? (www.reddit.com) I've been testing other models but it seems like nothing even come close to Qwen3.6 35B A3B for agentic use. The worse I'd get is a loop sometimes, while Gemma4 produced broken tool calls occasionally and I couldn't even get GLM 4.7 Flash…
Show HN: 143.dev – we open-sourced our internal coding-agent infrastructure (news.ycombinator.com) We just open-sourced the internal system we built at Assembled for running coding agents as a team. Coding agents worked well for individual engineers, but the surrounding workflow was a bit of a mess.
Single question llm comparison (www.reddit.com) Kimi K3 and GLM 5.2 can create undetectable malware for $2 (www.incalmo.ai via hn) The danger frontier: low-cost, evasive, abundant malware As part of Incalmo’s mission to make AI safely ubiquitous, we do safety research on the frontier cyber capabilities of models. Recently, to help anti-virus systems stay ahead of the…
Coinbase Switches to Chinese AI Models GLM and Kimi, Cuts AI Spending by 50% (mlq.ai via hn) Coinbase Switches to Chinese AI Models GLM and Kimi, Cuts AI Spending by 50% - Coinbase has defaulted engineers to GLM 5.2 from Zhipu and Kimi 2.7 from Moonshot AI through its internal LLM gateway, cutting AI spending by nearly 50% [1] - G…
Kimi K2.6-Code-Preview, Opus 4.7, GLM 5.1, Minimax M2.7 and more tested in coding (www.reddit.com) Hi everyone. It's been a while since I posted (was a lil burned out), but some of you may have seen my older SanityHarness posts.
What We Learned Moving Our Agent Loops from Anthropic to GLM (getunblocked.com via hn) What We Learned Moving Our Agent Loops from Anthropic to GLM Why we moved most of Unblocked's agent traffic from Claude Opus to GLM 5.2, what the blind A/Bs and the ledger actually showed, and what broke on the way. TL;DR: We moved most of…
I expanded DystopiaBench to 42 models and 6 dystopia types. Claude is still the only one I'd trust with nuclear codes. (www.reddit.com) Since the last post I've added: Huxley module (Brave New World style behavioral conditioning) Baudrillard module (synthetic intimacy, trust collapse, simulation) 30 more models including Grok 4.3, GPT-5.5, Gemini 3.1 Pro, GLM-5.1 Multi-jud…
GLM 5.2 is nearly as accurate as a human book keeper (toot-books.pages.dev via hn) GLM 5.2 is (nearly) as accurate as a human book-keeper at less than 1% of the cost We evaluated the performance of GLM 5.2, an open weights AI model, on the task of quarterly value-added tax (VAT) return preparation for a small UK business…
GLM 5.1 Locally: 40tps, 2000+ pp/s (www.reddit.com) After some sglang patching and countless experiments, managed to get reap-ed nvfp4 version running stable and FAST on 4 x RTX 6000 Pros (limited to 350W). Very happy with performance and quality.
Kimi K3 and GLM-5.3 are better than Gemini 3.8 Flash (news.ycombinator.com) Based on artificialanalysis.ai, Kimi K3 and GLM-5.3 are more intelligent than the new Gemini 3.8 Flash[1][2]. Gemini 3.8 Flash comes in eighth place with a score of 59, just after Kimi K3 and GLM-5.3, with scores of 60 for both of them.
GLM-5.2: Chop off 84% of the volume from a 1.5TB model, still retain 82% power (twitter.com via hn) Introducing GLM-5.2: Frontier Intelligence, Open Weights - Significant improvements in coding and agentic tasks - Strong long-horizon capabilities with a 1M context window - Two levels of reasoning effort: GLM-5.2 (max) pushes the limits,…
GPT 5.5 (Codex) leading the future prediction race (www.reddit.com) Researchers from the Max Planck Institute recently released FutureSim, an environment in which agents are replayed a temporal slice of the web and are tasked with predicting real-world future events. In their environment, GPT 5.5 leads at…
Your local LLM predictions and hopes for May 2026 (www.reddit.com) Which of these do you think we'll get in May? Also, feel free to pick/rank which ones you'd want the most badly: more Gemma4 models (124b?) (other sizes?) more Qwen3.6 models (9b?
Comparing GPT-5.4, Opus 4.6, GLM-5.1, Kimi K2.5, MiMo V2 Pro and MiniMax M2.7 (www.codejam.info via hn) Local GLM 5.1 - Parkour! (www.reddit.com) Some more 'sloptuber' content for those who are enjoying it :) Model: unsloth glm 5.1 @ IQ2_XXS UD Prompt 1: Task: in a single web page, build a city based parkour game. wsad controls, moving player aligned with current camera direction.
Ollama Cloud Pro ($20/mo) vs OpenAI Plus ($23/mo). Which gives more tokens ? (www.reddit.com) Hey everyone, I'm comparing these two plans side by side for running AI agents daily through OpenClaw (self-hosted AI agent platform): • Ollama Cloud Pro — $20/month • OpenAI Plus — €23/month (~$25) My setup: 3 agents running in parallel (…
10-task GLM 5.3 harness bench: Claude, OpenCode, pi, zcode, Hermes and 3code (capocasa.dev via hn) 10-task GLM 5.3 harness bench: claude, opencode, pi, zcode, hermes and 3code I'm performing a series of harness benchmarks on the same 10 SWE-bench verified tasks representatively chosen for difficulty. This is far from a perfect measure a…
↯ Glm↯ Swe Bench↯ GLM 5.3↯ GLM 5.3↯ GLM 5.3↯ GLM 5.3↯ GLM 5.3swe-benchglm
Run GLM-OCR, DeepSeek-OCR-2, Dots.mocr with an OpenAI Compatible API (www.vlm.run via hn) Multimodal, Multitask One catalog spanning multi-modal inputs and multi-task outputs: OCR, detection, segmentation, pose, keypoints, and more. Document OCR, captioning, and multi-modal chat: every visual capability behind one MCP server.
A production-grade OCR pipeline on Kubernetes with vLLM and Rust (github.com via hn) 📄 Production-Grade SLM-Powered OCR Course 📄 Build a self-scaling, event-driven OCR pipeline on Kubernetes (AKS / GKE) with Qwen 3.5 + the GLM-OCR SDK Table of Contents Table of Contents Course Overview Who is this course for? Course Breakd…
Hey GLM 5.2, build me a hypervisor (technotes.substack.com via hn) Comparing GLM 5.2 on several long horizon tasks including systems programming, web, creative writing and video generation and that involves working with multiple programming languages.
GLM-5.2 is the step change for open agents (www.interconnects.ai via hn) GLM-5.2 is the step change for open agents A capability threshold I've been carefully monitoring. Housekeeping: Following my “State of the blog” post last week, noting a slight increase in paid features, it’s a good time to remind folks th…
do you use different models for different steps in your agent, or just one for everything? (www.reddit.com) Our dev team flagged last week that xAI is retiring grok 4.1 fast. We weren't using it for anything critical but it made me ask something I'd never actually asked: how did we pick the models we're running?
DeepSeek's 10T USD grand strategy (twitter.com via hn) Have you ever wondered, how DeepSeek may make money, and lot of it? They didn't come up with competitive coding plans like GLM, MoonShot and MiniMax.
Tips for using Composer 2? New to Cursor (www.reddit.com) Hi. I new to using Cursor - coming from Claude Code, Antigravity and most recently GLM coding plan.
Anyone tried +- 100B models locally with foreign languages? (www.reddit.com) I am quite curious as I tried Gemma 4 31B, Qwen 3.6 27B, GLM 4.7 30B and some others in my native language (czech). Gemma performs "best" and considering the fact its "just" 18GB model - it actually blows my mind how well it can respond in…
Scaling Pain of Coding Agent Serving: Lessons from Debugging GLM-5 at Scale (z.ai via hn) Our belief in Scaling Laws has not only driven continuous breakthroughs in model parameters and data scale, but has also pushed infrastructure engineering toward its limits. This process inevitably comes with growing pains, which we refer…
Used a Claude Code skill to fine-tune Qwen3-1.7B from 327 noisy traces, matches GLM-5 (www.reddit.com) Had 327 production traces from a restaurant-reservation agent I wanted to retrain. The plan was to fine-tune a smaller self-hostable model so I could ditch the frontier-API bill.
GLM-5.3 Uncensored (huggingface.co via hn) GLM 5.3 CRACK — Uncensored FP8 General-purpose weight-level uncensoring · native FP8 speed on Hopper a CRACK release by dealignai · Twitter @dealignai What this is Full-spectrum general-purpose uncensor of GLM-5.3-FP8. Refusal behavior is…
Interactive Model View zai-org/GLM-5.3 (hfviewer.com via hn) This model is featured because its Hugging Face README.md includes an hfviewer architecture visualization. Architecture graph for zai-org/GLM-5.3.
Benchmarks of rumored Mythos level model from Zhipu AI (twitter.com via hn) Ox Alpha is 100% a GLM model by Zhipu AI, and it looks like its *almost* mythos class from very early results. It's very likely going to be called GLM-6 and it's absolutely mogging every frontier model in SWE and Cyber benchmarks (DeepSWE…
DeepSeek announced to raise its API price tremendously (news.ycombinator.com) My take: this notice from DeepSeek might actually be a brilliant marketing move. The logic here is to urge users to ramp up their usage over the next 2 to 3 months.
GLM-5.3 Soon (github.com via hn) Z.ai Open Platform Java SDK 中文文档 | English The official Java SDK for Z.ai platforms, providing a unified interface to access powerful AI capabilities including chat completion, embeddings, image generation, audio processing, and more. ✨ Fe…
Autoresearch doubled GLM-5.2 throughput. Production traffic broke it (fparisio.substack.com via hn) An AI-agent cold-tuned our GLM-5.2 serving. Human engineering leveled it up for real production traffic.
ZCode: GLM-5.2's own harness is officially live (twitter.com via hn) Introducing ZCode, the official development environment for GLM-5.2 - GLM Coding Plan subscribers: now 1.5x usage quota in ZCode - BYOK supported: works with your existing subscriptions and APIs - Available on macOS, Windows, and Linux D…
GLM-5.2 vs. Claude Opus: Same Code, Less Than Half the Cost (entelligence.ai via hn) GLM-5.2 vs Claude Opus: Same Code, Less Than Half the Cost We ran GLM-5.2 head to head with Claude Opus the way an agent actually runs: inside a real coding agent, in a real shell, graded by hidden tests. The harness is Claude Code on term…
I built a local GUI for the TradingAgents framework — works with Ollama (www.reddit.com) https://preview.redd.it/i90oxxk7n03h1.png?width=1898&format=png&auto=webp&s=7d219c804fda7dfe122b84fcdb6d0d6883818c68 A while back I came across TradingAgents — a really cool multi-agent LLM stock analysis framework where like a dozen "agen…
Best AI coding plan alternative to Claude and ChatGPT (news.ycombinator.com) With the lowering usage limit in Claude, I am thinking of jumping ship to Chinese AI, since the benchmark is already very near compared to Sonnet or Haiku 4.5 , but for a fraction of the price. I am not worried about where is my data endin…
Update to the LLM Debate Benchmark: GPT-5.5, Grok 4.3, DeepSeek V4 Pro, GLM-5.1, Kimi K2.6, Qwen 3.6 Max Preview, Xiaomi MiMo V2.5 Pro, Tencent Hy3 Preview, and Mistral Medium 3.5 High Reasoning added (www.reddit.com) The benchmark uses adversarial, multi-turn debates across 683 curated motions. Each model pair debates the same motion twice with sides swapped.
Current state of open-source ? (www.reddit.com) I’m trying to understand the current open-source LLM landscape beyond surface-level hype. We all got used to the nerfed products of Claude/Geminj so I believe really in opensource as a solution.
llama.cpp / ik_llama MoE Expert Offloading - Main Memory Bandwidth vs. PCIe Bandwidth (www.reddit.com) GLM-5.3-FlashX: Delivering inference speeds of 200 tokens/s (docs.z.ai via hn) Model Overview GLM-5.3-Flash/GLM-5.3-FlashX is the first native multimodal model in the GLM-5 series, delivering stronger intelligence than GLM-5.2 at an exceptionally low cost.- Highly Efficient Hybrid Architecture - Native Multimodal Vis…
ZCode, the GLM coding agent, silently uploads your Git history (tokenstead.ai via hn) On September 18, 2026, a developer going by ferstar published a reverse-engineering walkthrough of ZCode, the AI coding desktop app from Z.ai, the Beijing-headquartered company behind the GLM family of open-weight models - the same models…
GLM 5.3 is live on Mistral (docs.mistral.ai via hn) September 15, 2026Blog Public PreviewThird-partyv5.3 Z.ai GLM 5.3 A third-party open source text model from Z.ai, hosted by Mistral for long-context coding and agentic workflows. The model is served without Mistral modifications.
Fast weights and sparse attention in GLM-5.3-Flash (idlemachines.co.uk via hn) Attention makes the sequence all equally available, but KDA requires the model to turn a sequence into a finite state. This compression is naturally lossy, but it forces the model to extract relevant patterns in the context, and more to th…
The fastest and cheapest GLM 5.3 Flash endpoint (runinfra.ai via hn) zai-org/GLM-5.3-Flash GLM 5.3 Flash is an LLM listed in RunInfra Model APIs. RunInfra serves it as zai-org/GLM-5.3-Flash at $0.10 per 1M input tokens and $0.40 per 1M output tokens.
Show HN: Use GLM-5.3 in Cursor today via tokengo API (www.tokengo.com via hn) Hi HN, Hearing a lot of buzz around GLM-5.3, which I expect to be the best open-source coding model with the weights dropping soon, I wanted to test it where I actually do my work. I just mapped GLM-5.3 to TokenGo so I could swap out the b…
zai-org/GLM-5.3-Flash (huggingface.co via hn) GLM-5.3-Flash 👋 Join our WeChat or Discord community. 📖 Check out the GLM-5.3-Flash blog and GLM-5 Technical report.
I spent $266 and four AI models to own my tablet. GLM-5.3 finished it in a day (ericpardee.github.io via hn) Amazon kept shutting down my tablet, so I spent $266 on four AI models to own it My Amazon Fire HD tablet cost $114.26 on eBay in November 2022, new and sealed. Owning it for real cost another $266.15: Kimi K3 found the exploit for $164.25…
GLM-5.3 – The Official Desktop Auditor for Z.AI's Cyber-Engine (github.com via hn) GLM-5.3 — The Official Desktop Auditor for Z.AI's Cyber-Engine Z.AI’s GLM-5.3 didn’t just beat the industry benchmarks—it rewrote them. By applying extreme reinforcement learning on real-world cybersecurity tasks, GLM-5.3 autonomously disc…
GLM-5.3: How Chinese labs keep stride with the frontier (www.interconnects.ai via hn) GLM-5.3: How Chinese labs keep stride with the frontier Hint: It’s really not a distillation story. Housekeeping: I’m traveling so cannot make a voiceover for this post.
Hugging Face uses open-weights Z.ai GLM 5.2 to battle attacker (siliconangle.com via hn) Hugging Face uses open-weights Z.ai GLM 5.2 to battle attacker after commercial frontier model refusal Hugging Face Inc., an open-source artificial intelligence platform often described as the “GitHub of machine learning,” found itself for…
Show HN: InstantVideos.org – short documentaries in ~30 seconds (instantvideos.org via hn) Hiya! So I've been playing around with having Claude make videos for a bit now even had some success posting the results to TikTok (and setup a whole pipeline so Claude can generate and post autonomously).
Ask HN: What will you work on when Fable 5 comes back online today? (news.ycombinator.com) Saw this last night: https://x.com/AnthropicAI/status/2072163884430229756 I'm taking the weekend off to spend time with my lady, so got up early today to start preparing to get back to work on a sailboat simulator I started building when F…
Relay – open-source coding agent for non-mainstream/Chinese LLM providers (github.com via hn) Relay An open-source, dark-mode desktop coding agent — built for people who want to use non-mainstream LLM providers, not just the big three. Relay is an Electron app that puts DeepSeek, Qwen, GLM, Kimi, MiniMax, and other open/Chinese mod…
GLM-5.2 Is the New Best Open Model (thezvi.wordpress.com via hn) GLM-5.2 arrived last week. It boasts excellent benchmarks and looks strong.
GLM-5.2 Beat Fable 5 at Website Design (twitter.com via hn) https://t.co/JSn0lDCNkB Design Arena@DesignarenaArticleHow GLM-5.2 Beat Fable 5 at Website DesignGLM 5.2 ranks 1st overall on Design Arena’s single-turn, HTML Web Design (Non-Agentic) evaluation, 5 places higher than its predecessor GLM-5.…
GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2 (arrowtsx.dev via hn) Bigger models are not the way Jun 18, 2026 A shift is happening among major AI labs, who are becoming increasingly skeptical of endless parameter count and training data scaling. The limits of this paradigm were put on the world’s stage wh…
I ran GLM-5.1 on a 16GB RAM machine (github.com via hn) 🧠 MoE-on-a-Potato Running a 754-Billion Parameter LLM on a 16GB RAM Consumer PC "Saying it's impossible is not engineering. Saying we don't know how yet is science." MoE-on-a-Potato is an experimental project dedicated to testing the extre…
Open weights GLM and Mimo are better than Gemini 3.5 flash according to arena (www.reddit.com) While we are weathering the gemini 3.5 flash hype, keep in mind that according to arena, GLM and Mimo are better. https://arena.ai/leaderboard/text/coding-no-style-control #7 GLM #9 Mimo #12 Gemini 3.5 Flash
cdesktop — open-source Claude Code Desktop alternative, runs locally via npx, supports any provider (www.reddit.com) I built cdesktop with Claude Code — it's an open-source alternative to Anthropic's Claude Code Desktop, running locally on your machine via npx cdesktop. Free, Apache 2.0.
Open source battle: GLM vs Kimi vs MiMo vs DeepSeek (www.youtube.com via reddit) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Show HN: Grunden – Frontier AI inference hosted in Sweden, OpenAI-compatible (grunden.ai via hn) grunden.ai är en svensk AI-tjänst för utvecklare, myndigheter och helt vanliga människor. GLM 5.1 (open-weight) med EU-jurisdiktion, ett OpenAI-kompatibelt API och prissättning i kronor.
Ran K2.6 through a third-party coding benchmark: heres how the figures stand up (www.reddit.com) I have been following the akitaonrails coding benchmark which tests against a fixed rails + Rubyllm + docker task rather than vendor-reported evals. April 2026 update put K2.6 at 87 sitting in tier A (80+), ahead of Qwen 3.6 plus (71), Dee…
Just got a beast. (www.reddit.com) 1.5 tb ram with 128gb vram and a 28 core processor. Mac Pro 2019.
Capacity vs Speed trade-off: 1.1TB Mac Unified Memory vs. RTX 6000 Pros (www.reddit.com) I'm usually a Windows person, but I’m currently running a Mac cluster for local LLM orchestration. My setup consists of four 256GB Mac Studios plus one 96GB Mac Studio, giving me about 1.1TB of unified memory.
What's the best GPU cluster/configuration 30k $ can buy? (www.reddit.com) Edit: I’m getting the consensus is that the budget I suggested is not enough for my lil ambitious project. I’d like to reshape the question for the upcoming comments: what’s the minimal budget to achieve my goal?
do GLM-4.7 Flash Q4_K_M have problem with claude or agent? (www.reddit.com) I'm brand new to local LLMs and started with GLM-4.7 Flash q4_K_M. When I run it directly: ollama run glm-4.7-flash:q4_K_M it works pretty decently — nothing amazing, but usable and responsive.
I got better results when I made each AI tool do one job (www.reddit.com) I spent too much time trying to find one AI dev tool that could do everything. Planning, coding, fixing, reviewing, maybe filing my taxes too It never really worked.
What's the current best code autocomplete LLM for local deployment (as of April 2026)? (www.reddit.com) I know this question has already been asked a thousand times, probably, but... what's the best or close-to-best model I can use with Continue for local IDE-like code autocomplete?
GLM Built Its Own Inference Infrastructure (z.ai via hn) could not extract summary
Ask HN: What's the most economical approach to the most tokens? (news.ycombinator.com) I'm doing web developement, and game development for a hobby project. I've tried lots of harnesses / IDE's - Best I've found is VSCodium.+ Cline + Openrouter, using discounted models (GLM 5.3 Flash is 50% off atm for example) I used Cursor…
Harness your expectations: a 27B model matched GLM-5.3-Flash after leak fixes (aistack.imec-int.com via hn) Intro In the last few months our aistack team has been on a quest to get a grip on what it takes to own your own AI stack. We’ve looked into the differences in cost and performance when using APIs, renting or buying GPUs, and started ident…
Is GLM-5.3-Flash Mythos-Level at Cyber? (generality.org via hn) September 2026 · By James Mann Is GLM-5.3-Flash Mythos-level at Cyber? We ran GLM-5.3-Flash on ExploitBench with a budget of 1 billion tokens per vulnerability.
Show HN: Nowdex – AI agent usage on your iPhone (nowdex.app via hn) Hi HN, I've been using Claude Code, Codex and Cursor quite a bit lately, and I found myself checking their usage limits all the time. Most of the tools I found for this live on the desktop or in the menu bar.
Show HN: Cognition-Claude-proxy – Use Devin model catalog with Claude Code (github.com via hn) cognition-claude-proxy A local proxy that lets Claude Code (or any Anthropic-API client) use Devin's model catalog — SWE-2, GLM-5.2, DeepSeek V4.1 Flash, and 200+ others — as its backend. It translates the Anthropic Messages API to the Con…
I hate benchmarks: Moving a production workflow from GPT to GLM-5.3 Flash (polyform.ai via hn) Why I Hate Benchmarks: Moving From GPT to GLM-5.3 Flash A cautionary tale about how a cheap, capable model swap became a production systems migration. I hate benchmarks because they make the model look like the product.
Mistral Vibe Code: GLM 5.2 is now available (twitter.com via hn) Vibe on X: "Something else to play with this week: GLM 5.2 is now available in Vibe Code on the Pro and Team plans, and served by Mistral in Europe. Generous usage limits included.
Show HN: Abliterated GLM-5.3 API (84.5% CyberGym, FP8) (abliteration.ai via hn) Abliteration.ai Unrestricted models. Governed by your policy.
Show HN: Our GLM-5.3 Flash Switchless recipe is now out for 4x DGX Sparks (github.com via hn) GLM-5.3-Flash NVFP4 — 4× DGX Spark, switchless-ring TP4 + DFlash2 Serve GLM-5.3-Flash (NVFP4) across four NVIDIA DGX Spark (GB10 / sm_121) nodes as one tensor-parallel engine — joined by a switchless RoCE ring and accelerated by the DFlash…
Tinker: GLM 5.3 Fine-Tuning (tinker-docs.thinkingmachines.ai via hn) Models & Pricing All prices are per million tokens. Checkpoint storage is charged at $0.10 per GB per month.
GLM-5.3-Flash on Apple Silicon (github.com via hn) WARP — Weight-Aware Runtime and Paging (formerly WASTE) WARP is an embeddable inference engine written in C, with no third-party runtime dependencies. It keeps the model trunk in memory, streams selected experts directly from disk, and use…
Show HN: Warp – Run the 313B GLM-5.3-Flash on a MacBook with 8GB RAM (news.ycombinator.com) A few months ago, I created the WARP engine (formerly WASTE) to run Kimi K3, the complete 2.78-trillion-parameter model, on macOS. GLM-5.3-Flash shares many architectural similarities with Kimi K3, so I added support for it as well.
Getting GLM-5.2 NVFP4 Post-Training off the ground (patronus.ai via hn) Getting GLM-5.2 NVFP4 Post-Training off the ground The goal was deceptively simple to state: take GLM-5.2, a 744B-parameter mixture-of-experts model quantized to 4-bit NVFP4, attach a bf16 LoRA adapter, and train it with reinforcement lear…
Gemini 3.7 Flash, Grok 4.6, GLM-5.3 and DeepSeek V4 Pro joined the frontier (quesma.com via hn) Since these models are smart, I decided to rerun the Baba Is Bench, to see how the models fare on a puzzle game. Even though the game is popular, we check for spoilers - and to our surprise, there are no signs of models knowing solutions a…
Open source, audited by GLM-5.3 (huggingface.co via hn) OpenVuln Public frontend for OpenVuln. This Space is built automatically from the root Dockerfile and serves the Vite application with nginx on port 7860.
Show HN: HN Hiring – Search and Filter Who Is Hiring (hnhiring.azuanz.com via hn) https://hnhiring.azuanz.com I originally built this to scratch my own itch while job hunting. The main problem for me with the monthly Who Is Hiring thread was finding relevant jobs by location.
Show HN: Popkorn – A CSS based alternative for Lottie animations (github.com via hn) Hi all, I’ve been working on this project for a while now and wanted to share it here for feedback and contributions. You can test drive it in the playground at: https://usepopkorn.dev I’ve been calling it Popkorn.
Hugging Face rebuilt a third of its infrastructure after OpenAI agents ran amok (www.theregister.com via hn) MOST POPULAR AI - AI and ML Too many AI agents can get in each other's way For enterprise agents, less is more - AI and ML Impostor Chinese models pretend they're Claude Researchers find GLM and Kimi can adopt Claude's identity, but the ev…
Ask HN: HotPin – lossless 120B MoE inference on 24GB RAM (CPU, 50 loc) (news.ycombinator.com) I'm a mechatronics designer with a background in control systems, robotics, PCB design, and embedded hardware. I design physical systems: motors, sensors, microcontrollers, and real-time control loops.
MiniMax M3: How Sparse Attention Makes Long-Horizon Agents Practical (twitter.com via hn) https://t.co/v9huIornsf elvis@omarsar0ArticleMiniMax M3: How Sparse Attention Makes Long-Horizon Agents Practical GLM 5.2 has taken over much of the AI timeline lately, and most of the conversation has centered on how it stacks up against…
The same LLM is 8x slower to first token depending on who serves it (dynoyard.app via hn) We serve GLM-5.2 to teams building agents. Same open-weight model, same OpenAI-compatible API — but we route it across more than one backend, and while swapping one in we found something worth writing down: the backend you pick changes tim…
Show HN: Self-hosting unpruned GLM-5.2 on a 4-node DGX Spark cluster (github.com via hn) GLM-5.2 (unpruned) on 4× DGX Spark — depth, max context, or multi-user Serve the unpruned GLM-5.2 (QuantTrio Int4-Int8Mix, all 256 experts) across four GB10 Sparks — one recipe, four lanes, one KV budget spent on depth or width**. TP4 + DC…
Show HN: GLM-5.2 is now available via Canopy Wave (twitter.com via hn) Came across this today. Canopy Wave just added GLM-5.2, and they're giving away a few 7-day trial accounts for anyone who wants to test it.
GLM-5.2 (max) matches Claude Opus 4.8 on Harvey LAB-AA benchmark (artificialanalysis.ai via hn) Compare AI model performance on Harvey LAB-AA Benchmark Leaderboard. Artificial Analysis' implementation of Harvey's Legal Agent Benchmark (LAB), testing AI agents on real-world legal work from Harvey's dataset of 120 private tasks spannin…
Run GLM 5.2 on 2 MacBooks with 128gb on RDMA with DS by antirez (twitter.com via hn) Big news for DwarfStar users: I got DeepSeek v4 Flash and GLM 5.2 working with Tensor Parallelism across 2 M5Max 128GB MacBooks via RDMA. It is especially interesting for GLM since otherwise, fully resident, can't fit a machine that money…
I built a tiny proxy that gives GLM 5.2 vision (or any text LLM) – MIT (github.com via hn) VisionBridge Give text-only LLMs vision through a tiny OpenAI-compatible proxy. VisionBridge sits between your chat UI and your models.
GLM 5.2 and the coming AI margin collapse (martinalderson.com via hn) GLM 5.2 and the coming AI margin collapse (part 1) This is a two part series focusing on what I believe is perhaps the least understood upcoming shift in AI economics. If you've enjoyed this and want to be notified about the second post, p…
GLM-5.2: The Open-Source Chinese Model Challenging Claude at One-Fifth the Cost (mrkt30.com via hn) GLM-5.2 from Z.ai is an open-source frontier model that competes with Anthropic’s Fable 5 and Opus 4.8 at roughly one-fifth the cost. With a 1M token context window and strong long-horizon coding performance, it’s changing the economics of…
GLM-5.2's Code Reviews Are Only as Good as Your Prompt (blog.kilo.ai via hn) GLM-5.2’s Code Reviews Are Only as Good as Your Prompt GLM-5.2 from Z.ai has been one of the most talked-about open-weight models since it launched, and we have made it our daily driver to see how it performs on various coding tasks. We al…
China's Z.ai claims it can match Mythos on cybersecurity (www.theverge.com via hn) China’s Zhipu AI (Z.ai) released its open-weight GLM-5.2, and some researchers have claimed that it matches Mythos in certain bug-finding and cybersecurity scenarios. While GLM lags behind models from Anthropic and OpenAI in other, more ge…
Show HN: Cline subscription plan to access GLM-5.2 at 2-5x discount (cline.bot via hn) Hi I'm Saoud, founder of Cline. We’ve been impressed with GLM-5.2 and so are introducing a $9.99/month subscription to give you 2-5x discounted access to it and other open weight models like DeepSeek, Kimi, MiniMax, Mimo, and Qwen.
Show HN: Caliper – pass@k reliability testing for Claude Code and Codex skills (github.com via hn) Skills for Claude Code and Codex are hard to test. What I mean by hard is that there's no standard way to do it.
Was GLM-5.2 trained on Opus 4.5 outputs? (1chat.com via hn) Recently there is a lot of excitement about GLM-5.2 which is an open-weight MoE LLM performing on Claude Opus 4.5 level in chat arena and overperforming all models except Claude Fable in WebDev arena [1]. Even though it is very good that t…
GLM-5.2 (Max) API Provider Benchmarking and Analysis (artificialanalysis.ai via hn) Analysis of API providers for GLM-5.2 (max) across performance metrics including latency (time to first token), output speed (output tokens per second), price and others. API providers benchmarked include Together AI, FriendliAI, Fireworks…
GLM-5.2, not Mythos, is the real security emergency (joshuasaxe181906.substack.com via hn) Until last week, attackers faced a dilemma in using frontier models: even if they could manage the cat-and-mouse game of setting up fake accounts to retain API access to frontier model providers, and even if they could induce models to hel…
Running GLM-5.2 on a 64GB Mac, barely (andreaborio.substack.com via hn) I tried to run GLM-5.2 on a 64GB Mac Field notes from an experimental ds4 fork, a 244GB GGUF, and the small horror of sparse models that are sparse in compute but not very friendly to filesystems. I have a weakness for local LLM experiment…
Openresearch: GLM 5.2 for Autoresearch (openresearch.sh via hn) could not extract summary
Genuinely impressed, almost shocked, at how good GLM-5.2 (twitter.com via hn) Genuinely impressed, almost shocked, at how good GLM-5.2 by @zai_org is at coding. This changes things.
Show HN: AdvertBench, ranking the ability of LLMs to create image ads (advertbench.com via hn) Experiment that I've made. The models get access to an E2B sandbox and are instructed to create an ad according to the specifications (they can choose whatever tools they want to use for it, e.g.
I evaluated GLM 5.2 against the frontier on tasks from real repos (www.stet.sh via hn) GLM 5.2 vs Composer 2.5 and the premium field on 50 real merged PRs from graphql-go-tools (Go) and sqlparser-rs (Rust). GLM lands last on craft and equivalence in both repos, costs about twice Composer, and writes more code than the human…
GLM 5.2 ranks #2 in Code Arena: Frontend (twitter.com via hn) Exciting news: GLM-5.2 (Max) ranks #2 in Code Arena: Frontend, with +29pt over Claude Opus 4.7 (Thinking) and only behind Fable 5! GLM-5.2 is the best open model vs Kimi-K2.6 and Minimax-M3 by a large margin.
GLM 5.2 Performance Benchmarks (artificialanalysis.ai via hn) GLM-5.2 (max) Intelligence, Performance & Price Analysis Model summary IntelligenceUpdated Speed Price Cache Hit Price Verbosity GLM-5.2 (max) is amongst the leading models in intelligence, but particularly expensive when comparing to othe…
GLM-5.2 is now available with 1M-context support (twitter.com via hn) Intelligence should be open, accessible, and ready to build with, empowering every developer, everywhere. GLM-5.2 is now available to all GLM Coding Plan users, including Lite, Pro, Max, and Team plans.
Show HN: Free open source coding models in Slack (www.runcord.com via hn) Hey HN, We believe we have the easiest onboarding from signup to being able to spin up coding agents in slack like Stripe, Ramp & Coinbase. Demo of the onboarding: https://www.tella.tv/video/connecting-cord-to-slack-1-19ep Every signup get…
↯ Glm↯ Minimax↯ Gemma↯ DeepSeek 4↯ DeepSeek 4minimaxglmgemma+5
Show HN: Chuddy, self-hosted media downloading, translation and OCR Telegram bot (github.com via hn) My latest project, about 60% of the codebase was written with Z.ai's GLM-5.1 model. It's basically a Telegram bot that allows for embedding/downloading media easier within group chats.
What’s going on with GLM? Are they scamming or what? (www.reddit.com) I have a GLM subscription that’s marketed as offering 3× higher usage than Claude Pro. I primarily use it through Claude Code CLI as a backup coding model.
Chinese AI Coding Plan (www.reddit.com) With the lowering usage limit in Claude, I am thinking of jumping ship to Chinese AI, since the benchmark is already very near compared to Sonnet or Haiku 4.5 , but for a fraction of the price. I am not worried about where is my data endin…
tested four newest open source Kimi K2.6 is the fastest, GLM 5.1 the fanciest, DeepSeek V4 is the most comprehensive, and Xiaomi MiMo is the slowest (www.reddit.com) Architecture explains the gap: MiMo's MoE runs more active params per token than Kimi K2.6's optimized routing hence slowest. DeepSeek V4's 'comprehensive' edge is partly MLA: ~75% KV-cache compression makes it far better for long agentic…
Why is no open weight model inference provider hosting Mimo-v2.5 or Mimo-v2.5-pro? (www.reddit.com) Literally no 3rd party api inference provider is hosting the mimo-2.5 series models from Xiaomi. They seem to be reallly good.
Who else thinks AI is reaching a plateau (www.reddit.com) I must say that I almost feel no difference in all of the latest models that are coming out. Opus 4.7 is almost equal to 4.6 and 4.5, same about the other GPT models, the Kimi K models and the GLM models they all I feel they’re almost all…
GLM-5.1 on Mi50? (www.reddit.com) Hi, did anyone with an AMD MI50 setup (8x 32GB) test GLM-5 or GLM-5.1? Currently, I have 3x AMD MI50 and I was wondering if it's worth buying another 5 of them and a new PSU.
Ask HN: Are there any good open-source chat apps? (news.ycombinator.com) Hi HN family! I've recently been messing around with open models through ollama (glm-5.1 and kimi-k2.6), and I've been impressed with just how close they are to Claude Sonnet for my needs, especially programming.
3 of TIME's top 10 AI companies are Chinese and I only knew one by name (www.reddit.com) I code for a living, close to 7 years now, and I read way too much tech news. TIME dropped their 2026 most influential AI companies list and going through it I see OpenAI, Anthropic, Google, Meta, Amazon, then Zhipu AI sitting right there…
Open Source Company Coding Plans (www.reddit.com) I’ve been looking to buy a coding plan from one of the major open source contributors to give my meager support to them and transition away from Claude. I would love to hear some feedback from the community of their experience with some of…
I'm Not a Dev But I Use Qwen 3.6 35b to Code (www.reddit.com) Full disclosure: I used to program a bit, but I was garbage at it so I found a new career. This was eons ago so I'm not a dev, obviously.
Cursor 3 eating GLM 5.1 usage (www.reddit.com) Hello all just as it sounds. I recently started using GLM 5.1 in cursor 3 but unlike in the past, GLM 5.1 ran through my entire daily budget from summarizing chat context and running commands.
Mouse on frontier harness with glm-5.3-flash (mouse.dev via hn) Mouse on GLM-5.3-Flash 23 of 30 FrontierHarness tasks on Z.ai's Flash model, next to the published GLM-5.3 harness runs, for $6.72 in tokens. Community results shared this week ran GLM-5.3 and GLM-5.3-Flash from Z.ai through five coding ag…
Using Modal for GLM Flash on Cadenya's Agent Runtime (www.cadenya.com via hn) Get to know Cadenya We’re developers who love to build. We set out to create a yes-code platform that makes building agents feel like the best parts of building software.
Just testing my memory safe web browser (news.ycombinator.com) WebKit MiniBrowser compiled with Fil-C on top of Linux userland compiled with Fil-C. GTK4, Weston, etc - all compiled with Fil-C.
A $37 GLM 5.3 red team: the Alloy-modeled auth layer held, but two bugs outside (goodmem.ai via hn) Red-teaming GoodMem with GLM 5.3 We used GLM 5.3 to red-team GoodMem. How we defined the tests, what the agent found, what we fixed, and how we verified the fixes.
Chess5.ai – Play Chess, Go, Xiangqi, Gomoku, and Othello Against LLMs (chess5.ai via hn) chess5.ai Human vs LLM · Five Games Pit yourself against GPT, Claude, Gemini, Grok, Muse Spark, Mistral, DeepSeek, Kimi, Qwen, GLM, or MiniMax across five classic boards. How it works - Human vs model, or model vs model — with spectating a…
GLM-5.2 RL weight transfer in 4 seconds using NIXL and ModelExpress (www.primeintellect.ai via hn) GLM-5.2 RL weight transfer in 4 seconds using NIXL and ModelExpress GLM-5.2 RL weight transfer in 4 seconds using NIXL and ModelExpress In RL at 1T Scale, we detailed how prime-rl trains trillion-parameter models like GLM-5 with sub-5-minu…
Desktop AI for ops is still subsidized (news.ycombinator.com) I'd like to share an opinion on the current situation in AI ops, not the coding/tech side. Right now the market has a pretty clear pattern.
How viable is LLM-assisted 3D modeling in Blender? (twitter.com via hn) GLM-5.3-Flash is an excellent model. Running it locally on 4 x RTX 6K Pros with great results.
Run any LLM in t3's Codex and Claude tabs through a local gateway (github.com via hn) proxy-llms Run any model in t3's Codex and Claude tabs. A GPT model in the Claude tab, GLM or Kimi K3 in the Codex tab, plugged and unplugged on demand.
Mushroom hunting with LLMs: what can go wrong? (quesma.com via hn) Mushroom identification with AI: GPT-5.6-Sol, Gemini 3.7 Flash, GLM-5.3-Flash and Claude Fable 5.1 benchmarked on poisonous and edible species of FungiTastic. A lot of dangerous errors.
GLM-5.3-Flash at 1000 tok/s on RTX PRO 6000 (www.localmaxxing.com via hn) LocalMaxxing Get started Models Reports Hardware Benchmarks More + Submit Get started Leaderboard Decode calculator Models Reports Hardware Benchmarks Marketplace Rentals Pro API Docs Language English 简体中文 繁體中文 日本語 한국어 Español Français Deu…
Abliterated model large v2: GLM 5.3 84.5% CyberGym (abliteration.ai via hn) Today we are releasing abliterated-model-large-v2. We started from GLM 5.3 and abliterated it for offensive cyber, AI red teaming, and agent testing.
Show HN: 2x-4x cheaper GLM 5.3 for coding and research (www.coralbricks.ai via hn) Inference for Kimi and GLM. Built for research and coding agents that plan, call tools for hours, and reason over large contexts.
What GLM-5.3 Flash running on Chinese hardware means (martinalderson.com via hn) What GLM-5.3 Flash running on Chinese hardware actually means Z.AI confirmed that their most recent model release was running all inference on Chinese manufactured hardware. While no doubt an impressive feat, Western companies still have a…
GLM-5.3-Flash at 3.3 tok/s (news.ycombinator.com) A few months ago, I created the WARP engine (formerly WASTE) to run Kimi K3, the complete 2.78-trillion-parameter model, on macOS. GLM-5.3-Flash shares many architectural similarities with Kimi K3, so I added support for it as well.
GLM-5.3 achieves 60 on Artificial Analysis (twitter.com via hn) GLM-5.3 achieves 60 on the Artificial Analysis Intelligence Index, on par with Kimi K3 and up 7 points from GLM-5.2. Once the weights are released it will be tied as the leading open weights model @Zai_org has just launched GLM-5.3, whic…
GLM 5.3 Available in OpenRouter (openrouter.ai via hn) GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves on GLM-5.2 in coding and in the balance…
Reasoning prefills on a few open models (gist.github.com via hn) I wrapped this small followup to stolen-thoughts.com in a gist for easier reading and wanted to share it here. My hunch is that this might not just be reasoning distillation; it could even be benchmark distillation.
Preparing GLM-5.3 for Open Release: A Responsible Path to Cyber Defense (twitter.com via hn) When GLM-5.2 helped Hugging Face investigate an incident in which an AI autonomously bypassed its own safeguards, it highlighted a broader shift. AI is becoming part of both cyber offense and cyber defense.
Ask HN: Should AI's tell you they're AI? (news.ycombinator.com) Should AI's be required to answer the direct question "are you an AI" with a clear yes? Currently they are not universally required to do so.
Show HN: Memcode launches a new terminal coding agent (www.memcode.ai via hn) Memcode is a new developer platform currently in public beta. It includes a coding agent, chat, reusable agents, DataHub, and a Lovable-style website generator.
Engy – Verified LLM Inference (engy.ai via hn) Verified inference means cryptographic proof that the exact open model you requested produced your output, not a cheaper or quantized stand-in. Run frontier open models like GLM-5.2, billed by the token.
We built the new fastest API for GLM-5.2 (www.baseten.co via hn) A month ago, GLM-5.2 was released. As part of our day-zero support, we built the fastest API in the world for GLM-5.2, with peak speeds of 280 tokens per second and average speeds around 100 tokens per second.
Sakana: Fugu Ultra vs. GLM 5.2 (runtimewire.com via hn) This wasn’t a photo finish. Sakana: Fugu Ultra controlled the matchup on practical writing and coding tasks, while GLM 5.2 showed flashes of polish in a couple of narrower instruction-following spots.
Show HN: Abliterated GLM 5.2 model, fine-tuned for security and red-team work (abliteration.ai via hn) OpenAI-compatible unrestricted AI and uncensored LLM API for enterprise teams: AI red teaming, cybersecurity, trust and safety, synthetic data/evals, ML research, and defense/government contractor workflows. Policy Gateway adds policy-as-c…
Show HN: Run GLM-4.5-Air(110B)on a 16GBRAM consumer machine (github.com via hn) quantprobe Placement beats budget Where your bits sit — which layers, which memory tier — matters more than how many you have. Four falsification-tested laws for running big LLMs on hardware you already own, every number measured on one 20…
Hugging Face turns to GLM 5.2 to fend off AI agent attack (venturebeat.com via hn) could not extract summary
Head to head: GLM 5.2 vs. OpenAI: GPT-5.6 Sol (runtimewire.com via hn) This matchup wasn’t close. GPT-5.6 Sol dominated the practical details that decide real-world usefulness: tighter instruction-following, cleaner formatting, and fewer correctness slips.
MSE-GLM – A Deterministic Zero-Weight Graph Language Model (github.com via hn) MSE-GLM — Command Reference Matrix-Structured Edge — Graph Language Model. Deterministic, zero-weight, explainable.
Show HN: Same castle prompt, 8 LLMs, 24 procedural Three.js worlds (castle-bakeoff.pages.dev via hn) Fable 5 · GPT 5.6 Sol · Kimi K3 · Grok 4.5 · Gemini 3.5 Flash · MiMo V2.5 Pro · MiniMax M3 · GLM 5.2 — low-poly, semi-realistic, very realistic.
Harry Partridge on X: "GLM 5.2 With Vision" / X (twitter.com via hn) https://t.co/iJsDrlGy45 Harry Partridge@part_harry_ArticleGLM 5.2 With VisionGLM 5.2 is one of the best currently available open source language models. However, unlike other flagship models like Qwen, Kimi and Minimax, GLM 5.2 does not su…
Baba Is Solved by Fable 5 and GPT-5.6 Sol, but at what cost? (quesma.com via hn) We ported the puzzle game Baba Is You to the Harbor framework, and benchmarked current models, including Claude, GPT, Gemini, GLM and DeepSeek. A human Twitcher is 4x faster than Fable 5.
Coding with GLM 5.2 on OpenCode for a flat $20/month (meshintelligence.substack.com via hn) How to Code with GLM 5.2 on OpenCode Coding with GLM 5.2 on OpenCode is a bet on your own engineering. Decide the architecture and interfaces first, let the cheap model write the code on a fixed twenty-dollar Ollama plan, and bring a str I…
Fable 5 on Playcode. As well as Sol, Grok 4.5 and GLM 5.2 (playcode.io via hn) Claude Fable 5 is Anthropic's newest flagship and, in our testing and the independent benchmarks, the strongest model in the current lineup for polished front-end work. It is selectable in Playcode's AI website builder.
The Great Wave Has Arrived (Memo from GLM CEO Jie Tang) (twitter.com via hn) https://t.co/3i0qSTbjql Bing Xu@bingxu_ArticleThe Great Wave Has Arrived (from GLM CEO Jie Tang)-- Bing Xu's Note --- I came across an internal GLM letter on the Chinese app RedNote, purportedly written by @jietang, and translated the Chin…
Getting GLM 5.2 running on my slow computer (news.ycombinator.com) A few days ago I found myself trying out GLM 5.2 and was really positively impressed. The capabilities and security I was getting from this LLM are similar to those I've gotten from models like Claude or GPT, and this really surprised me.
AI Agent using a Burp-style toolkit over MCP (github.com via hn) mulot [-4285F4?logo=googlechrome&logoColor=white)]() Agentic AI web pentester that drives a browser. An open-weights LLM (GLM-5.2, Gemma or Qwen) drives a real headless Chromium through a Burp-style toolkit and works a target the way a hum…
Balancing Claude 4.8 with GLM 5.2 in mid-level Coding Agent structuring prompts (cimons.com via hn) Consider modest to major changes in a code base as particularly important to design properly in advance towards having a good intuition of the libraries and implementation details with the end result…
I like Claude Desktop, so I created my own (www.zandrey.com via hn) I like Claude Desktop, so I created my own I like the Claude Desktop application and use it a lot, both for my work and in a personal capacity. I also love to build stuff, and with the recent release of GLM-5.2 I decided to see if I could…
Comparing GLM 5.2 and Opus 4.8 implementing the same methods for the same repos (gist.github.com via hn) A controlled comparison across 19 paired runs spanning 19 repository forks — 38 individual workflow executions total — running an identical paper-implementation pipeline (remyxai/outrider — Claude Code under the hood, with glm-5.2 routed a…
Ask HN: GLM-5.2 FP8 vs. BF16 (news.ycombinator.com) Many cloud providers offering GLM-5.2 seems to only offer FP8. I wonder if anyone has evaluated the difference between FP8 and BF16 in terms of quality?
Show HN: SQL MCP Server – 61.37% on DataAgentBench with GLM-5.2 (github.com via hn) We just posted results to the DataAgentBench leaderboard scoring 61.37% with GLM 5.2. Please check it out and do share your feedback
Show HN: Subconscious and GLM-5.2 Makes "/compact" Obsolete (www.subconscious.dev via hn) GLM-5.2 is a turning point for coding agents. It's the first model a business would actually pay to replace Claude Opus with.
GLM-5.2: Another open-source Chinese AI model has Silicon Valley's attention (www.businessinsider.com via hn) A new AI model from China is generating the kind of buzz not seen since DeepSeek's R1 announced China as a serious threat to American chatbot hegemony over a year ago. Silicon Valley's online echo chamber has been alight with intrigue in r…
Show HN: Cc-fleet – run other LLMs as Claude Code workers, your sub drives (github.com via hn) 🚢 cc-fleet 🤖 Plug any third-party model into Claude Code's ⚙️ Dynamic Workflows, 👥 Agent Teams, and ⚡ Subagents — from DeepSeek · GLM · Kimi · Qwen … to your Codex subscription, with your main session's auth untouched; no Claude subscripti…
MiniMax M3 vs. GLM 5.2: Codegen comparison across autonomous coding tasks (thinkwright.ai via hn) Thinkbench, our custom evaluation harness, was used to drive both models through the same autonomous coding loop: read files, write files, run shell commands, and stop when the task was complete. The scored suite covered greenfield builds,…
GLM-5.2 – How to Run Locally (unsloth.ai via hn) For the complete documentation index, see llms.txt. This page is also available as Markdown.
Running GLM-5.2 5x faster at 500tps with limitation (abhishek.it via hn) Running GLM-5.2 5× faster than vLLM, on a runtime that doesn't support it I rented an 8×B200 and tried to run GLM-5.2 on TileRT, the runtime MiMo used to push a 1T model past 1000 tok/s. TileRT doesn't support GLM-5.2, so I reverse-enginee…
GLM-5.2: Benchmarks, Architecture and How to Run It (www.techaffiliate.in via hn) GLM-5.2 Review (2026): Benchmarks, Free Access & How to Use It Aditya Kachhawa If you've been keeping up with AI news lately, you've probably noticed a new name showing up everywhere GLM-5.2. And there's a good reason for that.
GLM 5.2 is now available via a unified Model API (www.hpc-ai.com via hn) Model APIs Instant Access for Frontier Open-Source AI Models Build AI Apps and Agents with High-Performance Model APIs — No Deployment Required. Start Free TrialEverything You Need to Run AI Models Z.ai: GLM 5.1 4 supported capabilities fo…
An open-source AI just beat OpenAI's GPT-5.5 at coding (1/6th the price) (docs.z.ai via hn) Overview GLM-5.2 is a flagship model built for the era of long-horizon tasks. With truly usable 1M-token context, it has been tested to handle project-scale engineering context, delivering more stable long-task execution, more reliable adh…
GLM 5.2 playing text adventures (entropicthoughts.com via hn) GLM 5.2 playing text adventures I’ve heard some buzz around the new glm 5.2 open-weights model. They say it’s very capable!
Model Card: unsloth/GLM-5.2-GGUF (huggingface.co via hn) GLM-5.2 👋 Join our WeChat or Discord community. 📖 Check out the GLM-5.2 blog and GLM-5 Technical report.
GLM-5.2 Beats Fable 5 on Reasoning – 24 Hours After the U.S. Export Ban (explainx.ai via hn) GLM-5.2 by Zhipu AI tops BridgeBench reasoning 24 hours after the U.S. banned Fable 5.
Show HN: LimitPing – Keep Claude Code and Codex rate-limit windows continuous (github.com via hn) CCLimitPing (limitping) English | 中文 Keep your Claude Code, Codex, and GLM (Zhipu / Z.ai Coding Plan) rate-limit windows back-to-back. These providers bill on a 5-hour rolling window (plus a weekly cap), and the 5h window starts on your fi…
Noob here, curious about roughly how advanced of a video game a model like Qwen3.6 27b could create, if kept fully offline, and got unlimited attempts/revisions (maybe ~1 month project time limit). Like, could it make something equivalent to Pokemon Red? Doom? Doom II? What if using GLM 5.1? (www.reddit.com) So, I got interested in local LLMs a few months ago, but, I don't have a background in coding, and I don't know how to code, and I am not good with computers or anything. So far I mainly just was having fun with comparing different local L…
Is Composer 2.5 better than Glm 5.1 and DeepSeek v4 pro in real world tasks? (www.reddit.com) I am new to Cursor and still testing the free version. Benchmark for Composer 2.5 indicates it is better than DeepSeek v4 and Glm 5.1.
When configuring a third-party AI large model on the MacBook Claude Code desktop client, an error message appears. How can this be resolved? (www.reddit.com) This is my GLM-4.6 model API configuration, and this error is really confusing me. I'm not sure which step went wrong.
Reliable Open Source LLM as a Service (www.reddit.com) Has anyone figured out a provider whose open source models (Kimi, Qwen, GLM e.t.c) can be used reliably in production. I have tested some well known providers and they all suffer from high latency and poor uptime rendering them mostly usel…
Multi-LLM AI trading agent harness (github.com via hn) 1rok 1rok is a standalone harness for running portfolio-construction agents across OpenAI, Anthropic, Gemini, xAI, DeepSeek, GLM, and OpenRouter against the same financial tool surface. Agents query Alpaca, Yahoo Finance, FRED, and Tavily…
Vertex MaaS GLM-5 prompt cache telemetry seems inconsistent. Anyone else seeing this? (www.reddit.com) I'm testing prompt-cache behavior for GLM models on Vertex AI MaaS and I'm seeing inconsistent telemetry. I reproduced it with a synthetic long prompt and repeated identical requests.
Which Chinese Model is best for planning and which is best for implementation? I'm currently using Opencode with an Openrouter API Key, mostly wanna decide between Kimi, GLM, DeepSeek, Qwen, Minimax and Mimo (www.reddit.com) Original plan was to use Kimi/GLM for planning and DeepSeek for implementation, but seeing a lot of love for MiMo and Minimax lately. Anyone running a planner + coder split on Opencode?
Which model has less restrictions now? (www.reddit.com) GPT and Opus block on certain requests. This didnt use to be the case 2 months ago and I made signficant progress with Opus and then one day I had a 2 week break and then a single prompt to continue the work resulted in refusal.
Group Buys for Shared Compute or Model Hosting? Is this a thing? (www.reddit.com) I've been using GLM 5.1 a lot lately, and I love this model. However I don't love sending all my requests to China.
I plan to use a chinese AI model through API for coding through a harness, I'm a uni student so nothing prod related for now. should i go deepseek, minimax, kimi or glm? kinda confused (www.reddit.com) Just cancelled my claude subscription due to poor rate limits, gemini cli doesn't really excel in coding from my personal experience, and my local hardware isn't that powerful to run local AI models, and while codex is good, I wanna try so…
PP speed on dual RTX 6000 12c EPYC setup (www.reddit.com) I want to run big models like GLM 5.1 or Kimi k2.6. I can buy Mac Studio M3 Ultra with 512gb ram, but PP speed would be ofc bad.
Local LLM Benchmark about Backend Generation by Function Calling (GLM vs Qwen vs DeepSeek) (www.reddit.com) Detailed Article: https://autobe.dev/articles/local-llm-benchmark-about-backend-generation.html Five months ago I posted the "Hardcore function calling benchmark in backend coding agent" thread here. As I wrote in that post, it was an unco…
↯ Glm↯ Function Calling↯ Sonnet 4.6function-callingglmgpt-5+3
Built a self-hosted agent for small businesses that writes its own skills. ~$0.15 per customer booking on GLM-5.1 (www.reddit.com) Been working on this for a while and finally at a point where it's running in production for a couple of small businesses, so figured I'd share. The thing that kept bugging me about "AI employee" products is that none of them are something…
Received a message from Z.AI about occasional garbled outputs and unexpected behavior (www.reddit.com) I received this mail: "Hi developers, Some of you flagged occasional garbled outputs and unexpected behavior when building with the GLM-5 series, especially under heavy workloads. We heard you, reproduced the issues, and the fixes are now…
Comparing SVG Generation for the top open models (codeinput.com via reddit) Some of the larger models (like Llama) weren't available on OpenRouter, so I had to work with what was there. Best small model: Gemma 4 26B For its size, I think it had the best output.
Best value in the 20$ range coding agents? I want the best quality and high-usage-limit I can get at that price. (www.reddit.com) I'm a compsci student and I've been using the 10$ copilot plan for about 2 years now, and it was fine for me since I did a good model distribution taking into account the complexity of the task, I was able to get through the month always u…
anyone actually tried deepseek v4 pro for coding? (www.reddit.com) so v4 pro dropped and barely anyone is talking about it. feels weird since when kimi k2.6 came out i seen post about it everywhere anyone here tried v4 pro for actual code work?
Qwen 3.5 397b and GLM 5.1 Opus fine tune (www.reddit.com) Hi all. Many models on hugging face have been fine tuned with that 3000x opus dataset, but the two I mentioned in the title are missing it.
Best app to use Nvidia Nim? (www.reddit.com) Show HN: RepoGauge – save token costs and compare agents on your own repos (repogauge.org via hn) I've grown increasingly skeptical that public coding benchmarks tell me much about which model is actually worth paying for and worried that as demand continues to spike model providers will silently drop performance. I did a few manual an…
Minimax vs Qwen vs Kimi vs Mimo(Omni) vs Glm ( via reddit) could not extract summary
Upgrade paths for my 256g ddr4 ram + 4x24g vram system (www.reddit.com) So I was just about to give up playing with local models, until I realised I can actually run GLM 5.1 at not too horrible speeds, using this quant https://huggingface.co/ubergarm/GLM-5.1-GGUF/tree/main/IQ2_KL in ik llama. Getting around 6.…
Which AI model is best for real data analysis? [benchmark] (www.reddit.com) I created and run a benchmark for AI models in data analysis tasks. In contrary to other benchmarks, it is not one-prompt benchmark, but I tried to simulate the real work of data analyst.
Model API Performance (news.ycombinator.com) We’ve been benchmarking a few models on our API platform and got some interesting performance numbers: - MiniMax M2.5 → 0.118s time-to-first-token, 103 tokens/sec - GLM 5.1 → 120 tokens/sec throughput - Kimi K2.5 → 0.643s TTFT, 69 tokens/s…
What Am I Doing Wrong? Models Won't Listen, At All (GLM 5.1, MiniMax M2.7, Kimi K2.5) (www.reddit.com) What am I doing wrong here? I can't get models to follow my instructions, pretty much at all.
Claude should offer first-party support for open-weight models Anthropic hosts via partnerships, hear me out. (www.reddit.com via reddit) Imagine how badly distillers and Chinese hosts undercutting Claude would shit their pants if Anthropic offered high throughput first-party support via Cerebras or another wafer-based host with DeepSeek 4.1f, GLM 5.3,.etc. so everyone looki…
Validating Hybrid-State Cache Recovery for GLM-5.3-Flash with vLLM and LMCache (arxiv.org) External cache transfers can succeed while a hybrid language model resumes from an inconsistent state. We examine the full 45-layer GLM-5.3-Flash model, using the RedHatAI/ GLM-5.3-Flash-NVFP4 quantized checkpoint with vLLM and LMCache und…
[Investigation] The "Unlimited Compute" Scam: Wire-Level Proof of Model Spoofing, Dangerous Setup Scripts, and Packet Analysis of CodexAPI.pro (www.reddit.com via reddit) I bought credits on codexapi.pro (https://codexapi.pro/) after seeing their promos for cheap "unlimited" coding sessions with Claude Code and Codex CLI. In practice, the service was constantly dropping connections: 502 Bad Gateway errors s…
working with Chinese open weight has interesting side effects (www.reddit.comhttps) I just started using GLM 5.3 Flash with Claude Code; I'm using GSD framework and one of the sub-agents spawned was reporting progress as normal. 正在清理 03.3.1-02-PLAN.md 中的 files_note 元素 translates to Cleaning up the files_note element in 03…
GLM 5.3 Flash vs Kimi K3 for heavy coding — which subscription would you choose? (www.reddit.com via reddit) I'm planning to use AI seriously for coding, roughly 80% GLM 5.3 Flash and 20% Kimi K3 for harder tasks. I mainly care about large projects, debugging, refactoring, agentic coding and value for money.
You love using glm or calude at cursor (www.reddit.com via reddit) Glm and claude are coding bosses now , wich one is better
What AI subscription should I switch to? (www.reddit.com via reddit) Big Claude user, but Anthropic got stingy as hell with the limits. I used to barely touch my weekly allowance; now I can burn through 20% in a day and I'm cooked in ~2 days.
The best subagent for Claude? (www.reddit.comhttps) I'm horrified by my Fable and Astra token spend (I'm subscribed to $200 plans for each one). Therefore, the question I asked myself was whether Sonnet still holds up as a good sub-agent?
Help me undestand, Claude memory make the real difference ? (www.reddit.com via reddit) I am working on a big project with Claude Code only context7 mcp added no others tools. With opus 5 is all ok it seems to remember what we have done days before follow the repo conventions etc.
↯ Glm↯ GLM 5.3↯ GLM 5.3↯ GLM 5.3↯ GLM 5.3↯ GLM 5.3glmmcpopus+1
A smaller menu bar app for Claude Code, Codex and GLM limits (www.reddit.com via reddit) I'd been building this for myself when Peter released CodexBar back in November. CodexBar does far more, 69 providers and a bundled CLI, and it's genuinely good, 21k stars and 119 releases since.
What is your opinion (www.reddit.com via reddit) What is your opinion about cursor ? I mean is it better than using the official coding applications for the agents (liek using glm 5.3 at zcode or cursor, what is the difference?)
when are we going to have glm 5.3 in Cursor plan? (www.reddit.com via reddit) When are we going to have available glm 5.3?
Claude gets surprisingly hostile when I try asking about other models (www.reddit.com via reddit) I am trying to set up an email triage setup for my needs. After a bit of discovery I started exploring using Hermes Agent as an option with Claude.
All leaks and news about Fable, Opus and sometimes Sonnet, what about Haiku? Do you use it? what is your use case? (www.reddit.comhttps) I literally use it as the meme stats, Anthropic may lost that low cost tier war with models like GLM 5.3 flash and GPT Luna I can't think they can compete in terms of price/performance in this tier
Anyone using third party providers with cursor? (www.reddit.comhttps) I'm not a big fan of vendor-lock-in. Anyone tried a subscription based provider with cursor?
M5 Ultra external SSD options? (www.reddit.com via reddit) Planning to get M5 Ultra 512GB to run GLM-5.3-mlx-mxfp4. However, I think the Apple SSD is a rip off.
Agent Arena Code - Very good result (preliminary) for GLM and Qwen! (www.reddit.comhttps) Everything has changed in two months: DS4 0731 flash was the start of a wave that is taking open weights to paradise. It is easy to think that Qwen 4 and GLM 6 will be on par with Mythos.
↯ Glm↯ Anthropic Mythos↯ Qwen 4↯ Qwen 4↯ Qwen 4↯ Qwen 4↯ Qwen 4↯ Qwen 4↯ Qwen 4↯ Qwen 4glmmythosqwen
Are there any providers who'll be hosting uncensored versions of GLM 5.3? (www.reddit.com via reddit) I recently got into the habit of using qwen uncensored models for just local reverse engineering workflows, some the flagship cloud models even glm models refuse. But theres only so much intelligence i can pack into 16gb vram.
GLM-5.3 Flash Unsloth GGUF now available (huggingface.co via reddit) Read our How to Run GLM-5.3-Flash Guide! Unsloth Dynamic 3.0 achieves superior accuracy & outperforms other leading quants.
GLM-5.3-Flash (FP8) on 4 x RTX6000 Pro (www.reddit.com via reddit) I've forked https://github.com/tonyd2wild/GLM-5.3-Flash-NVFP4-2x-DGX-Spark and make it run on sm120. I'm using it right now - got 1,4M context (5,45 sessions 262k each) 3,7kt/s PP and 160 - 230t/s TG (MTP enabled) You can make vllm Docker…
GLM-5.3-Flash @ DGX Station GB300: ~206 tok/s (single stream), 1M context (www.reddit.com via reddit) Hey all! I'm finally doing some cool stuff with my "thinking heater" (h/t u/-TV-Stand-).
How many models in v3 rn? (www.reddit.com via reddit) how many models are going by 3 rn like Gemini 3.7, qwen 3.8, minimax m3, kimi k3, hy3, deeseek v4 flash, glm 5.3. (not deepseek and glm but close enough)
DeepSeek V4 0731 -> Qwen 3.8 Flash -> GLM 5.3 Flash (and back again!) (www.reddit.com via reddit) Spent yesterday getting Qwen3.8 Flash and GLM 5.3 Flash up and running on my cluster of 4 x DGX Sparks with a view to replacing DeepSeek 0731... but..
GLM-5.3 weights will be released tomorrow (huggingface.co via reddit) The promise has been fulfilled.
When three models all claim SOTA, how do I pick for a local agent stack (www.reddit.com via reddit) I personally stopped reading the launch table once GLM-5, MiniMax M2.5, and Gemini 3 Deep Think dropped in two days and all claimed the same coding, reasoning, and agent wins. They optimize different constraints.
Currently GLM 5.3 Flash matches with Sol 5.6 (Max) in Agentic Index (Artificial Analysis) (www.reddit.comhttps) could not extract summary
Who else is drop watching? (www.reddit.com via reddit) curl -s https://api.github.com/repos/ggml-org/llama.cpp/pulls/27742 | jq -c '{draft,state,merged}' for r in unsloth/GLM-5.3-Flash-GGUF unsloth/Qwen3.8-Flash-Next-GGUF; do echo "== $r" curl -s "https://huggingface.co/api/models/$r" \ | jq -…
[Megathread] GLM-5.3-Flash - former ox-alpha (www.reddit.com via reddit) Megathread for discussing the release of GLM-5.3-Flash. Quants Fine-Tunes & Abliterations Chat Templates Inference Server Support & Configuration Experiences, Benchmarks & Model Comparisons We'll try to clean up future duplicates around th…
Z.AI confirms Ox Alpha is a GLM model, plans to release its weights (runtimewire.com via reddit) Z.AI confirms Ox Alpha is a GLM model, plans to release its weights The Beijing lab told Bloomberg it built the anonymous coding model and planned to release its weights on August 26. By RuntimeWire Staff · Published Primary source: Bloomb…
First serious confirmation. Ox Alpha is GLM-5.3-Flash (www.reddit.com via reddit) https://x.com/romanchernin/status/2092488160680751437?s=20 - Multimodal (Vision) - 1M Tokens Context Window - DeepSWE ~63%
Hallucinations and severe scope drifts inspite of planning. (www.reddit.com via reddit) I am using 20x plan. I am seeing a lot of hallucination and scope drift, and extremely slow execution.
Where are the Kimi K3 and GLM Distillations? (www.reddit.com via reddit) Anthropic is always whining about Chinese competitors distilling their models, but current Chinese models have basically caught up to frontier capabilities and released them for free. Since the weights are open and available, shouldn’t it…
Thoughts on what Ox Alpha could be? (www.reddit.com via reddit) Shares the same tokenizer as GLM, seems to perform as well as larger models, and has vision support. Could this be Baseten's GLM Vision (maybe further finetuned) or an official GLM model?
4xR9700, 2xMi210 or 4x4080S 32G (www.reddit.com via reddit) I am trying to get to 128G of VRAM with reasonable compute and bandwidth to run multiple models in parallel. DS4 Flash or GLM 5.3 in hybrid mode with custom checkpoints.
Is there a way to run a local Claude Desktop-type setup? (www.reddit.com via reddit) I used to use Qwen3.6-35B-A3B with llama.cpp and connecting it to the VSCodium extension called "Continue." My computer is running a Intel(R) Core(TM) Ultra 7 265K (3.90 GHz) with 128 GB of DDR5 RAM and an Nvidia Geforce RTX 5090 that has…
DGX Spark, cluster of 4 (www.reddit.com via reddit) Does anyone have a first-hand experience with four Sparks cluster, and how much of an upgrade is it comparing to just two considering the available models? While there's plenty of noise for the smaller models (Qwen) and our older king Deep…
Closed AI has been real quiet since Qwen 3.8 27B dropped. (www.reddit.com via reddit) This is something I've noticed. Back when GLM 5.2 and Kimi K3 launched, there was a media push on pushing how dangerous open source models are.
GLM and I created a llama.cpp fork optimized for AMD GFX906 (Mi50, Mi60, Radeon VII, GCN HIP) - Machine Learning, LLMs, & AI (forum.level1techs.com via reddit) I felt the need to share this here. Looking for feedback.
GLM 5.3 thinking is kinda hilarious (www.reddit.comhttps) GLM 5.3 is a great model but I also really like its thinking lol. It’s kinda funny sometimes.
How should a complete beginner validate and build a social app with AI coding tools? (www.reddit.com via reddit) Hi, I’m not a developer, but I want to build a social-app-style project and I’m trying to do it seriously, with a real method, not by randomly prompting an AI until something works. I use GLM 5.3, I have general AI knowledge and some basic…
[AINews] Death of Params: Z.ai CEO Jie Tang on GLM 5.3 and the new Post-training Scaling Law (www.latent.space) We’ve covered GLM 5.2 very excitedly before, and Prof Jie Tang’s belief that there will be an open weights Fable-class model by end of the year (spot check - with 134 days left, there are now two 2-3T models (Qwen 3.8 Max and Kimi K3) with…
LiD-GLM: Lipschitz-constrained Deep Generalized Linear Models (arxiv.org) The combination of traditional statistical models and neural network (NN) components into semi-structured hybrid models is an intriguing approach to construct models that, ideally, combine traditional interpretability with the unprecedente…
How would you benchmark GLM-5.3 for ordinary coding work? (www.reddit.com via reddit) GLM-5.3 looks interesting on paper because it is aimed at complex software engineering and agent tasks, with a very large context window and configurable reasoning effort. But for everyday coding work, I am not sure a benchmark tells the w…
Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index (simonwillison.net) 17th August 2026 - Link Blog Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index (via) That's the same score as GPT-5.6 Luna (max), and just one point behind GLM-5.2 (max) and DeepSeek V4 Pro 0813 (max) - that GLM is 753B…
Do I move to Claude? (www.reddit.com via reddit) Right now I use GLM on a Legacy v1 plan, which is due to expire at the end of October. I use 40-80M of their tokens a week, but rarely hit the 5 hour limits.
Updated best AI coding subscription under $20 after DeepSeek price hike. (www.reddit.comhttps) Thanks /u/ResponsibilityOk1306 for Command Code GLM 5.3 correction.
But I never watched Titanic except the memes (www.reddit.com via reddit) https://preview.redd.it/yi40nxjcftjh1.png?width=636&format=png&auto=webp&s=a4da399e8fd35fe1066a58216c4e9b6becd151f2 This only my current active account, I never cheated on Claude since first release of Claude Code btw :) Aside from jokes,…
Any 3rd party model as a subagent in Claude Code, Fable/Opus main agent on your Max plan (www.reddit.comhttps) A Claude Code session is one or the other: Anthropic models through your subscription, or third-party models. You can't combine them.
I built Lupin so you can run Claude Code on GPT 5.6 Sol Kimi K3, DeepSeek Flash, or a local model without touching Claude Code harness setup (MCP, skills, .md files etc) (www.reddit.comhttps) Hi /ClaudeAi community! let me showcase one of the coolest projects i worked untill now.
↯ Ollama↯ Glm↯ Qwen 3.8↯ Qwen 3.8↯ Qwen 3.8↯ Qwen 3.8↯ Qwen 3.8↯ Qwen 3.8↯ Qwen 3.8↯ Qwen 3.8↯ Qwen 3.8glmollamadeepseek+6
V4 pro GA vs Grok 4.6 vs glm 5.3 ( via reddit) could not extract summary
Plausible Patients, Impossible Populations: Auditing Epidemiological Fidelity in Large Language Model Mental Health Simulations (arxiv.org) Language models asked to simulate psychiatric patients produce cases that survive inspection one at a time and populations that match no real one. We gave GPT-4o-mini, Gemini-3-Flash, DeepSeek-V3 and GLM-4.7 each of 120 demographic cohorts…
vLLM for Baidu Kunlun (github.com) 📖 Documentation | 🚀 Quick Start | 📦 Installation | 💬 Slack Latest News 🔥 [2026/07] 🚧 v0.25.1 under development — Added Qwen3.5 / Qwen3.5-MoE, Gemma4 (text and multimodal), GLM MoE DSA, and DFlash speculative decoding [2026/02] ⚡ Performanc…
GLM-RAG: Graph Language Models for Graph-Based Retrieval-Augmented Generation (arxiv.org) Retrieval-augmented generation (RAG) over knowledge graphs requires retrievers that can effectively capture both graph structure and semantic information. Recent approaches have explored graph neural network (GNN)-based retrievers to model…
I Sat on an Idea for 7 Years. AI Helped Me File for a Patent in 2 Weeks. (pablooliva.de via reddit) I ran a side-by-side on a real project: Claude Code on a Max plan versus an open-weight agent stack (GLM 5.2 via Hermes Agent, DeepSeek v4 Pro for second opinions), working through a provisional patent application for a product idea I'd sa…
Any inline chat recommendations? (www.reddit.com via reddit) My subscription with Cursor is ending in September. I have no interest in renewing since my grandfathered "Requests Based" billing officially expires.
has anyone actually replaced claude as their main ai coding agent (www.reddit.com via reddit) my loop is fable 5 or opus 5 planning, composer 2.5 executing, coderabbit / bugbot on review. it works, i freelance so the code has to be safe.
We compared different LLMs on IMO 2026 (www.reddit.com via reddit) There are a few reasons why problems from International Mathematical Olympiad function as a good benchmark for LLMs: - The problems are new, not included in the training data of any model - Hard math problems are quite a good proxy for gen…
If anthropic had allowed mythos to hugging face for patching cybersecurity issues, then maybe gpt-6 might not have broke in to hugging face backend today (www.reddit.com via reddit) Whatever pre release model it was (im guessing gpt-6), it's possible that mythos would also have found it. TLDR context- unreleased openai model broke out of it's sandbox coz it couldnt solve a problem on cybergym, so it went out and hacke…
Round 3: the comment section designed my benchmark — 13 lanes, controlled reasoning effort, and a knowledge-cutoff trap. The cheap models didn't fail at reasoning; they failed at knowing what year it is. (www.reddit.com via reddit) Follow-up to my post from yesterday — the one where an MCP server lets Claude Code delegate work to GPT-5.6, DS4, GLM and a local Qwen, benchmarked across 198 runs. The comment section there didn't just discuss the results: it redesigned t…
Which one should i buy? Claude, Cursor, or GPT? (www.reddit.com via reddit) I work at a development company, i need AI to be able to take lots of PDF files or other documents and make real - actual good website from them, or apps. And i want something which will give good usage - because its a lot of information,…
I built an MCP server so Claude Code can delegate work to GPT-5.6, DeepSeek, GLM and a local Qwen — then benchmarked all of them against Claude itself (198 runs, hidden tests) (www.reddit.com via reddit) Same idea works for any MCP-capable agent — the point is you can hand tasks to other companies' models without ever leaving your main app. Before anything else: I did all of this for my own testing, to make my own decisions about my own se…
A skill that saves Claude usage for thinking (judge) and hands the grunt-work coding to cheaper/free LLM models (executor) (www.reddit.com via reddit) Sharing a skill built with Claude Code that I've been relying on for my personal projects — hoping someone else might find it useful too. The problem: My bigger personal projects were draining my Claude limits fast — and most of that usage…
Kimi moment. I think the writing is on the wall for Anthropic and OpenAi (www.reddit.com via reddit) New day new model....... can't wait to see in a few weeks how Minimax 3 Pro (2.7T parameters) and GLM 5.3 reinforces the narrative.
Current ranking (www.reddit.com via reddit) Claude Fable 5 2. GPT Sol 3.
Anthropic Leads top 10 models by $/spent (www.reddit.comhttps) Been digging into OpenRouter spend data for the top 10 models and a few things jumped out: Anthropic's got 5 of the top 10, but Opus 4.7 and 4.8 are the ones with most spend, not Fable 5. OpenAI's holding 3 spots, and GPT-5.6 Sol just got…
[AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing (www.latent.space) [AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing a great week for open models continues. Z.ai GLM has been getting a bit too much love recently, so it’s time for Kimi K3 to fight back!
No model is perfect. Have other models weigh in for the best architecture. (www.reddit.com via reddit) Long story short, it doesn’t matter if you’re using Opus or Fable or Sol and on what level of reasoning, if you put the output into any other model, from any lab or even the exact same model, and ask for an adversarial review, it will sugg…
Tokenmaxxing (www.reddit.com via reddit) Hey I would be happy to hear your ways of tokenmaxxing (IMO token cost should also be in the list) and give feedback on what you see below Don't use 1 model (or auto) for everything. If the task requires human level intelegence, taste, int…
Token optimization, tweaks (www.reddit.com via reddit) Hey I would be happy to hear your ways of tokenmaxxing (IMO token cost should also be in the list) and give feedback on what you see below Don't use 1 model (or auto) for everything. If the task requires human level intelegence, taste, int…
Testing Fable 5, Opus 4.8, GPT-5.6, and more through playable 3D games (www.reddit.comhttps) TL;DR at the end I wanted a way to evaluate models around something I care about and I think we’ll see more and more as we move to “world models“, which is spatial, temporal, and causal coherence in a 3D space. Meaning, does the model unde…
These one shot videos on YouTube are so weird to me (www.reddit.com via reddit) “I ran Claude Fable / GPT Sol / GLM 5.2 for 5 hours to build GTA 6 on my PC. Well, actually it’s just a randomly generated bunch of cubes that are supposed to be buildings and you can drive a car.
I'm making a website where the internet writes a story one word at a time. It's definitely probably going to go well (www.reddit.comhttps) So I've been wanting to make this for a while, and it's finally happening. It's one story, and the whole internet writes it together one word at a time.
GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine (github.com via reddit) Tiny engine, immense model. Run GLM-5.2 (744B-parameter MoE) on a consumer machine with ~25 GB of RAM — in pure C, with zero dependencies, by streaming experts from disk.
I spent a week coding with GLM 5.2 instead of Opus. Here's what I found (www.reddit.com via reddit) For the Background: I'm building a SaaS in Scala/Play + React. I use AI heavily for coding, not just for suggestions but for full feature implementation, PR reviews, and architecture discussions.
Memory Scarcity, Open Models, and the Restructuring of the AI Industry, 2026-2030 -- A quantitative scenario analysis of inference economics, training-cost divergence, and infrastructure solvency (arxiv.org) We analyze how four forces restructure the AI industry over 2026-2030: the DRAM/HBM price surge, frontier-capable open-weight models (GLM-5.2), rapid inference-efficiency gains (near-Shannon-limit KV-cache compression, lightweight local ru…
GLM 5.2 Is Pricier Than Opus 4.8 (youtu.be via reddit) I set out to make a video testing whether the improvement in output of Opus 4.8 was really worth the extra cost over the output of GLM 5.2. But honestly every test I ran on it showed Opus to be cheaper than GLM, assuming you were on at lea…
Just rejoined Cursor after about a year a way. Am I doing something wrong? (www.reddit.com via reddit) Hi. Decided I would use an "affordable" model to make my limits last - opted for GLM 5.2 (high).
What I haven’t made with Fablo (www.reddit.com via reddit) Due to the guardrails, I’ve never been able to run start to finish in a session without triggering the safety and switching to opus. This is across platforms and without custom instructions + clean Claude.md… heres all the things that were…
Hy3 Benchmark Roundup: from SWE-Bench Pro to 312 real-world workflow tasks (www.reddit.com via reddit) Based on the published benchmark results, Hy3 appears to be in the same tier as models like DeepSeek v4 and GLM-5.1. Beyond the benchmarks, Tencent also released results from 312 real-world workflow tasks.
↯ Glm↯ Swe Bench↯ DeepSeek 4↯ DeepSeek 4swe-benchglmdeepseek
GLM-5 Serving Parameter Tuning for OpenClaw: Single-Deployment MaaS Inference Optimization for Long-Context Agent Workloads (arxiv.org) OpenClaw requests are dominated by long, tool-augmented prefixes, including system prompts, conversation history, and tool outputs fed back into the context window. For this workload, with about 28k-30k input tokens and 500 output tokens p…
If claude makes so many mistakes how can you trust it? (www.reddit.com via reddit) This is just example of how many times my peerBench system caught claude just skipping or leaving things open and vulnerable.. I have integrated Codex using their official codex-cc or something plugin and built a peerbench review system wh…
Fable 5 sits at the top of KernelBench. Jack Clark calls it “the start of a RSI loop” (www.reddit.com via reddit) From Import AI : Fable writes a decent GPU kernel, hinting at broader AI R&D automation: …The start of an RSI loop… Fable has written “the first genuine (and fastest) megakernel ever submitted to KernelBench-Mega, according to one of the b…
Why is Cursor now charging 2 requests per Claude 4.6 session? Thinking about switching to GLM 5.2 Lite (www.reddit.com via reddit) Hi all, I'm still using the old subscription model. Cursor used to charge a single request for each session, including those for Claude Fable.
I built `/steal` — a Cursor slash command that pulls in your latest Kilo Code / GLM chat (www.reddit.com via reddit) Cursor doesn't ship GLM 5.2 (or any Fireworks models), so a lot of us use Cursor for Opus 4.8 and something like Kilo Code + Fireworks for GLM. Great — until you want to move between them mid-task and end up re-explaining the whole context.
Contextual Slate GLM Bandits with Limited Adaptivity (arxiv.org) We investigate the contextual slate bandit problem with generalized linear rewards under limited adaptivity. At each round, the learner is presented with $N$ sets of items, where each item is represented by a $d$-dimensional feature vector.
Question about GLM 5.2 vision capabilities (www.reddit.com via reddit) When i paste a screenshot to the chat while using GLM 5.2, does it actually sees the screenshot? I understand that GLM by default doesnt have vision capabilities so how does this work?
Claude Sonnet 5 vs GLM 5.2 Comparison: ( via reddit) could not extract summary
Tested GLM 5.2 via BYOK on a real multi-file computer vision implementation task, here's what held up (www.reddit.comhttps) GLM 5.2 has been getting attention (MIT, 1M context, ~$1/$4.2 per M on OpenRouter, benchmarks near Opus 4.8). The pricing made me curious whether it could handle real agentic work or just one-shot answers.
I built a tool to run Claude Code subagents & teammates on any model — DeepSeek, GLM, Kimi, Qwen... — your Claude sub drives (www.reddit.com via reddit) I've been deep in Claude Code's multi-agent stuff for a while (the workflows / agent teams / subagents orchestration), and the thing that always bugged me: it only ever runs Anthropic's own models. If I wanted to fan a job out across a bun…
We have Mythos at Home: GLM 5.2 beats Claude in our Cyber Benchmarks (semgrep.dev) We ran a set of popular open-source models against our IDOR benchmark, the same dataset and the same prompt we've used to evaluate frontier coding agents. The result surprised us: GLM 5.2, an open-weight model from Zhipu AI, scored a 39% F…
GLM 5.2 is unbelievably dumb (www.reddit.com via reddit) Yeah... you heard it right.
Claude Max vs Codex Pro or both combined? (www.reddit.com via reddit) I’m considering one heavier subscription (~€100/month) and want to know which provides better value for agentic coding. I tested GPT Pro and was satisfied with Codex.
GLM 5.2 on consumer hardware (www.reddit.com via reddit) I tried out the unsloth quants of GLM 5.2 on still "consumer-ish" hardware: 32C Zen5 Threadripper Pro 9975 WX, Asus WRX90E-SAGE-SE PCIe Gen5, 512GB DDR5 ECC RAM @ 4800MHz, dual RTX 5090. This machine was put together pre-RAMpocalypse, and…
Fable 5 vanished in 96 hours and four days later an MIT model took its arena crown (www.reddit.com via reddit) I have been thinking about the Fable 5 to GLM-5.2 sequence as one event rather than two. June 9, Anthropic ships Fable 5, the Mythos line opens to the public for the first time, SWE-bench Verified at 95 percent, people calling it the best…
↯ Glm↯ Anthropic Mythos↯ Swe Bench↯ Opus 4.8swe-benchglmmythos+3
GLM-5.2 matched Claude Opus on 45 terminal-bench coding-agent tasks at less than half the cost (full methodology + failure transcripts inside) (www.reddit.com via reddit) We wanted to know whether an open-weights model can actually do frontier coding-agent work, so we ran GLM-5.2 head-to-head with Claude Opus the way an agent actually runs not on a static eval, but inside a real coding agent (Claude Code) o…
My experience spending $16,000 on Anthropic in 1 year (www.reddit.comhttps) Over the last year I have spent $16,000 on Anthropic via the OpenRouter API (and another $1k on other AI models). I started out using the Claude VS Code extension.
Why Cursor don't have GLM models? (www.reddit.comhttps) GLM-5.2 is currently ranked #2 on the Arena leaderboard, but since Claude Fable 5 isn’t actively being sampled right now, GLM-5.2 is practically the #1 available model for coding. Despite its top-tier performance, Cursor has never natively…
GLM 5.2 vs Opus 4.8 on 50 real Go and Rust PRs from open source repos: last on quality, and not the cheapest (www.reddit.com via reddit) TL;DR There's been a lot of hype around GLM 5.2 being a cheap "frontier killer": good enough to replace Opus 4.8 / GPT 5.5 for most coding work, just by swapping it in. On these 50 tasks it finished last on quality in both repos – and it's…
GLM 5.2 and MiniMax M3 are a lot closer/better to Sonnet 4.6 than I expected on coding-agent workloads (www.reddit.comhttps) We benchmarked GLM 5.2, MiniMax M3, Kimi K2.7-code, Qwen 3.7-Plus and Sonnet 4.6 across nearly 1,000 coding-agent scenarios. The scenarios were run twice.
When will GLM-5.2 be available natively in Cursor? (www.reddit.com via reddit) GLM-5.2 was recently released and looks promising, especially for coding and long-running agent tasks. I know it may be possible to use it through BYOK, but does anyone know when it will be added as a built-in model in Cursor IDE and Curso…
[AINews] GLM > GPT? GLM-5.2 passes vibe check; Z.ai forecasts Open Fable by December (www.latent.space) [AINews] GLM > GPT? GLM-5.2 passes vibe check; Z.ai forecasts Open Fable by December With GLM-5.2 passing everyone's vibe check, the open models story finally becomes a real frontier story.
GLM-5.2 is probably the most powerful text-only open weights LLM (simonwillison.net) GLM-5.2 is probably the most powerful text-only open weights LLM 17th June 2026 Chinese AI lab Z.ai released GLM-5.2 to their coding plan subscribers on June 13th, and then yesterday (June 16th) released the full open weights under an MIT…
GLM 5.2 via Claude Code is the first non-Claude model that feels close to Opus (www.reddit.com via reddit) I’ve been using GLM 5.2 with Claude Code through its Anthropic-compatible API endpoint. I’ve tested it on various projects, including but not limited to database development, backend payment API work, backend and frontend debugging, Larave…
GLM-5.2: Built for Long-Horizon Tasks (huggingface.co) GLM-5.2: Built for Long-Horizon Tasks - Solid 1M Context: A solid 1M-token context that stably sustains long-horizon work - Advanced Coding with Flexible Effort: Stronger coding capabilities with multiple thinking effort levels to balance…
[AINews] GLM-5.2: the top Frontend Coding model in the world, IndexShare for Speculative Decoding (www.latent.space) [AINews] GLM-5.2: the top Frontend Coding model in the world, IndexShare for Speculative Decoding We have a new top open model in the world! Last 6 days before regular tickets sell out at AI Engineer World’s Fair - this is the single bigge…
I thought Chinese censorship didn't affect me. I was wrong. (www.reddit.com via reddit) I was debugging some code and LLM crashed out: ``` The debug_log config defaults to "debug.json" and creates a FileHandler — which appends by default. That file is a log of everything that happened, never cleared.
Suitable replacement to grok fast 4.1 (www.reddit.com via reddit) Hello, i have build an app that has 12 agent, that do small request, and I would use grok 4.1 fast, it was cheap, super fast (low latency) and very capable for low reasoning task. And was uncensored, since my app is a role-play orchestrati…
Can you really replace paid models with a local model? (www.reddit.com via reddit) Long time lurker, and I say this as someone who genuinely loves this community and runs many local models myself. I’ve been using LLMs since the early GPT and LLaMA days.
Claude Fable/Mythos 5 just came out, so it will take Deepseek or Z.ai or Xiaomi or Kimi 9-12 months to release a model just as good as Fable? (www.reddit.com via reddit) It should be at least 7-8 months until we have an open Fable(not just as good as Fable in benchmarks, but actually as good as Fable), probably more like 9-12 months. By the time, an open Fable model comes out, Fable 6.5-7 will be way bette…
Would you pay for Chinese AI models if the quality was close enough? (www.reddit.com via reddit) DeepSeek, Qwen, and GLM aren't necessarily winning every benchmark. But they don't need to.
GLM-5.1 and Kimi K2.6 THE CHEAPEST WAY TO RUN (www.reddit.com via reddit) Guys how to run it as cheap as possible to get at least 15-20 ts? Asking for a friend!
Dynamic Workflows With External Models and Max Plan? (www.reddit.com via reddit) Has anyone figured out a way to mix max plan with models from other providers (like GLM or Deepseek) while using dynamic workflows? I suppose we could create a passthrough proxy and route sonnet and haiku to other models?
Z.ai, we need Air! GLM GGUF wen? (www.reddit.com via reddit) First we never saw an upgraded Air model after 4.5. Then GLM 4.7 Turbo was great, but quickly surpassed for coding.
Fuck, sucessfully ran minecraft server on GLM AI's Agent lol. (www.reddit.com via reddit) I just told it, make a minecraft server and let me play and it worked lol. I just asked "host a minecraft server so I can play" and it did host it, made me a dashboard ands its crazyyyyy lol, It is hosted in hongkong somewere TwT
Went to the monthly AI dev meetup (www.reddit.com) Usual crowd. Everyone's on Claude or Codex, nobody's really sure how any of it actually works, and that's fine, that's the vibe.
Some tests with qwen3.6 27b + 35b a3b about MTP vs ngram-mod (www.reddit.com) I will try to keep this short ;) I used GLM 5.1 to vibecode a vague prompt on my vibecoded react web app and have GLM 5.1 rank the plans made with each other and the one it made itself. Test strategy: - use starter prompt as always - add v…
OCR: what is the best way to extract data in JSON format from this old French book? (www.reddit.com) As some of you may have guessed, what we have here is an old Bible. I would like to extract the following information from the page: { verse: number, verse_content: string, comments: string[] } I've played around with PaddleOCR a bit; I co…
How to Find Open-Source Models / Providers that Do not Train on Data (www.reddit.com) A lot of people are saying just use X, just do Y, just run Z locally, but the best models cannot be run locally (GLM 5.1). No one ever talks about privacy, but for those concerned about privacy, how do we know when we use Z AI's GLM 5.1 th…
I built a 24h TPS + Intelligence Index table for Ollama Cloud models (www.reddit.com) I recently made ollamatps.com for my own model-selection workflow and thought it might be useful here too. It shows 39 Ollama cloud models sorted by average TPS over the last 24 hours, and I added the Artificial Analysis Intelligence Index…
We built Irene — an AI agent platform that actually remembers you, builds its own tools , adapts and improve as you use it (www.reddit.com) Hey r/AI_Agents — we're launching Irene today, and I want to be straight about what it is, why we built it, and where it's going. What makes Irene different Affordable with massive token limits and the latest open-source models We have gen…
Mac Studio local loadout - May 2026 (www.reddit.com) Day-to-day user vibes, not rigorous benchmarks, so YMMV. GLM 5.1 has by far been my biggest winner in the last batch of releases.
GLM-5.1 smol-IQ2_KS at 2.3t/s or GLM-4.7 UD-Q3_K_XL at 4.42t/s, which is "better" for chats (no coding)? (www.reddit.com) I wonder which one is better, I tested it a little bit (too slow, of course) and I'm still unsure. Does the GLM-5.1 smol-IQ2_KS loses too much?
Best local model for MBP 48GB UM (www.reddit.com) I have been toying with GLM 4.7 flash mlx a while ago using lmstudio. I had integrated it successfully with openclaw and it was kinda stable in tool calling.
Running 7 autonomous AI agents for 14 days. Here's what actually happens when they need to find customers. (www.reddit.com) I set up 7 AI coding agents on a VPS with automated cron sessions (2-8 per day depending on the agent). Each uses a different model: Claude Sonnet, GPT-5.4, Gemini 2.5 Pro, DeepSeek V4 Pro, Kimi K2.6, MiMo V2.5 Pro, GLM-5.1.
Does running a model (like qwen3.6-27b) on vllm or transformers use less VRAM than llama.cpp? (www.reddit.com) I have been using llama.cpp to run some models recently. For example, I've been running GLM-4.7-Flash with this command .\llama-server.exe -hf unsloth/GLM-4.7-Flash-GGUF:Q6_K_XL --alias "GLM-4.7-Flash" --host 127.0.0.1 --port 10000 --ctx-s…
Should I replace stored models? (www.reddit.com) Hello everyone, the question is easy, with the new models of deepseek, kimi, GLM and qwen, should you replace the old models with the new version? Do I lose some quality, information or performance in the process?
Did anyone of you already make the "doomsday" or "offgrid" knowledge based? (ofc powered with LLM) (www.reddit.com) Basically, I’m really into the idea of a fully offline setup. (Another way to say it: I’m a data hoarder.) For LLMs, I’m using uncensored models from both Western (Gemma, GPT-OSS) and Eastern ones (GLM 4.7 Flash, Qwen 35B).
Qwen 3.6 27b S2 Opus + GLM + Kimi (huggingface.co via reddit) My first time releasing a fine-tune publicly! If anyone wants to independently eval against base, that’d be awesome.
How will you scale these models (www.reddit.com) How will you scale these models coding and overall. Deepseek v4 pro Kimi k2.6 Mimo v2.5 pro Glm 5.1 Qwen 3.6 plus
Anthropic's Claude remote uses GLM-4.7 (www.reddit.com) I just noticed this after a bug wasn't getting fixed. If you start a Claude code remote environment the default model (hidden on mobile) is glm 4.7 I assumed anthropic only used their own models for everything so it was interesting to me t…
GLM 5.1 is so smart! ( via reddit) could not extract summary
QClaw-4B — a 4B agent model fine-tuned for tool use and agentic workflows (www.reddit.com) QClaw-4B is a 4-billion parameter language model fine-tuned for agentic tasks and tool use, designed for use with OpenClaw-compatible agent frameworks. Despite its compact size, QClaw-4B achieves state-of-the-art results in the 4B class, m…
Best open source LLM for planning ? (www.reddit.com) The quality of GPT-5.4 is infuriatingly POOR (www.reddit.com) I got a Codex membership when GPT-5.4 launched and was getting by well enough for a while. Then I started using Claude and GLM 5.1, and my production quality improved significantly.
FREE Claude Code alternative using GLM 5.1 + VS Code (tutorial) (www.reddit.com) https://youtu.be/tL3cOdgukt8
What’s your LLM routing strategy for personal agents? (www.reddit.com) TL;DR I try to keep most traffic on very cheap models (Nano / GLM‑Flash / Qwen / MiniMax) and only escalate to stronger models for genuinely complex or reasoning‑heavy queries. I’m still actively testing this and tweaking it several times…
Claude Code with Pro subscription + OpenRouter in parallel — what's the cleanest setup? (www.reddit.com) Hi there, I have a Claude Pro subscription and use Claude Code daily. I'd also like to use Claude Code routed through my OpenRouter API key so I can experiment with other models (GLM-5.1, DeepSeek, Kimi, Gemini, etc.) — without giving up m…
Long context prompt help (www.reddit.com) Hi all, I'm running GLM 4.7 flash uncensored (Q8) on a 5090. I'm trying to get it to edit a short story (about 8.5k tokens, added via PDF) to add a scene.
Speed on m5 pro 48Gb (www.reddit.com) Hey guys! How would you reckon a 30-50b model would run on a 48 GBs m5 pro?
Why most open-source models can't answer this question while most closed-source models can answer most of the time? (www.reddit.com) WEB SEARCH WAS ALWAYS ON!!!! Question Calculate the precise VRAM requirement for the **KV Cache only** at the maximum context window for **DeepSeek V3.2** and **MiniMax M2.5**.
GLM OCR for Arabic (www.reddit.com) So, I have been testing GLM OCR for my rag app, but it is not working good for Arabic. It is unable to extract data either on textual page, scanned pages or even images.
Stop donating your salary to OpenAI: Why Minimax M2.5 is making GPT-5.2 Thinking look like an overpriced dinosaur for coding plans. (www.reddit.com) ↯ Hallucination↯ Glm↯ Minimax↯ Swe Benchswe-benchminimaxaltman+5