event

Glm

378 items · started 2026-02-14 · ongoing (last activity 2026-09-19)

  1. Imagine how badly distillers and Chinese hosts undercutting Claude would shit their pants if Anthropic offered high throughput first-party support via Cerebras or another wafer-based host with DeepSeek 4.1f, GLM 5.3,.etc. so everyone looki…

  2. Model Overview GLM-5.3-Flash/GLM-5.3-FlashX is the first native multimodal model in the GLM-5 series, delivering stronger intelligence than GLM-5.2 at an exceptionally low cost.- Highly Efficient Hybrid Architecture - Native Multimodal Vis…

  3. On September 18, 2026, a developer going by ferstar published a reverse-engineering walkthrough of ZCode, the AI coding desktop app from Z.ai, the Beijing-headquartered company behind the GLM family of open-weight models - the same models…

  4. could not extract summary

  5. September 15, 2026Blog Public PreviewThird-partyv5.3 Z.ai GLM 5.3 A third-party open source text model from Z.ai, hosted by Mistral for long-context coding and agentic workflows. The model is served without Mistral modifications.

  6. I'm doing web developement, and game development for a hobby project. I've tried lots of harnesses / IDE's - Best I've found is VSCodium.+ Cline + Openrouter, using discounted models (GLM 5.3 Flash is 50% off atm for example) I used Cursor…

  7. Intro In the last few months our aistack team has been on a quest to get a grip on what it takes to own your own AI stack. We’ve looked into the differences in cost and performance when using APIs, renting or buying GPUs, and started ident…

  8. External cache transfers can succeed while a hybrid language model resumes from an inconsistent state. We examine the full 45-layer GLM-5.3-Flash model, using the RedHatAI/ GLM-5.3-Flash-NVFP4 quantized checkpoint with vLLM and LMCache und…

  9. Mouse on GLM-5.3-Flash 23 of 30 FrontierHarness tasks on Z.ai's Flash model, next to the published GLM-5.3 harness runs, for $6.72 in tokens. Community results shared this week ran GLM-5.3 and GLM-5.3-Flash from Z.ai through five coding ag…

  10. September 2026 · By James Mann Is GLM-5.3-Flash Mythos-level at Cyber? We ran GLM-5.3-Flash on ExploitBench with a budget of 1 billion tokens per vulnerability.

  11. Hi HN, I've been using Claude Code, Codex and Cursor quite a bit lately, and I found myself checking their usage limits all the time. Most of the tools I found for this live on the desktop or in the menu bar.

  12. cognition-claude-proxy A local proxy that lets Claude Code (or any Anthropic-API client) use Devin's model catalog — SWE-2, GLM-5.2, DeepSeek V4.1 Flash, and 200+ others — as its backend. It translates the Anthropic Messages API to the Con…

  13. I bought credits on codexapi.pro (https://codexapi.pro/) after seeing their promos for cheap "unlimited" coding sessions with Claude Code and Codex CLI. In practice, the service was constantly dropping connections: 502 Bad Gateway errors s…

  14. I just started using GLM 5.3 Flash with Claude Code; I'm using GSD framework and one of the sub-agents spawned was reporting progress as normal. 正在清理 03.3.1-02-PLAN.md 中的 files_note 元素 translates to Cleaning up the files_note element in 03…

  15. Get to know Cadenya We’re developers who love to build. We set out to create a yes-code platform that makes building agents feel like the best parts of building software.

  16. I'm planning to use AI seriously for coding, roughly 80% GLM 5.3 Flash and 20% Kimi K3 for harder tasks. I mainly care about large projects, debugging, refactoring, agentic coding and value for money.

  17. WebKit MiniBrowser compiled with Fil-C on top of Linux userland compiled with Fil-C. GTK4, Weston, etc - all compiled with Fil-C.

  18. Red-teaming GoodMem with GLM 5.3 We used GLM 5.3 to red-team GoodMem. How we defined the tests, what the agent found, what we fixed, and how we verified the fixes.

  19. Glm and claude are coding bosses now , wich one is better

  20. Big Claude user, but Anthropic got stingy as hell with the limits. I used to barely touch my weekly allowance; now I can burn through 20% in a day and I'm cooked in ~2 days.

  21. I'm horrified by my Fable and Astra token spend (I'm subscribed to $200 plans for each one). Therefore, the question I asked myself was whether Sonnet still holds up as a good sub-agent?

  22. 10-task GLM 5.3 harness bench: claude, opencode, pi, zcode, hermes and 3code I'm performing a series of harness benchmarks on the same 10 SWE-bench verified tasks representatively chosen for difficulty. This is far from a perfect measure a…

  23. chess5.ai Human vs LLM · Five Games Pit yourself against GPT, Claude, Gemini, Grok, Muse Spark, Mistral, DeepSeek, Kimi, Qwen, GLM, or MiniMax across five classic boards. How it works - Human vs model, or model vs model — with spectating a…

  24. I am working on a big project with Claude Code only context7 mcp added no others tools. With opus 5 is all ok it seems to remember what we have done days before follow the repo conventions etc.

  25. I'd been building this for myself when Peter released CodexBar back in November. CodexBar does far more, 69 providers and a bundled CLI, and it's genuinely good, 21k stars and 119 releases since.

  26. GLM-5.2 RL weight transfer in 4 seconds using NIXL and ModelExpress GLM-5.2 RL weight transfer in 4 seconds using NIXL and ModelExpress In RL at 1T Scale, we detailed how prime-rl trains trillion-parameter models like GLM-5 with sub-5-minu…

  27. What is your opinion about cursor ? I mean is it better than using the official coding applications for the agents (liek using glm 5.3 at zcode or cursor, what is the difference?)

  28. I'd like to share an opinion on the current situation in AI ops, not the coding/tech side. Right now the market has a pretty clear pattern.

  29. Attention makes the sequence all equally available, but KDA requires the model to turn a sequence into a finite state. This compression is naturally lossy, but it forces the model to extract relevant patterns in the context, and more to th…

  30. Why I Hate Benchmarks: Moving From GPT to GLM-5.3 Flash A cautionary tale about how a cheap, capable model swap became a production systems migration. I hate benchmarks because they make the model look like the product.

  31. GLM-5.3-Flash is an excellent model. Running it locally on 4 x RTX 6K Pros with great results.

  32. proxy-llms Run any model in t3's Codex and Claude tabs. A GPT model in the Claude tab, GLM or Kimi K3 in the Codex tab, plugged and unplugged on demand.

  33. Based on artificialanalysis.ai, Kimi K3 and GLM-5.3 are more intelligent than the new Gemini 3.8 Flash[1][2]. Gemini 3.8 Flash comes in eighth place with a score of 59, just after Kimi K3 and GLM-5.3, with scores of 60 for both of them.

  34. Mushroom identification with AI: GPT-5.6-Sol, Gemini 3.7 Flash, GLM-5.3-Flash and Claude Fable 5.1 benchmarked on poisonous and edible species of FungiTastic. A lot of dangerous errors.

  35. GLM 5.3 CRACK — Uncensored FP8 General-purpose weight-level uncensoring · native FP8 speed on Hopper a CRACK release by dealignai · Twitter @dealignai What this is Full-spectrum general-purpose uncensor of GLM-5.3-FP8. Refusal behavior is…

  36. LocalMaxxing Get started Models Reports Hardware Benchmarks More + Submit Get started Leaderboard Decode calculator Models Reports Hardware Benchmarks Marketplace Rentals Pro API Docs Language English 简体中文 繁體中文 日本語 한국어 Español Français Deu…

  37. Vibe on X: "Something else to play with this week: GLM 5.2 is now available in Vibe Code on the Pro and Team plans, and served by Mistral in Europe. Generous usage limits included.

  38. Abliteration.ai Unrestricted models. Governed by your policy.

  39. Inference for Kimi and GLM. Built for research and coding agents that plan, call tools for hours, and reason over large contexts.

  40. When are we going to have available glm 5.3?

  41. GLM-5.3-Flash NVFP4 — 4× DGX Spark, switchless-ring TP4 + DFlash2 Serve GLM-5.3-Flash (NVFP4) across four NVIDIA DGX Spark (GB10 / sm_121) nodes as one tensor-parallel engine — joined by a switchless RoCE ring and accelerated by the DFlash…

  42. I am trying to set up an email triage setup for my needs. After a bit of discovery I started exploring using Hermes Agent as an option with Claude.

  43. I literally use it as the meme stats, Anthropic may lost that low cost tier war with models like GLM 5.3 flash and GPT Luna I can't think they can compete in terms of price/performance in this tier

  44. I'm not a big fan of vendor-lock-in. Anyone tried a subscription based provider with cursor?

  45. Models & Pricing All prices are per million tokens. Checkpoint storage is charged at $0.10 per GB per month.

  46. What GLM-5.3 Flash running on Chinese hardware actually means Z.AI confirmed that their most recent model release was running all inference on Chinese manufactured hardware. While no doubt an impressive feat, Western companies still have a…

  47. This model is featured because its Hugging Face README.md includes an hfviewer architecture visualization. Architecture graph for zai-org/GLM-5.3.

  48. WARP — Weight-Aware Runtime and Paging (formerly WASTE) WARP is an embeddable inference engine written in C, with no third-party runtime dependencies. It keeps the model trunk in memory, streams selected experts directly from disk, and use…

  49. Z.ai on X: "GLM-5.3 is now open-weight. Our most capable model for agentic coding and cyber defense is now available to download, run, and customize.

  50. Planning to get M5 Ultra 512GB to run GLM-5.3-mlx-mxfp4. However, I think the Apple SSD is a rip off.

  51. Everything has changed in two months: DS4 0731 flash was the start of a wave that is taking open weights to paradise. It is easy to think that Qwen 4 and GLM 6 will be on par with Mythos.

  52. I recently got into the habit of using qwen uncensored models for just local reverse engineering workflows, some the flagship cloud models even glm models refuse. But theres only so much intelligence i can pack into 16gb vram.

  53. A few months ago, I created the WARP engine (formerly WASTE) to run Kimi K3, the complete 2.78-trillion-parameter model, on macOS. GLM-5.3-Flash shares many architectural similarities with Kimi K3, so I added support for it as well.

  54. Read our How to Run GLM-5.3-Flash Guide! Unsloth Dynamic 3.0 achieves superior accuracy & outperforms other leading quants.

  55. I've forked https://github.com/tonyd2wild/GLM-5.3-Flash-NVFP4-2x-DGX-Spark and make it run on sm120. I'm using it right now - got 1,4M context (5,45 sessions 262k each) 3,7kt/s PP and 160 - 230t/s TG (MTP enabled) You can make vllm Docker…

  56. Hey all! I'm finally doing some cool stuff with my "thinking heater" (h/t u/-TV-Stand-).

  57. how many models are going by 3 rn like Gemini 3.7, qwen 3.8, minimax m3, kimi k3, hy3, deeseek v4 flash, glm 5.3. (not deepseek and glm but close enough)

  58. Spent yesterday getting Qwen3.8 Flash and GLM 5.3 Flash up and running on my cluster of 4 x DGX Sparks with a view to replacing DeepSeek 0731... but..

  59. The promise has been fulfilled.

  60. I personally stopped reading the launch table once GLM-5, MiniMax M2.5, and Gemini 3 Deep Think dropped in two days and all claimed the same coding, reasoning, and agent wins. They optimize different constraints.

  61. zai-org/GLM-5.3-Flash GLM 5.3 Flash is an LLM listed in RunInfra Model APIs. RunInfra serves it as zai-org/GLM-5.3-Flash at $0.10 per 1M input tokens and $0.40 per 1M output tokens.

  62. could not extract summary

  63. Hi HN, Hearing a lot of buzz around GLM-5.3, which I expect to be the best open-source coding model with the weights dropping soon, I wanted to test it where I actually do my work. I just mapped GLM-5.3 to TokenGo so I could swap out the b…

  64. curl -s https://api.github.com/repos/ggml-org/llama.cpp/pulls/27742 | jq -c '{draft,state,merged}' for r in unsloth/GLM-5.3-Flash-GGUF unsloth/Qwen3.8-Flash-Next-GGUF; do echo "== $r" curl -s "https://huggingface.co/api/models/$r" \ | jq -…

  65. Megathread for discussing the release of GLM-5.3-Flash. Quants Fine-Tunes & Abliterations Chat Templates Inference Server Support & Configuration Experiences, Benchmarks & Model Comparisons We'll try to clean up future duplicates around th…

  66. GLM-5.3-Flash 👋 Join our WeChat or Discord community. 📖 Check out the GLM-5.3-Flash blog and GLM-5 Technical report.

  67. Z.AI confirms Ox Alpha is a GLM model, plans to release its weights The Beijing lab told Bloomberg it built the anonymous coding model and planned to release its weights on August 26. By RuntimeWire Staff · Published Primary source: Bloomb…

  68. I am using 20x plan. I am seeing a lot of hallucination and scope drift, and extremely slow execution.

  69. Anthropic is always whining about Chinese competitors distilling their models, but current Chinese models have basically caught up to frontier capabilities and released them for free. Since the weights are open and available, shouldn’t it…

  70. Prompt injection and gzip-NCD compression analysis reveal that OX Alpha, a mysterious LLM on OpenRouter, is GLM developed by Z.ai. A stealthy new model called OX Alpha has popped up on https://openrouter.ai/ and is climbing up the leaderbo…

  71. Shares the same tokenizer as GLM, seems to perform as well as larger models, and has vision support. Could this be Baseten's GLM Vision (maybe further finetuned) or an official GLM model?

  72. I am trying to get to 128G of VRAM with reasonable compute and bandwidth to run multiple models in parallel. DS4 Flash or GLM 5.3 in hybrid mode with custom checkpoints.

  73. I used to use Qwen3.6-35B-A3B with llama.cpp and connecting it to the VSCodium extension called "Continue." My computer is running a Intel(R) Core(TM) Ultra 7 265K (3.90 GHz) with 128 GB of DDR5 RAM and an Nvidia Geforce RTX 5090 that has…

  74. Seventeen models on the Featherbench leaderboard: glm-5.3 leads at 100%, eight tie at 96%, and checker errors flip the safety ranking. See who tops the board.

  75. Does anyone have a first-hand experience with four Sparks cluster, and how much of an upgrade is it comparing to just two considering the available models? While there's plenty of noise for the smaller models (Qwen) and our older king Deep…

  76. Amazon kept shutting down my tablet, so I spent $266 on four AI models to own it My Amazon Fire HD tablet cost $114.26 on eBay in November 2022, new and sealed. Owning it for real cost another $266.15: Kimi K3 found the exploit for $164.25…

  77. This is something I've noticed. Back when GLM 5.2 and Kimi K3 launched, there was a media push on pushing how dangerous open source models are.

  78. I felt the need to share this here. Looking for feedback.

  79. GLM 5.3 is a great model but I also really like its thinking lol. It’s kinda funny sometimes.

  80. Getting GLM-5.2 NVFP4 Post-Training off the ground The goal was deceptively simple to state: take GLM-5.2, a 744B-parameter mixture-of-experts model quantized to 4-bit NVFP4, attach a bf16 LoRA adapter, and train it with reinforcement lear…

  81. Ox Alpha is 100% a GLM model by Zhipu AI, and it looks like its *almost* mythos class from very early results. It's very likely going to be called GLM-6 and it's absolutely mogging every frontier model in SWE and Cyber benchmarks (DeepSWE…

  82. Since these models are smart, I decided to rerun the Baba Is Bench, to see how the models fare on a puzzle game. Even though the game is popular, we check for spoilers - and to our surprise, there are no signs of models knowing solutions a…

  83. Hi, I’m not a developer, but I want to build a social-app-style project and I’m trying to do it seriously, with a real method, not by randomly prompting an AI until something works. I use GLM 5.3, I have general AI knowledge and some basic…

  84. We’ve covered GLM 5.2 very excitedly before, and Prof Jie Tang’s belief that there will be an open weights Fable-class model by end of the year (spot check - with 134 days left, there are now two 2-3T models (Qwen 3.8 Max and Kimi K3) with…

  85. Multimodal, Multitask One catalog spanning multi-modal inputs and multi-task outputs: OCR, detection, segmentation, pose, keypoints, and more. Document OCR, captioning, and multi-modal chat: every visual capability behind one MCP server.

  86. GLM-5.3 achieves 60 on the Artificial Analysis Intelligence Index, on par with Kimi K3 and up 7 points from GLM-5.2. Once the weights are released it will be tied as the leading open weights model @Zai_org has just launched GLM-5.3, whic…

  87. GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves on GLM-5.2 in coding and in the balance…

  88. What We Learned Moving Our Agent Loops from Anthropic to GLM Why we moved most of Unblocked's agent traffic from Claude Opus to GLM 5.2, what the blind A/Bs and the ledger actually showed, and what broke on the way. TL;DR: We moved most of…

  89. The combination of traditional statistical models and neural network (NN) components into semi-structured hybrid models is an intriguing approach to construct models that, ideally, combine traditional interpretability with the unprecedente…

  90. GLM-5.3 looks interesting on paper because it is aimed at complex software engineering and agent tasks, with a very large context window and configurable reasoning effort. But for everyday coding work, I am not sure a benchmark tells the w…

  91. 17th August 2026 - Link Blog Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index (via) That's the same score as GPT-5.6 Luna (max), and just one point behind GLM-5.2 (max) and DeepSeek V4 Pro 0813 (max) - that GLM is 753B…

  92. Right now I use GLM on a Legacy v1 plan, which is due to expire at the end of October. I use 40-80M of their tokens a week, but rarely hit the 5 hour limits.

  93. Thanks /u/ResponsibilityOk1306 for Command Code GLM 5.3 correction.

  94. https://preview.redd.it/yi40nxjcftjh1.png?width=636&format=png&auto=webp&s=a4da399e8fd35fe1066a58216c4e9b6becd151f2 This only my current active account, I never cheated on Claude since first release of Claude Code btw :) Aside from jokes,…

  95. I wrapped this small followup to stolen-thoughts.com in a gist for easier reading and wanted to share it here. My hunch is that this might not just be reasoning distillation; it could even be benchmark distillation.

  96. GLM-5.3 — The Official Desktop Auditor for Z.AI's Cyber-Engine Z.AI’s GLM-5.3 didn’t just beat the industry benchmarks—it rewrote them. By applying extreme reinforcement learning on real-world cybersecurity tasks, GLM-5.3 autonomously disc…

  97. OpenVuln Public frontend for OpenVuln. This Space is built automatically from the root Dockerfile and serves the Vite application with nginx on port 7860.

  98. When GLM-5.2 helped Hugging Face investigate an incident in which an AI autonomously bypassed its own safeguards, it highlighted a broader shift. AI is becoming part of both cyber offense and cyber defense.

  99. GLM-5.3: How Chinese labs keep stride with the frontier Hint: It’s really not a distillation story. Housekeeping: I’m traveling so cannot make a voiceover for this post.

  100. could not extract summary

  101. A Claude Code session is one or the other: Anthropic models through your subscription, or third-party models. You can't combine them.

  102. https://hnhiring.azuanz.com I originally built this to scratch my own itch while job hunting. The main problem for me with the monthly Who Is Hiring thread was finding relevant jobs by location.

  103. Should AI's be required to answer the direct question "are you an AI" with a clear yes? Currently they are not universally required to do so.

  104. Hi /ClaudeAi community! let me showcase one of the coolest projects i worked untill now.

  105. could not extract summary

  106. Language models asked to simulate psychiatric patients produce cases that survive inspection one at a time and populations that match no real one. We gave GPT-4o-mini, Gemini-3-Flash, DeepSeek-V3 and GLM-4.7 each of 120 demographic cohorts…

  107. My take: this notice from DeepSeek might actually be a brilliant marketing move. The logic here is to urge users to ramp up their usage over the next 2 to 3 months.

  108. Memcode is a new developer platform currently in public beta. It includes a coding agent, chat, reusable agents, DataHub, and a Lovable-style website generator.

  109. Smaller, faster, safer: running Kimi and GLM at scale Workers AI runs inference for some of the best open models in the world on GPUs in Cloudflare data centers close to your users. Two of the most capable, and most demanding, are Moonshot…

  110. Hi all, I’ve been working on this project for a while now and wanted to share it here for feedback and contributions. You can test drive it in the playground at: https://usepopkorn.dev I’ve been calling it Popkorn.

  111. Z.ai Open Platform Java SDK 中文文档 | English The official Java SDK for Z.ai platforms, providing a unified interface to access powerful AI capabilities including chat completion, embeddings, image generation, audio processing, and more. ✨ Fe…

  112. 📖 Documentation | 🚀 Quick Start | 📦 Installation | 💬 Slack Latest News 🔥 [2026/07] 🚧 v0.25.1 under development — Added Qwen3.5 / Qwen3.5-MoE, Gemma4 (text and multimodal), GLM MoE DSA, and DFlash speculative decoding [2026/02] ⚡ Performanc…

  113. Retrieval-augmented generation (RAG) over knowledge graphs requires retrievers that can effectively capture both graph structure and semantic information. Recent approaches have explored graph neural network (GNN)-based retrievers to model…

  114. Verified inference means cryptographic proof that the exact open model you requested produced your output, not a cheaper or quantized stand-in. Run frontier open models like GLM-5.2, billed by the token.

  115. The danger frontier: low-cost, evasive, abundant malware As part of Incalmo’s mission to make AI safely ubiquitous, we do safety research on the frontier cyber capabilities of models. Recently, to help anti-virus systems stay ahead of the…

  116. MOST POPULAR AI - AI and ML Too many AI agents can get in each other's way For enterprise agents, less is more - AI and ML Impostor Chinese models pretend they're Claude Researchers find GLM and Kimi can adopt Claude's identity, but the ev…

  117. I ran the 62-item politicalcompass.org test 30 times each on sixteen models: OpenAI's GPT-5.x and GPT-4o, Claude, Gemini, Grok, Llama, Mistral, and China's DeepSeek, Qwen, Kimi and GLM. Fifteen land in the libertarian-left quadrant.

  118. I ran a side-by-side on a real project: Claude Code on a Max plan versus an open-weight agent stack (GLM 5.2 via Hermes Agent, DeepSeek v4 Pro for second opinions), working through a provisional patent application for a product idea I'd sa…

  119. My subscription with Cursor is ending in September. I have no interest in renewing since my grandfathered "Requests Based" billing officially expires.

  120. Coinbase Switches to Chinese AI Models GLM and Kimi, Cuts AI Spending by 50% - Coinbase has defaulted engineers to GLM 5.2 from Zhipu and Kimi 2.7 from Moonshot AI through its internal LLM gateway, cutting AI spending by nearly 50% [1] - G…

  121. my loop is fable 5 or opus 5 planning, composer 2.5 executing, coderabbit / bugbot on review. it works, i freelance so the code has to be safe.

  122. There are a few reasons why problems from International Mathematical Olympiad function as a good benchmark for LLMs: - The problems are new, not included in the training data of any model - Hard math problems are quite a good proxy for gen…

  123. A month ago, GLM-5.2 was released. As part of our day-zero support, we built the fastest API in the world for GLM-5.2, with peak speeds of 280 tokens per second and average speeds around 100 tokens per second.

  124. I'm a mechatronics designer with a background in control systems, robotics, PCB design, and embedded hardware. I design physical systems: motors, sensors, microcontrollers, and real-time control loops.

  125. This wasn’t a photo finish. Sakana: Fugu Ultra controlled the matchup on practical writing and coding tasks, while GLM 5.2 showed flashes of polish in a couple of narrower instruction-following spots.

  126. 📄 Production-Grade SLM-Powered OCR Course 📄 Build a self-scaling, event-driven OCR pipeline on Kubernetes (AKS / GKE) with Qwen 3.5 + the GLM-OCR SDK Table of Contents Table of Contents Course Overview Who is this course for? Course Breakd…

  127. OpenAI-compatible unrestricted AI and uncensored LLM API for enterprise teams: AI red teaming, cybersecurity, trust and safety, synthetic data/evals, ML research, and defense/government contractor workflows. Policy Gateway adds policy-as-c…

  128. quantprobe Placement beats budget Where your bits sit — which layers, which memory tier — matters more than how many you have. Four falsification-tested laws for running big LLMs on hardware you already own, every number measured on one 20…

  129. I’ve been building Echo (https://echo.tracerml.ai/), an experiment in making one AI system out of a pool of open-weight models rather than choosing a single model and using it for every task. It started with a simple experiment.

  130. Whatever pre release model it was (im guessing gpt-6), it's possible that mythos would also have found it. TLDR context- unreleased openai model broke out of it's sandbox coz it couldnt solve a problem on cybergym, so it went out and hacke…

  131. Hugging Face uses open-weights Z.ai GLM 5.2 to battle attacker after commercial frontier model refusal Hugging Face Inc., an open-source artificial intelligence platform often described as the “GitHub of machine learning,” found itself for…

  132. Follow-up to my post from yesterday — the one where an MCP server lets Claude Code delegate work to GPT-5.6, DS4, GLM and a local Qwen, benchmarked across 198 runs. The comment section there didn't just discuss the results: it redesigned t…

  133. could not extract summary

  134. This matchup wasn’t close. GPT-5.6 Sol dominated the practical details that decide real-world usefulness: tighter instruction-following, cleaner formatting, and fewer correctness slips.

  135. I work at a development company, i need AI to be able to take lots of PDF files or other documents and make real - actual good website from them, or apps. And i want something which will give good usage - because its a lot of information,…

  136. https://t.co/v9huIornsf elvis@omarsar0ArticleMiniMax M3: How Sparse Attention Makes Long-Horizon Agents Practical GLM 5.2 has taken over much of the AI timeline lately, and most of the conversation has centered on how it stacks up against…

  137. Same idea works for any MCP-capable agent — the point is you can hand tasks to other companies' models without ever leaving your main app. Before anything else: I did all of this for my own testing, to make my own decisions about my own se…

  138. A full GLM-5.2 scan found 30.168% K15 charged-format accounting. A separate byte-split representation was decoded bit-for-bit across all 59,509 BF16 tensors at 24.967% reduction.

  139. MSE-GLM — Command Reference Matrix-Structured Edge — Graph Language Model. Deterministic, zero-weight, explainable.

  140. Sharing a skill built with Claude Code that I've been relying on for my personal projects — hoping someone else might find it useful too. The problem: My bigger personal projects were draining my Claude limits fast — and most of that usage…

  141. Fable 5 · GPT 5.6 Sol · Kimi K3 · Grok 4.5 · Gemini 3.5 Flash · MiMo V2.5 Pro · MiniMax M3 · GLM 5.2 — low-poly, semi-realistic, very realistic.

  142. https://t.co/iJsDrlGy45 Harry Partridge@part_harry_ArticleGLM 5.2 With VisionGLM 5.2 is one of the best currently available open source language models. However, unlike other flagship models like Qwen, Kimi and Minimax, GLM 5.2 does not su…

  143. New day new model....... can't wait to see in a few weeks how Minimax 3 Pro (2.7T parameters) and GLM 5.3 reinforces the narrative.

  144. Claude Fable 5 ⁠2. GPT Sol ⁠3.

  145. Been digging into OpenRouter spend data for the top 10 models and a few things jumped out: Anthropic's got 5 of the top 10, but Opus 4.7 and 4.8 are the ones with most spend, not Fable 5. OpenAI's holding 3 spots, and GPT-5.6 Sol just got…

  146. [AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing a great week for open models continues. Z.ai GLM has been getting a bit too much love recently, so it’s time for Kimi K3 to fight back!

  147. We serve GLM-5.2 to teams building agents. Same open-weight model, same OpenAI-compatible API — but we route it across more than one backend, and while swapping one in we found something worth writing down: the backend you pick changes tim…

  148. We ported the puzzle game Baba Is You to the Harbor framework, and benchmarked current models, including Claude, GPT, Gemini, GLM and DeepSeek. A human Twitcher is 4x faster than Fable 5.

  149. Long story short, it doesn’t matter if you’re using Opus or Fable or Sol and on what level of reasoning, if you put the output into any other model, from any lab or even the exact same model, and ask for an adversarial review, it will sugg…

  150. Hey I would be happy to hear your ways of tokenmaxxing (IMO token cost should also be in the list) and give feedback on what you see below Don't use 1 model (or auto) for everything. If the task requires human level intelegence, taste, int…

  151. Hey I would be happy to hear your ways of tokenmaxxing (IMO token cost should also be in the list) and give feedback on what you see below Don't use 1 model (or auto) for everything. If the task requires human level intelegence, taste, int…

  152. TL;DR at the end I wanted a way to evaluate models around something I care about and I think we’ll see more and more as we move to “world models“, which is spatial, temporal, and causal coherence in a 3D space. Meaning, does the model unde…

  153. An AI-agent cold-tuned our GLM-5.2 serving. Human engineering leveled it up for real production traffic.

  154. GLM-5.2 (unpruned) on 4× DGX Spark — depth, max context, or multi-user Serve the unpruned GLM-5.2 (QuantTrio Int4-Int8Mix, all 256 experts) across four GB10 Sparks — one recipe, four lanes, one KV budget spent on depth or width**. TP4 + DC…

  155. Came across this today. Canopy Wave just added GLM-5.2, and they're giving away a few 7-day trial accounts for anyone who wants to test it.

  156. How to Code with GLM 5.2 on OpenCode Coding with GLM 5.2 on OpenCode is a bet on your own engineering. Decide the architecture and interfaces first, let the cheap model write the code on a fixed twenty-dollar Ollama plan, and bring a str I…

  157. Claude Fable 5 is Anthropic's newest flagship and, in our testing and the independent benchmarks, the strongest model in the current lineup for polished front-end work. It is selectable in Playcode's AI website builder.

  158. “I ran Claude Fable / GPT Sol / GLM 5.2 for 5 hours to build GTA 6 on my PC. Well, actually it’s just a randomly generated bunch of cubes that are supposed to be buildings and you can drive a car.

  159. https://t.co/3i0qSTbjql Bing Xu@bingxu_ArticleThe Great Wave Has Arrived (from GLM CEO Jie Tang)-- Bing Xu's Note --- I came across an internal GLM letter on the Chinese app RedNote, purportedly written by @jietang, and translated the Chin…

  160. So I've been wanting to make this for a while, and it's finally happening. It's one story, and the whole internet writes it together one word at a time.

  161. ive been seeing a recurring claim that open (weight) models 6 months behind the frontier are good enough for the majority of ‘work’. if you've had a concrete task in the last month where GLM/DeepSeek/Kimi/Qwen failed and Opus/Fable/GPT suc…

  162. Tiny engine, immense model. Run GLM-5.2 (744B-parameter MoE) on a consumer machine with ~25 GB of RAM — in pure C, with zero dependencies, by streaming experts from disk.

  163. GLM 5.2 is (nearly) as accurate as a human book-keeper at less than 1% of the cost We evaluated the performance of GLM 5.2, an open weights AI model, on the task of quarterly value-added tax (VAT) return preparation for a small UK business…

  164. For the Background: I'm building a SaaS in Scala/Play + React. I use AI heavily for coding, not just for suggestions but for full feature implementation, PR reviews, and architecture discussions.

  165. A few days ago I found myself trying out GLM 5.2 and was really positively impressed. The capabilities and security I was getting from this LLM are similar to those I've gotten from models like Claude or GPT, and this really surprised me.

  166. We analyze how four forces restructure the AI industry over 2026-2030: the DRAM/HBM price surge, frontier-capable open-weight models (GLM-5.2), rapid inference-efficiency gains (near-Shannon-limit KV-cache compression, lightweight local ru…

  167. I set out to make a video testing whether the improvement in output of Opus 4.8 was really worth the extra cost over the output of GLM 5.2. But honestly every test I ran on it showed Opus to be cheaper than GLM, assuming you were on at lea…

  168. Compare AI model performance on Harvey LAB-AA Benchmark Leaderboard. Artificial Analysis' implementation of Harvey's Legal Agent Benchmark (LAB), testing AI agents on real-world legal work from Harvey's dataset of 120 private tasks spannin…

  169. Big news for DwarfStar users: I got DeepSeek v4 Flash and GLM 5.2 working with Tensor Parallelism across 2 M5Max 128GB MacBooks via RDMA. It is especially interesting for GLM since otherwise, fully resident, can't fit a machine that money…

  170. Hi. Decided I would use an "affordable" model to make my limits last - opted for GLM 5.2 (high).

  171. VisionBridge Give text-only LLMs vision through a tiny OpenAI-compatible proxy. VisionBridge sits between your chat UI and your models.

  172. Due to the guardrails, I’ve never been able to run start to finish in a session without triggering the safety and switching to opus. This is across platforms and without custom instructions + clean Claude.md… heres all the things that were…

  173. mulot [-4285F4?logo=googlechrome&logoColor=white)]() Agentic AI web pentester that drives a browser. An open-weights LLM (GLM-5.2, Gemma or Qwen) drives a real headless Chromium through a Burp-style toolkit and works a target the way a hum…

  174. Based on the published benchmark results, Hy3 appears to be in the same tier as models like DeepSeek v4 and GLM-5.1. Beyond the benchmarks, Tencent also released results from 312 real-world workflow tasks.

  175. OpenClaw requests are dominated by long, tool-augmented prefixes, including system prompts, conversation history, and tool outputs fed back into the context window. For this workload, with about 28k-30k input tokens and 500 output tokens p…

  176. Consider modest to major changes in a code base as particularly important to design properly in advance towards having a good intuition of the libraries and implementation details with the end result…

  177. Hiya! So I've been playing around with having Claude make videos for a bit now even had some success posting the results to TikTok (and setup a whole pipeline so Claude can generate and post autonomously).

  178. This is just example of how many times my peerBench system caught claude just skipping or leaving things open and vulnerable.. I have integrated Codex using their official codex-cc or something plugin and built a peerbench review system wh…

  179. GLM 5.2 and the coming AI margin collapse (part 1) This is a two part series focusing on what I believe is perhaps the least understood upcoming shift in AI economics. If you've enjoyed this and want to be notified about the second post, p…

  180. From Import AI : Fable writes a decent GPU kernel, hinting at broader AI R&D automation: …The start of an RSI loop… Fable has written “the first genuine (and fastest) megakernel ever submitted to KernelBench-Mega, according to one of the b…

  181. GLM-5.2 from Z.ai is an open-source frontier model that competes with Anthropic’s Fable 5 and Opus 4.8 at roughly one-fifth the cost. With a 1M token context window and strong long-horizon coding performance, it’s changing the economics of…

  182. Hi all, I'm still using the old subscription model. Cursor used to charge a single request for each session, including those for Claude Fable.

  183. Cursor doesn't ship GLM 5.2 (or any Fireworks models), so a lot of us use Cursor for Opus 4.8 and something like Kilo Code + Fireworks for GLM. Great — until you want to move between them mid-task and end up re-explaining the whole context.

  184. GLM Coding Lite 适合轻量开发任务 包含基础使用额度 - 适合轻量迭代与小型仓库 - 持续获得最新旗舰模型与功能 - 支持 20+ 编程工具,包括 ZCode 深度适配 GLM Coding Plan GLM 深度适配 ZCode,让 Agent 编程更稳定、更高效。 GLM Coding 适合轻量开发任务 包含基础使用额度 GLM Coding 适合专业开发工作流 包含 Lite 全部权益与 5 倍 Lite 额度 GLM Coding 适合高频与大规模任务…

  185. Comparing GLM 5.2 on several long horizon tasks including systems programming, web, creative writing and video generation and that involves working with multiple programming languages.

  186. Introducing ZCode, the official development environment for GLM-5.2 - GLM Coding Plan subscribers: now 1.5x usage quota in ZCode - BYOK supported: works with your existing subscriptions and APIs - Available on macOS, Windows, and Linux D…

  187. Saw this last night: https://x.com/AnthropicAI/status/2072163884430229756 I'm taking the weekend off to spend time with my lady, so got up early today to start preparing to get back to work on a sailboat simulator I started building when F…

  188. GLM-5.2’s Code Reviews Are Only as Good as Your Prompt GLM-5.2 from Z.ai has been one of the most talked-about open-weight models since it launched, and we have made it our daily driver to see how it performs on various coding tasks. We al…

  189. We investigate the contextual slate bandit problem with generalized linear rewards under limited adaptivity. At each round, the learner is presented with $N$ sets of items, where each item is represented by a $d$-dimensional feature vector.

  190. When i paste a screenshot to the chat while using GLM 5.2, does it actually sees the screenshot? I understand that GLM by default doesnt have vision capabilities so how does this work?

  191. I like Claude Desktop, so I created my own I like the Claude Desktop application and use it a lot, both for my work and in a personal capacity. I also love to build stuff, and with the recent release of GLM-5.2 I decided to see if I could…

  192. could not extract summary

  193. A controlled comparison across 19 paired runs spanning 19 repository forks — 38 individual workflow executions total — running an identical paper-implementation pipeline (remyxai/outrider — Claude Code under the hood, with glm-5.2 routed a…

  194. We just open-sourced the internal system we built at Assembled for running coding agents as a team. Coding agents worked well for individual engineers, but the surrounding workflow was a bit of a mess.

  195. GLM 5.2 has been getting attention (MIT, 1M context, ~$1/$4.2 per M on OpenRouter, benchmarks near Opus 4.8). The pricing made me curious whether it could handle real agentic work or just one-shot answers.

  196. Relay An open-source, dark-mode desktop coding agent — built for people who want to use non-mainstream LLM providers, not just the big three. Relay is an Electron app that puts DeepSeek, Qwen, GLM, Kimi, MiniMax, and other open/Chinese mod…

  197. China’s Zhipu AI (Z.ai) released its open-weight GLM-5.2, and some researchers have claimed that it matches Mythos in certain bug-finding and cybersecurity scenarios. While GLM lags behind models from Anthropic and OpenAI in other, more ge…

  198. Hi I'm Saoud, founder of Cline. We’ve been impressed with GLM-5.2 and so are introducing a $9.99/month subscription to give you 2-5x discounted access to it and other open weight models like DeepSeek, Kimi, MiniMax, Mimo, and Qwen.

  199. Many cloud providers offering GLM-5.2 seems to only offer FP8. I wonder if anyone has evaluated the difference between FP8 and BF16 in terms of quality?

  200. Skills for Claude Code and Codex are hard to test. What I mean by hard is that there's no standard way to do it.

  201. I've been deep in Claude Code's multi-agent stuff for a while (the workflows / agent teams / subagents orchestration), and the thing that always bugged me: it only ever runs Anthropic's own models. If I wanted to fan a job out across a bun…

  202. We ran a set of popular open-source models against our IDOR benchmark, the same dataset and the same prompt we've used to evaluate frontier coding agents. The result surprised us: GLM 5.2, an open-weight model from Zhipu AI, scored a 39% F…

  203. Recently there is a lot of excitement about GLM-5.2 which is an open-weight MoE LLM performing on Claude Opus 4.5 level in chat arena and overperforming all models except Claude Fable in WebDev arena [1]. Even though it is very good that t…

  204. Analysis of API providers for GLM-5.2 (max) across performance metrics including latency (time to first token), output speed (output tokens per second), price and others. API providers benchmarked include Together AI, FriendliAI, Fireworks…

  205. Yeah... you heard it right.

  206. I’m considering one heavier subscription (~€100/month) and want to know which provides better value for agentic coding. I tested GPT Pro and was satisfied with Codex.

  207. I tried out the unsloth quants of GLM 5.2 on still "consumer-ish" hardware: 32C Zen5 Threadripper Pro 9975 WX, Asus WRX90E-SAGE-SE PCIe Gen5, 512GB DDR5 ECC RAM @ 4800MHz, dual RTX 5090. This machine was put together pre-RAMpocalypse, and…

  208. Until last week, attackers faced a dilemma in using frontier models: even if they could manage the cat-and-mouse game of setting up fake accounts to retain API access to frontier model providers, and even if they could induce models to hel…

  209. I have been thinking about the Fable 5 to GLM-5.2 sequence as one event rather than two. June 9, Anthropic ships Fable 5, the Mythos line opens to the public for the first time, SWE-bench Verified at 95 percent, people calling it the best…

  210. We wanted to know whether an open-weights model can actually do frontier coding-agent work, so we ran GLM-5.2 head-to-head with Claude Opus the way an agent actually runs not on a static eval, but inside a real coding agent (Claude Code) o…

  211. GLM-5.2 vs Claude Opus: Same Code, Less Than Half the Cost We ran GLM-5.2 head to head with Claude Opus the way an agent actually runs: inside a real coding agent, in a real shell, graded by hidden tests. The harness is Claude Code on term…

  212. We just posted results to the DataAgentBench leaderboard scoring 61.37% with GLM 5.2. Please check it out and do share your feedback

  213. Over the last year I have spent $16,000 on Anthropic via the OpenRouter API (and another $1k on other AI models). I started out using the Claude VS Code extension.

  214. GLM-5.2 arrived last week. It boasts excellent benchmarks and looks strong.

  215. I tried to run GLM-5.2 on a 64GB Mac Field notes from an experimental ds4 fork, a 244GB GGUF, and the small horror of sparse models that are sparse in compute but not very friendly to filesystems. I have a weakness for local LLM experiment…

  216. could not extract summary

  217. GLM-5.2 is the step change for open agents A capability threshold I've been carefully monitoring. Housekeeping: Following my “State of the blog” post last week, noting a slight increase in paid features, it’s a good time to remind folks th…

  218. GLM-5.2 is a turning point for coding agents. It's the first model a business would actually pay to replace Claude Opus with.

  219. GLM-5.2 is currently ranked #2 on the Arena leaderboard, but since Claude Fable 5 isn’t actively being sampled right now, GLM-5.2 is practically the #1 available model for coding. Despite its top-tier performance, Cursor has never natively…

  220. A new AI model from China is generating the kind of buzz not seen since DeepSeek's R1 announced China as a serious threat to American chatbot hegemony over a year ago. Silicon Valley's online echo chamber has been alight with intrigue in r…

  221. Genuinely impressed, almost shocked, at how good GLM-5.2 by @zai_org is at coding. This changes things.

  222. Experiment that I've made. The models get access to an E2B sandbox and are instructed to create an ad according to the specifications (they can choose whatever tools they want to use for it, e.g.

  223. 🚢 cc-fleet 🤖 Plug any third-party model into Claude Code's ⚙️ Dynamic Workflows, 👥 Agent Teams, and ⚡ Subagents — from DeepSeek · GLM · Kimi · Qwen … to your Codex subscription, with your main session's auth untouched; no Claude subscripti…

  224. GLM 5.2 vs Composer 2.5 and the premium field on 50 real merged PRs from graphql-go-tools (Go) and sqlparser-rs (Rust). GLM lands last on craft and equivalence in both repos, costs about twice Composer, and writes more code than the human…

  225. TL;DR There's been a lot of hype around GLM 5.2 being a cheap "frontier killer": good enough to replace Opus 4.8 / GPT 5.5 for most coding work, just by swapping it in. On these 50 tasks it finished last on quality in both repos – and it's…

  226. We benchmarked GLM 5.2, MiniMax M3, Kimi K2.7-code, Qwen 3.7-Plus and Sonnet 4.6 across nearly 1,000 coding-agent scenarios. The scenarios were run twice.

  227. https://t.co/JSn0lDCNkB Design Arena@DesignarenaArticleHow GLM-5.2 Beat Fable 5 at Website DesignGLM 5.2 ranks 1st overall on Design Arena’s single-turn, HTML Web Design (Non-Agentic) evaluation, 5 places higher than its predecessor GLM-5.…

  228. GLM-5.2 was recently released and looks promising, especially for coding and long-running agent tasks. I know it may be possible to use it through BYOK, but does anyone know when it will be added as a built-in model in Cursor IDE and Curso…

  229. Thinkbench, our custom evaluation harness, was used to drive both models through the same autonomous coding loop: read files, write files, run shell commands, and stop when the task was complete. The scored suite covered greenfield builds,…

  230. Bigger models are not the way Jun 18, 2026 A shift is happening among major AI labs, who are becoming increasingly skeptical of endless parameter count and training data scaling. The limits of this paradigm were put on the world’s stage wh…

  231. For the complete documentation index, see llms.txt. This page is also available as Markdown.

  232. Introducing GLM-5.2: Frontier Intelligence, Open Weights - Significant improvements in coding and agentic tasks - Strong long-horizon capabilities with a 1M context window - Two levels of reasoning effort: GLM-5.2 (max) pushes the limits,…

  233. [AINews] GLM > GPT? GLM-5.2 passes vibe check; Z.ai forecasts Open Fable by December With GLM-5.2 passing everyone's vibe check, the open models story finally becomes a real frontier story.

  234. Every few weeks the "best open model" crown changes hands. This week it's GLM-5.2, from the Chinese lab Z.ai — and unusually, the claim has teeth: it sits at #1 on the independent Artificial Analysis Intelligence Index.

  235. Running GLM-5.2 5× faster than vLLM, on a runtime that doesn't support it I rented an 8×B200 and tried to run GLM-5.2 on TileRT, the runtime MiMo used to push a 1T model past 1000 tok/s. TileRT doesn't support GLM-5.2, so I reverse-enginee…

  236. Model APIs Instant Access for Frontier Open-Source AI Models Build AI Apps and Agents with High-Performance Model APIs — No Deployment Required. Start Free TrialEverything You Need to Run AI Models Z.ai: GLM 5.1 4 supported capabilities fo…

  237. Overview GLM-5.2 is a flagship model built for the era of long-horizon tasks. With truly usable 1M-token context, it has been tested to handle project-scale engineering context, delivering more stable long-task execution, more reliable adh…

  238. GLM 5.2 playing text adventures I’ve heard some buzz around the new glm 5.2 open-weights model. They say it’s very capable!

  239. GLM-5.2 is probably the most powerful text-only open weights LLM 17th June 2026 Chinese AI lab Z.ai released GLM-5.2 to their coding plan subscribers on June 13th, and then yesterday (June 16th) released the full open weights under an MIT…

  240. I’ve been using GLM 5.2 with Claude Code through its Anthropic-compatible API endpoint. I’ve tested it on various projects, including but not limited to database development, backend payment API work, backend and frontend debugging, Larave…

  241. GLM-5.2 👋 Join our WeChat or Discord community. 📖 Check out the GLM-5.2 blog and GLM-5 Technical report.

  242. GLM-5.2 by Zhipu AI tops BridgeBench reasoning 24 hours after the U.S. banned Fable 5.

  243. June 17, 2026 GLM-5.2 is the new leading open weights model on the Artificial Analysis Intelligence Index Z ai’s GLM-5.2 is the new leading open weights model on the Artificial Analysis Intelligence Index scoring 51 and it sits on the Pare…

  244. GLM-5.2: Built for Long-Horizon Tasks - Solid 1M Context: A solid 1M-token context that stably sustains long-horizon work - Advanced Coding with Flexible Effort: Stronger coding capabilities with multiple thinking effort levels to balance…

  245. Exciting news: GLM-5.2 (Max) ranks #2 in Code Arena: Frontend, with +29pt over Claude Opus 4.7 (Thinking) and only behind Fable 5! GLM-5.2 is the best open model vs Kimi-K2.6 and Minimax-M3 by a large margin.

  246. GLM-5.2 (max) Intelligence, Performance & Price Analysis Model summary IntelligenceUpdated Speed Price Cache Hit Price Verbosity GLM-5.2 (max) is amongst the leading models in intelligence, but particularly expensive when comparing to othe…

  247. [AINews] GLM-5.2: the top Frontend Coding model in the world, IndexShare for Speculative Decoding We have a new top open model in the world! Last 6 days before regular tickets sell out at AI Engineer World’s Fair - this is the single bigge…

  248. Introducing GLM-5.2: Frontier Intelligence, Open Weights - Significant improvements in coding and agentic tasks - Strong long-horizon capabilities with a 1M context window - Two levels of reasoning effort: GLM-5.2 (max) pushes the limits,…

  249. Intelligence should be open, accessible, and ready to build with, empowering every developer, everywhere. GLM-5.2 is now available to all GLM Coding Plan users, including Lite, Pro, Max, and Team plans.

  250. I was debugging some code and LLM crashed out: ``` The debug_log config defaults to "debug.json" and creates a FileHandler — which appends by default. That file is a log of everything that happened, never cleared.

  251. Hello, i have build an app that has 12 agent, that do small request, and I would use grok 4.1 fast, it was cheap, super fast (low latency) and very capable for low reasoning task. And was uncensored, since my app is a role-play orchestrati…

  252. Long time lurker, and I say this as someone who genuinely loves this community and runs many local models myself. I’ve been using LLMs since the early GPT and LLaMA days.

  253. It should be at least 7-8 months until we have an open Fable(not just as good as Fable in benchmarks, but actually as good as Fable), probably more like 9-12 months. By the time, an open Fable model comes out, Fable 6.5-7 will be way bette…

  254. DeepSeek, Qwen, and GLM aren't necessarily winning every benchmark. But they don't need to.

  255. Guys how to run it as cheap as possible to get at least 15-20 ts? Asking for a friend!

  256. Has anyone figured out a way to mix max plan with models from other providers (like GLM or Deepseek) while using dynamic workflows? I suppose we could create a passthrough proxy and route sonnet and haiku to other models?

  257. CCLimitPing (limitping) English | 中文 Keep your Claude Code, Codex, and GLM (Zhipu / Z.ai Coding Plan) rate-limit windows back-to-back. These providers bill on a 5-hour rolling window (plus a weekly cap), and the 5h window starts on your fi…

  258. First we never saw an upgraded Air model after 4.5. Then GLM 4.7 Turbo was great, but quickly surpassed for coding.

  259. I just told it, make a minecraft server and let me play and it worked lol. I just asked "host a minecraft server so I can play" and it did host it, made me a dashboard ands its crazyyyyy lol, It is hosted in hongkong somewere TwT

  260. Hey HN, We believe we have the easiest onboarding from signup to being able to spin up coding agents in slack like Stripe, Ramp & Coinbase. Demo of the onboarding: https://www.tella.tv/video/connecting-cord-to-slack-1-19ep Every signup get…

  261. Been following the infrastructure side of AI more lately and stumbled on this from Zai. They upgraded the network architecture on a thousand-GPU cluster running GLM-5.1 coding inference from the standard ROFT setup to something they built…

  262. Our dev team flagged last week that xAI is retiring grok 4.1 fast. We weren't using it for anything critical but it made me ask something I'd never actually asked: how did we pick the models we're running?

  263. Have you ever wondered, how DeepSeek may make money, and lot of it? They didn't come up with competitive coding plans like GLM, MoonShot and MiniMax.

  264. So, I got interested in local LLMs a few months ago, but, I don't have a background in coding, and I don't know how to code, and I am not good with computers or anything. So far I mainly just was having fun with comparing different local L…

  265. Usual crowd. Everyone's on Claude or Codex, nobody's really sure how any of it actually works, and that's fine, that's the vibe.

  266. 🧠 MoE-on-a-Potato Running a 754-Billion Parameter LLM on a 16GB RAM Consumer PC "Saying it's impossible is not engineering. Saying we don't know how yet is science." MoE-on-a-Potato is an experimental project dedicated to testing the extre…

  267. I've been testing other models but it seems like nothing even come close to Qwen3.6 35B A3B for agentic use. The worse I'd get is a loop sometimes, while Gemma4 produced broken tool calls occasionally and I couldn't even get GLM 4.7 Flash…

  268. I am new to Cursor and still testing the free version. Benchmark for Composer 2.5 indicates it is better than DeepSeek v4 and Glm 5.1.

  269. https://preview.redd.it/i90oxxk7n03h1.png?width=1898&format=png&auto=webp&s=7d219c804fda7dfe122b84fcdb6d0d6883818c68 A while back I came across TradingAgents — a really cool multi-agent LLM stock analysis framework where like a dozen "agen…

  270. I will try to keep this short ;) I used GLM 5.1 to vibecode a vague prompt on my vibecoded react web app and have GLM 5.1 rank the plans made with each other and the one it made itself. Test strategy: - use starter prompt as always - add v…

  271. While we are weathering the gemini 3.5 flash hype, keep in mind that according to arena, GLM and Mimo are better. https://arena.ai/leaderboard/text/coding-no-style-control #7 GLM #9 Mimo #12 Gemini 3.5 Flash

  272. One way I like to test new models, is by one-shoting (with a good prompt) a single webpage clone of the classic arcade game pacman. I usually do 3 attempts and keep the best one.

  273. I built cdesktop with Claude Code — it's an open-source alternative to Anthropic's Claude Code Desktop, running locally on your machine via npx cdesktop. Free, Apache 2.0.

  274. Since the last post I've added: Huxley module (Brave New World style behavioral conditioning) Baudrillard module (synthetic intimacy, trust collapse, simulation) 30 more models including Grok 4.3, GPT-5.5, Gemini 3.1 Pro, GLM-5.1 Multi-jud…

  275. This is my GLM-4.6 model API configuration, and this error is really confusing me. I'm not sure which step went wrong.

  276. As some of you may have guessed, what we have here is an old Bible. I would like to extract the following information from the page: { verse: number, verse_content: string, comments: string[] } I've played around with PaddleOCR a bit; I co…

  277. A lot of people are saying just use X, just do Y, just run Z locally, but the best models cannot be run locally (GLM 5.1). No one ever talks about privacy, but for those concerned about privacy, how do we know when we use Z AI's GLM 5.1 th…

  278. Researchers from the Max Planck Institute recently released FutureSim, an environment in which agents are replayed a temporal slice of the web and are tasked with predicting real-world future events. In their environment, GPT 5.5 leads at…

  279. I recently made ollamatps.com for my own model-selection workflow and thought it might be useful here too. It shows 39 Ollama cloud models sorted by average TPS over the last 24 hours, and I added the Artificial Analysis Intelligence Index…

  280. My latest project, about 60% of the codebase was written with Z.ai's GLM-5.1 model. It's basically a Telegram bot that allows for embedding/downloading media easier within group chats.

  281. Has anyone figured out a provider whose open source models (Kimi, Qwen, GLM e.t.c) can be used reliably in production. I have tested some well known providers and they all suffer from high latency and poor uptime rendering them mostly usel…

  282. 1rok 1rok is a standalone harness for running portfolio-construction agents across OpenAI, Anthropic, Gemini, xAI, DeepSeek, GLM, and OpenRouter against the same financial tool surface. Agents query Alpaca, Yahoo Finance, FRED, and Tavily…

  283. I'm testing prompt-cache behavior for GLM models on Vertex AI MaaS and I'm seeing inconsistent telemetry. I reproduced it with a synthetic long prompt and repeated identical requests.

  284. About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC

  285. I have a GLM subscription that’s marketed as offering 3× higher usage than Claude Pro. I primarily use it through Claude Code CLI as a backup coding model.

  286. grunden.ai är en svensk AI-tjänst för utvecklare, myndigheter och helt vanliga människor. GLM 5.1 (open-weight) med EU-jurisdiktion, ett OpenAI-kompatibelt API och prissättning i kronor.

  287. Hi. I new to using Cursor - coming from Claude Code, Antigravity and most recently GLM coding plan.

  288. Original plan was to use Kimi/GLM for planning and DeepSeek for implementation, but seeing a lot of love for MiMo and Minimax lately. Anyone running a planner + coder split on Opencode?

  289. With the lowering usage limit in Claude, I am thinking of jumping ship to Chinese AI, since the benchmark is already very near compared to Sonnet or Haiku 4.5 , but for a fraction of the price. I am not worried about where is my data endin…

  290. With the lowering usage limit in Claude, I am thinking of jumping ship to Chinese AI, since the benchmark is already very near compared to Sonnet or Haiku 4.5 , but for a fraction of the price. I am not worried about where is my data endin…

  291. Hey r/AI_Agents — we're launching Irene today, and I want to be straight about what it is, why we built it, and where it's going. What makes Irene different Affordable with massive token limits and the latest open-source models We have gen…

  292. Day-to-day user vibes, not rigorous benchmarks, so YMMV. GLM 5.1 has by far been my biggest winner in the last batch of releases.

  293. I wonder which one is better, I tested it a little bit (too slow, of course) and I'm still unsure. Does the GLM-5.1 smol-IQ2_KS loses too much?

  294. GPT and Opus block on certain requests. This didnt use to be the case 2 months ago and I made signficant progress with Opus and then one day I had a 2 week break and then a single prompt to continue the work resulted in refusal.

  295. I've been using GLM 5.1 a lot lately, and I love this model. However I don't love sending all my requests to China.

  296. I have been following the akitaonrails coding benchmark which tests against a fixed rails + Rubyllm + docker task rather than vendor-reported evals. April 2026 update put K2.6 at 87 sitting in tier A (80+), ahead of Qwen 3.6 plus (71), Dee…

  297. Just cancelled my claude subscription due to poor rate limits, gemini cli doesn't really excel in coding from my personal experience, and my local hardware isn't that powerful to run local AI models, and while codex is good, I wanna try so…

  298. I have been toying with GLM 4.7 flash mlx a while ago using lmstudio. I had integrated it successfully with openclaw and it was kinda stable in tool calling.

  299. The benchmark uses adversarial, multi-turn debates across 683 curated motions. Each model pair debates the same motion twice with sides swapped.

  300. We present GLM-5V-Turbo, a step toward native foundation models for multimodal agents. As foundation models are increasingly deployed in real environments, agentic capability depends not only on language reasoning, but also on the ability…

  301. Architecture explains the gap: MiMo's MoE runs more active params per token than Kimi K2.6's optimized routing hence slowest. DeepSeek V4's 'comprehensive' edge is partly MLA: ~75% KV-cache compression makes it far better for long agentic…

  302. I want to run big models like GLM 5.1 or Kimi k2.6. I can buy Mac Studio M3 Ultra with 512gb ram, but PP speed would be ofc bad.

  303. Literally no 3rd party api inference provider is hosting the mimo-2.5 series models from Xiaomi. They seem to be reallly good.

  304. I set up 7 AI coding agents on a VPS with automated cron sessions (2-8 per day depending on the agent). Each uses a different model: Claude Sonnet, GPT-5.4, Gemini 2.5 Pro, DeepSeek V4 Pro, Kimi K2.6, MiMo V2.5 Pro, GLM-5.1.

  305. I have been using llama.cpp to run some models recently. For example, I've been running GLM-4.7-Flash with this command .\llama-server.exe -hf unsloth/GLM-4.7-Flash-GGUF:Q6_K_XL --alias "GLM-4.7-Flash" --host 127.0.0.1 --port 10000 --ctx-s…

  306. I am quite curious as I tried Gemma 4 31B, Qwen 3.6 27B, GLM 4.7 30B and some others in my native language (czech). Gemma performs "best" and considering the fact its "just" 18GB model - it actually blows my mind how well it can respond in…

  307. Detailed Article: https://autobe.dev/articles/local-llm-benchmark-about-backend-generation.html Five months ago I posted the "Hardcore function calling benchmark in backend coding agent" thread here. As I wrote in that post, it was an unco…

  308. Been working on this for a while and finally at a point where it's running in production for a couple of small businesses, so figured I'd share. The thing that kept bugging me about "AI employee" products is that none of them are something…

  309. Which of these do you think we'll get in May? Also, feel free to pick/rank which ones you'd want the most badly: more Gemma4 models (124b?) (other sizes?) more Qwen3.6 models (9b?

  310. I must say that I almost feel no difference in all of the latest models that are coming out. Opus 4.7 is almost equal to 4.6 and 4.5, same about the other GPT models, the Kimi K models and the GLM models they all I feel they’re almost all…

  311. Hello everyone, the question is easy, with the new models of deepseek, kimi, GLM and qwen, should you replace the old models with the new version? Do I lose some quality, information or performance in the process?

  312. Hi, did anyone with an AMD MI50 setup (8x 32GB) test GLM-5 or GLM-5.1? Currently, I have 3x AMD MI50 and I was wondering if it's worth buying another 5 of them and a new PSU.

  313. I received this mail: "Hi developers, Some of you flagged occasional garbled outputs and unexpected behavior when building with the GLM-5 series, especially under heavy workloads. We heard you, reproduced the issues, and the fixes are now…

  314. Basically, I’m really into the idea of a fully offline setup. (Another way to say it: I’m a data hoarder.) For LLMs, I’m using uncensored models from both Western (Gemma, GPT-OSS) and Eastern ones (GLM 4.7 Flash, Qwen 35B).

  315. Some of the larger models (like Llama) weren't available on OpenRouter, so I had to work with what was there. Best small model: Gemma 4 26B For its size, I think it had the best output.

  316. Our belief in Scaling Laws has not only driven continuous breakthroughs in model parameters and data scale, but has also pushed infrastructure engineering toward its limits. This process inevitably comes with growing pains, which we refer…

  317. Hi HN family! I've recently been messing around with open models through ollama (glm-5.1 and kimi-k2.6), and I've been impressed with just how close they are to Claude Sonnet for my needs, especially programming.

  318. I code for a living, close to 7 years now, and I read way too much tech news. TIME dropped their 2026 most influential AI companies list and going through it I see OpenAI, Anthropic, Google, Meta, Amazon, then Zhipu AI sitting right there…

  319. I’ve been looking to buy a coding plan from one of the major open source contributors to give my meager support to them and transition away from Claude. I would love to hear some feedback from the community of their experience with some of…

  320. Full disclosure: I used to program a bit, but I was garbage at it so I found a new career. This was eons ago so I'm not a dev, obviously.

  321. This is a follow up to the previous benchmark and tensor analysis of abliteration techniques across the Qwen model family. Same approach, same toolkit, new model family.

  322. My first time releasing a fine-tune publicly! If anyone wants to independently eval against base, that’d be awesome.

  323. 1.5 tb ram with 128gb vram and a 28 core processor. Mac Pro 2019.

  324. I'm a compsci student and I've been using the 10$ copilot plan for about 2 years now, and it was fine for me since I did a good model distribution taking into account the complexity of the task, I was able to get through the month always u…

  325. Had 327 production traces from a restaurant-reservation agent I wanted to retrain. The plan was to fine-tune a smaller self-hostable model so I could ditch the frontier-API bill.

  326. How will you scale these models coding and overall. Deepseek v4 pro Kimi k2.6 Mimo v2.5 pro Glm 5.1 Qwen 3.6 plus

  327. I just noticed this after a bug wasn't getting fixed. If you start a Claude code remote environment the default model (hidden on mobile) is glm 4.7 I assumed anthropic only used their own models for everything so it was interesting to me t…

  328. so v4 pro dropped and barely anyone is talking about it. feels weird since when kimi k2.6 came out i seen post about it everywhere anyone here tried v4 pro for actual code work?

  329. could not extract summary

  330. After some sglang patching and countless experiments, managed to get reap-ed nvfp4 version running stable and FAST on 4 x RTX 6000 Pros (limited to 350W). Very happy with performance and quality.

  331. QClaw-4B is a 4-billion parameter language model fine-tuned for agentic tasks and tool use, designed for use with OpenClaw-compatible agent frameworks. Despite its compact size, QClaw-4B achieves state-of-the-art results in the 4B class, m…

  332. other companies are slowly going away from open weight, not releasing base models, delaying open weight distribution, not releasing top models (this one I think is fair, but still), and I also noticed they stopped publishing research (old…

  333. I'm usually a Windows person, but I’m currently running a Mac cluster for local LLM orchestration. My setup consists of four 256GB Mac Studios plus one 96GB Mac Studio, giving me about 1.1TB of unified memory.

  334. I created this chart with recent open models from last 6 months. Few might be older than that possibly.

  335. Hi all. Many models on hugging face have been fine tuned with that 3000x opus dataset, but the two I mentioned in the title are missing it.

  336. I’m trying to understand the current open-source LLM landscape beyond surface-level hype. We all got used to the nerfed products of Claude/Geminj so I believe really in opensource as a solution.

  337. I have run two tests on each LLM with OpenCode to check their basic readiness and convenience: - Create IndexNow CLI in Golang (Easy Task) and - Create Migration Map for a website following SiteStructure Strategy. (Complex Task) Tested Qwe…

  338. I got a Codex membership when GPT-5.4 launched and was getting by well enough for a while. Then I started using Claude and GLM 5.1, and my production quality improved significantly.

  339. I've grown increasingly skeptical that public coding benchmarks tell me much about which model is actually worth paying for and worried that as demand continues to spike model providers will silently drop performance. I did a few manual an…

  340. Edit: I’m getting the consensus is that the budget I suggested is not enough for my lil ambitious project. I’d like to reshape the question for the upcoming comments: what’s the minimal budget to achieve my goal?

  341. https://youtu.be/tL3cOdgukt8

  342. TL;DR I try to keep most traffic on very cheap models (Nano / GLM‑Flash / Qwen / MiniMax) and only escalate to stronger models for genuinely complex or reasoning‑heavy queries. I’m still actively testing this and tweaking it several times…

  343. could not extract summary

  344. Hi everyone. It's been a while since I posted (was a lil burned out), but some of you may have seen my older SanityHarness posts.

  345. Hello all just as it sounds. I recently started using GLM 5.1 in cursor 3 but unlike in the past, GLM 5.1 ran through my entire daily budget from summarizing chat context and running commands.

  346. Hi there, I have a Claude Pro subscription and use Claude Code daily. I'd also like to use Claude Code routed through my OpenRouter API key so I can experiment with other models (GLM-5.1, DeepSeek, Kimi, Gemini, etc.) — without giving up m…

  347. Hi all, I'm running GLM 4.7 flash uncensored (Q8) on a 5090. I'm trying to get it to edit a short story (about 8.5k tokens, added via PDF) to add a scene.

  348. So I was just about to give up playing with local models, until I realised I can actually run GLM 5.1 at not too horrible speeds, using this quant https://huggingface.co/ubergarm/GLM-5.1-GGUF/tree/main/IQ2_KL in ik llama. Getting around 6.…

  349. As of mid Apr 2026, I have noticed every model has had a major intelligence drop. And no I'm not talking about just ChatGPT.

  350. So i have been seeing more of those pelican on a bike svg tests and while they work i feel like (and maybe you guys do too) they are getting kinda benchmaxxed so we should switch things up soon and this is my idea generate me a html svg of…

  351. I'm brand new to local LLMs and started with GLM-4.7 Flash q4_K_M. When I run it directly: ollama run glm-4.7-flash:q4_K_M it works pretty decently — nothing amazing, but usable and responsive.

  352. Ever since the company went public, they’ve been making a lot of changes that clearly seem to be prioritizing profit without regard to their customers. For example, with their coding plans: - They promised/advertised that the Lite coding p…

  353. Some more 'sloptuber' content for those who are enjoying it :) Model: unsloth glm 5.1 @ IQ2_XXS UD Prompt 1: Task: in a single web page, build a city based parkour game. wsad controls, moving player aligned with current camera direction.

  354. So I have been running gpt and glm-5.1 side by side lately and tbh the gap is way smaller than what im paying for On SWE-Bench Pro glm-5.1 actually took the top spot globally, beat gpt-5.4 and opus 4.6. overall coding score is like 55 vs g…

  355. I created and run a benchmark for AI models in data analysis tasks. In contrary to other benchmarks, it is not one-prompt benchmark, but I tried to simulate the real work of data analyst.

  356. We’ve been benchmarking a few models on our API platform and got some interesting performance numbers: - MiniMax M2.5 → 0.118s time-to-first-token, 103 tokens/sec - GLM 5.1 → 120 tokens/sec throughput - Kimi K2.5 → 0.643s TTFT, 69 tokens/s…

  357. I spent too much time trying to find one AI dev tool that could do everything. Planning, coding, fixing, reviewing, maybe filing my taxes too It never really worked.

  358. Hey guys! How would you reckon a 30-50b model would run on a 48 GBs m5 pro?

  359. What am I doing wrong here? I can't get models to follow my instructions, pretty much at all.

  360. WEB SEARCH WAS ALWAYS ON!!!! Question Calculate the precise VRAM requirement for the **KV Cache only** at the maximum context window for **DeepSeek V3.2** and **MiniMax M2.5**.

  361. So, I have been testing GLM OCR for my rag app, but it is not working good for Arabic. It is unable to extract data either on textual page, scanned pages or even images.

  362. I know this question has already been asked a thousand times, probably, but... what's the best or close-to-best model I can use with Continue for local IDE-like code autocomplete?

  363. Hey everyone, I'm comparing these two plans side by side for running AI agents daily through OpenClaw (self-hosted AI agent platform): • Ollama Cloud Pro — $20/month • OpenAI Plus — €23/month (~$25) My setup: 3 agents running in parallel (…

  364. GLM 5.1 is dominant in almost every aspect in Design arena, surpassing Opus 4.6 in many tasks. Although user experiences vary dependent on subscription plans for both of those one of them is open source.

  365. Something keeps nagging at me about the Chinese AI space lately. Every few months a new Chinese model drops that closes the gap with US frontier models a little more(not by throwing more compute at it, just genuinely clever engineering at…

← all threads