model roundup
GLM 5.2
-
Coinbase Switches to Chinese AI Models GLM and Kimi, Cuts AI Spending by 50% - Coinbase has defaulted engineers to GLM 5.2 from Zhipu and Kimi 2.7 from Moonshot AI through its internal LLM gateway, cutting AI spending by nearly 50% [1] - G…
-
We built the new fastest API for GLM-5.2 (www.baseten.co via hn)
A month ago, GLM-5.2 was released. As part of our day-zero support, we built the fastest API in the world for GLM-5.2, with peak speeds of 280 tokens per second and average speeds around 100 tokens per second.
-
Sakana: Fugu Ultra vs. GLM 5.2 (runtimewire.com via hn)
This wasn’t a photo finish. Sakana: Fugu Ultra controlled the matchup on practical writing and coding tasks, while GLM 5.2 showed flashes of polish in a couple of narrower instruction-following spots.
-
Show HN: Abliterated GLM 5.2 model, fine-tuned for security and red-team work (abliteration.ai via hn)
OpenAI-compatible unrestricted AI and uncensored LLM API for enterprise teams: AI red teaming, cybersecurity, trust and safety, synthetic data/evals, ML research, and defense/government contractor workflows. Policy Gateway adds policy-as-c…
-
Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models (news.ycombinator.com)
I’ve been building Echo (https://echo.tracerml.ai/), an experiment in making one AI system out of a pool of open-weight models rather than choosing a single model and using it for every task. It started with a simple experiment.
-
Whatever pre release model it was (im guessing gpt-6), it's possible that mythos would also have found it. TLDR context- unreleased openai model broke out of it's sandbox coz it couldnt solve a problem on cybergym, so it went out and hacke…
-
Hugging Face uses open-weights Z.ai GLM 5.2 to battle attacker (siliconangle.com via hn)
Hugging Face uses open-weights Z.ai GLM 5.2 to battle attacker after commercial frontier model refusal Hugging Face Inc., an open-source artificial intelligence platform often described as the “GitHub of machine learning,” found itself for…
-
Hugging Face turns to GLM 5.2 to fend off AI agent attack (venturebeat.com via hn)
could not extract summary
-
MiniMax M3: How Sparse Attention Makes Long-Horizon Agents Practical (twitter.com via hn)
https://t.co/v9huIornsf elvis@omarsar0ArticleMiniMax M3: How Sparse Attention Makes Long-Horizon Agents Practical GLM 5.2 has taken over much of the AI timeline lately, and most of the conversation has centered on how it stacks up against…
-
Lossless model compression experiment: GLM-5.2 in 25% less memory (brianbell-x.github.io via hn)
A full GLM-5.2 scan found 30.168% K15 charged-format accounting. A separate byte-split representation was decoded bit-for-bit across all 59,509 BF16 tensors at 24.967% reduction.
-
Harry Partridge on X: "GLM 5.2 With Vision" / X (twitter.com via hn)
https://t.co/iJsDrlGy45 Harry Partridge@part_harry_ArticleGLM 5.2 With VisionGLM 5.2 is one of the best currently available open source language models. However, unlike other flagship models like Qwen, Kimi and Minimax, GLM 5.2 does not su…
-
The same LLM is 8x slower to first token depending on who serves it (dynoyard.app via hn)
We serve GLM-5.2 to teams building agents. Same open-weight model, same OpenAI-compatible API — but we route it across more than one backend, and while swapping one in we found something worth writing down: the backend you pick changes tim…
-
Autoresearch doubled GLM-5.2 throughput. Production traffic broke it (fparisio.substack.com via hn)
An AI-agent cold-tuned our GLM-5.2 serving. Human engineering leveled it up for real production traffic.
-
Show HN: Self-hosting unpruned GLM-5.2 on a 4-node DGX Spark cluster (github.com via hn)
GLM-5.2 (unpruned) on 4× DGX Spark — depth, max context, or multi-user Serve the unpruned GLM-5.2 (QuantTrio Int4-Int8Mix, all 256 experts) across four GB10 Sparks — one recipe, four lanes, one KV budget spent on depth or width**. TP4 + DC…
-
Show HN: GLM-5.2 is now available via Canopy Wave (twitter.com via hn)
Came across this today. Canopy Wave just added GLM-5.2, and they're giving away a few 7-day trial accounts for anyone who wants to test it.
-
Coding with GLM 5.2 on OpenCode for a flat $20/month (meshintelligence.substack.com via hn)
How to Code with GLM 5.2 on OpenCode Coding with GLM 5.2 on OpenCode is a bet on your own engineering. Decide the architecture and interfaces first, let the cheap model write the code on a fixed twenty-dollar Ollama plan, and bring a str I…
-
These one shot videos on YouTube are so weird to me (www.reddit.com via reddit)
“I ran Claude Fable / GPT Sol / GLM 5.2 for 5 hours to build GTA 6 on my PC. Well, actually it’s just a randomly generated bunch of cubes that are supposed to be buildings and you can drive a car.
-
GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine (github.com via reddit)
Tiny engine, immense model. Run GLM-5.2 (744B-parameter MoE) on a consumer machine with ~25 GB of RAM — in pure C, with zero dependencies, by streaming experts from disk.
-
GLM 5.2 is nearly as accurate as a human book keeper (toot-books.pages.dev via hn)
GLM 5.2 is (nearly) as accurate as a human book-keeper at less than 1% of the cost We evaluated the performance of GLM 5.2, an open weights AI model, on the task of quarterly value-added tax (VAT) return preparation for a small UK business…
-
I spent a week coding with GLM 5.2 instead of Opus. Here's what I found (www.reddit.com via reddit)
For the Background: I'm building a SaaS in Scala/Play + React. I use AI heavily for coding, not just for suggestions but for full feature implementation, PR reviews, and architecture discussions.
-
Getting GLM 5.2 running on my slow computer (news.ycombinator.com)
A few days ago I found myself trying out GLM 5.2 and was really positively impressed. The capabilities and security I was getting from this LLM are similar to those I've gotten from models like Claude or GPT, and this really surprised me.
-
We analyze how four forces restructure the AI industry over 2026-2030: the DRAM/HBM price surge, frontier-capable open-weight models (GLM-5.2), rapid inference-efficiency gains (near-Shannon-limit KV-cache compression, lightweight local ru…
-
I built a tiny proxy that gives GLM 5.2 vision (or any text LLM) – MIT (github.com via hn)
VisionBridge Give text-only LLMs vision through a tiny OpenAI-compatible proxy. VisionBridge sits between your chat UI and your models.
-
AI Agent using a Burp-style toolkit over MCP (github.com via hn)
mulot [-4285F4?logo=googlechrome&logoColor=white)]() Agentic AI web pentester that drives a browser. An open-weights LLM (GLM-5.2, Gemma or Qwen) drives a real headless Chromium through a Burp-style toolkit and works a target the way a hum…
-
Show HN: InstantVideos.org – short documentaries in ~30 seconds (instantvideos.org via hn)
Hiya! So I've been playing around with having Claude make videos for a bit now even had some success posting the results to TikTok (and setup a whole pipeline so Claude can generate and post autonomously).
-
GLM 5.2 and the coming AI margin collapse (martinalderson.com via hn)
GLM 5.2 and the coming AI margin collapse (part 1) This is a two part series focusing on what I believe is perhaps the least understood upcoming shift in AI economics. If you've enjoyed this and want to be notified about the second post, p…