How To Write With An LLM (simonwillison.net)
17th September 2026 - Link Blog How To Write With An LLM. Thomas Ptacek on using LLMs as copyeditors, not as writing assistants: Rule Number One: You may not use a single word an LLM suggests to you.
- How to Write with an LLM (sockpuppet.org via hn)
Very impressive take by Claude on watermarking (www.reddit.com via reddit)
❯ If I provide all inference and I ask you to serialize it, then you stamp it with your imprimatur, you are saying that the value is in serialization and not the collections of facts used to infer Yes, and the inversion is sharper than tha…
Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the effectiveness of indivi…
Link: https://x.com/Lon/status/2101034933284417614 A 65-day analysis of 43,000+ Claude Code invocations found that 39% of Fable 5 calls get zero thinking tokens and the median invocation gets just 123 — while benchmarks use 16K-128K. The m…
-
868 items
event
CoworkIssues with Claude Cowork have been reported, including errors and disruptions for some users on April 16, 2026. Additionally, Google has developed its own desktop Agent to compete with Cowork, while users continue to explore alternatives and troubleshoot bugs in the platform.
- 5m Usage shot up to 100% under "Other"
- 8h New symbol next to conversation title?
- 8h I built a Mac panel that shows my Claude Code and Cowork sessions, which ones need me and which are done
- 15h I love Cowork and Chat, but I love them SEPARATE and UNEQUAL, does anyone else feel the same?
- 17h How can I disable Cowork again?
551 itemsmodel roundup
Opus 5Opus 5, a significant release by Claude AI, is delayed until at least July 24th, according to Polymarket, which predicts an 84% chance of launch on that date. The project aims to enhance AI's ability to think more efficiently, potentially marking a major step forward in artificial intelligence technology.
- 9m You want a disheartening experience? Have Claude walk you through a Windows bloatware and privacy checkup.
- 4h gave sonnet 5, opus 5, astra and fable 5.1 the same "lighthouse at night" svg prompt. then fable did it again today and… what?
- 4h Told Claude Code to build a YouTube plugin, it decided on its own to Rickroll me
- 12h What do you feel when you talk to agents?
- 12h Asked Claude to waste my remaining usage before weekly reset. Very satisfied with the result.
How Cooley is accelerating IPO work with ChatGPT (openai.com)
could not extract summary
Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit discriminatory conte…
OpenAI researcher on AI communicating across air-gaps via thermal side-channels [video] (www.youtube.com via hn)
About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
AIUC first got our attention with the NFDG backing, and have just announced a $40M series A today, with the most impressive industry advisor list we may have ever seen for an early startup behind AIUC-1, their agent standard backed by real…
-
374 items
event
GlmRecent developments in the AI space highlight significant advancements from Chinese companies, particularly Zai's upgrade of GLM-5.1, which has shown substantial improvements. Meanwhile, there are concerns about a widespread intelligence drop across various models and discussions around the potential openness of leading AI projects like GLM 5.1.
- 16m Chinese coding AI and Astra is dogshit compared to Claude. ABSOLUTE middle finger to the haters
- 1h [Help] Token Max Math Help
- 1d Claude should offer first-party support for open-weight models Anthropic hosts via partnerships, hear me out.
- 1d GLM-5.3-FlashX: Delivering inference speeds of 200 tokens/s
- 1d ZCode, the GLM coding agent, silently uploads your Git history
22 itemsmodel roundup
DeepSeek 4.1DeepSeek has released v4.1 Flash, an updated version of their API with enhanced features like multimodal support and improved speed, set to expire on September 10, 2026. The internal beta testing is now available, with input costing $0.22 and output $0.66 per 1M tokens.
- 22h Ask HN: DeepSeek v4.1 Flash on Local hardware, what tok/s do you see?
- 1d DeepSeek v4.1 Flash avg 102 tps on 4x RTX6000 pro max-q, 2.1x up from v4-flash
- 3d Setting Up Pi with DeepSeek v4.1 Flash on OpenRouter
- 3d DeepSeek v4.1 Flash Is Now Our Best Hacking Model
- 4d SGLang and Miles Add Day-0 Support for DeepSeek-v4.1
Your Agent Aced the Task. Will It Do It Again? (huggingface.co)
Your Agent Aced the Task. Will It Do It Again?
Show HN: The Smallest LLM (gist.github.com via hn)
Created September 20, 2026 04:04 - - Save skorotkiewicz/dedc3b5a857be7d0f2b378334721713c to your computer and use it in GitHub Desktop. the smallest LLM This file contains hidden or bidirectional Unicode text that may be interpreted or com…
- Show HN: WiFi-LLM (github.com via hn)
- Show HN: LLM for Dummies (ronreiter.github.io via hn)
- Show HN: When the LLM Accidentally (news.ycombinator.com)
The scaling laws hold that a language model grows more capable with more parameters and more training data, and Mixture-of-Experts (MoE) architectures have ridden these laws to remarkable results, activating only a fraction of an enormous…
Ternary Large Language Models (LLM) store every weight as one of three symbols $\{-1,0,+1\}$, so the cost of a ternary model is conventionally referenced to the information-theoretic $\log_2 3 \approx 1.585$ bits per weight. The prevailing…
-
119 items
model roundup
GPT 6On September 14, 2023, OpenAI launched GPT-6 Astra, claiming it represents a significant step towards advanced general intelligence and could revolutionize areas like cybersecurity and software engineering.
688 itemsevent
SecurityOpenAI has released GPT-5.4-Cyber for testing as part of its Trusted Access for Cyber Defense program, aiming to compete with Anthropic's Claude Mythos in the cybersecurity domain. Meanwhile, concerns are rising over the potential risks associated with advanced AI models like Mythos, prompting calls for improved defenses before wider releases.
- 16h When "Review" Becomes Permission: A Prompt Injection Lab
- 17h Liability and Prompt Injection
- 1d How do you handle tool sprawl and untrusted scripts when building with Claude Code and custom agents?
- 1d Show HN: An OSS Python dependency scanner for exploited, unmaintained packages
- 1d PhantomFix: A fake bug to Sentry Seer gets a coding agent to run attacker code
Show HN: Migrating from Claude Code CLI to Claude Desktop (github.com via hn)
Claude Profiles Run several isolated Claude Desktop accounts side by side, and move Claude Code session history between them. Two parts: claude-profiles, a CLI that creates profiles, launches them, and moves session history into them.
- Migrating from Claude Code (www.reddit.com via reddit)
- Claude in VS Code vs Desktop (www.reddit.com via reddit)
- Migrating Claude Code Desktop from old Laptop to new (www.reddit.com via reddit)
+5 more
- Show HN: Claude Code desktop alternative for macOS (skeezo.dev via hn)
- Is Claude Code on desktop still worse than CLI? (www.reddit.com via reddit)
- Claude Code Desktop vs Claude CLI (www.reddit.com via reddit)
- Show HN: Makememe – a meme CLI for your Claude Code (github.com via hn)
- Claude Code Desktop vs CLI (www.reddit.com)
We introduce and release ScienceBuddy, an interactive scientific research workspace that brings continually improving scientific agents into researchers' everyday workflows. ScienceBuddy supports researchers in carrying out scientific task…
- Recursive Self Improvement for Coding Agents (cline.ghost.io via hn)
AIEi Paris (Sep 23-24) and AIE NYC (Oct 12-14) is >50% sold out, AIE CODE (Nov 10-12 in SF) and AIEi Shanghai (Nov 5-6) are next on deck before AIEi Sydney (Dec 7-8 alongside NeurIPS) closes the year! It’s very rare that a new startup laun…
We last highlighted the pacing debate in July when Pacing the Frontier first emerged: And it seems that we’re in for round 2 as Dario, lead author on the original, wrote a rare personal blogpost to spell out how he sees pacing pan out spec…
-
163 items
event
HallucinationClaude Opus 4.6, Anthropic's flagship model, saw its accuracy drop on the BridgeBench hallucination test from 83% to 68%, highlighting a significant regression in handling certain tasks. Meanwhile, biologists are revisiting cases of mushroom-induced hallucinations in China, suggesting ongoing research into natural causes of similar phenomena.
318 itemsevent
HaikuClaude is introducing an advisor strategy with Haiku and Opus, allowing users to consult Opus for mid-task decisions while Haiku handles execution. A nursing student built a large pharmaceutical database using Claude Haiku on the side, highlighting its utility in complex tasks.
Capable open-weight models make local coding and reasoning attractive, but their context and execution state strain laptop memory. We present JustFit, an MLX-based inference runtime that combines KVExec for compressed KV execution, PhaseSw…
TIFU by forgetting to use Fable (www.reddit.comhttps)
I was trying to save Fable for the important projects and used too much Opus, now I’m in this situation…
- Why can’t I use Fable 5? (www.reddit.com via reddit)
Why AI is never good with vector designs? (www.reddit.comhttps)
This is across the bear type issue. This happens not just with Claude, but with Gemini, AntiGravity, GrokBot, ChatGPT.
- This is never good… (www.reddit.comhttps)
LLMs respond differently to harmful prompts when AI watermarking is used (arstechnica.com)
In response to a new European Union law, AI platforms are implementing new schemes for watermarking the content they generate. Anthropic recently disclosed its future Claude models will use SynthID-Text, an approach Google created and rele…
Software development follows an implementation-verification loop in which developers or agents iteratively revise an implementation until an evaluator, such as a test suite, accepts it. The evaluator checks the implementation against a set…
could not extract summary