$100 Claude Max & $100 Codex or $200 Claude Max (www.reddit.com via reddit)
Curious. If you had a budget of $200/mo.
- Cursor vs claude vs codex for 200$ (www.reddit.com via reddit)
- How's codex with Plus tier compared to 100$ Claude max? (www.reddit.com via reddit)
- Thinking about moving from $200 Codex to Claude Max (www.reddit.com via reddit)
How To Write With An LLM (simonwillison.net)
17th September 2026 - Link Blog How To Write With An LLM. Thomas Ptacek on using LLMs as copyeditors, not as writing assistants: Rule Number One: You may not use a single word an LLM suggests to you.
- How to Write with an LLM (sockpuppet.org via hn)
Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit discriminatory conte…
AIUC first got our attention with the NFDG backing, and have just announced a $40M series A today, with the most impressive industry advisor list we may have ever seen for an early startup behind AIUC-1, their agent standard backed by real…
-
694 items
event
CopilotMicrosoft is keeping its Copilot tool for Windows 11 but renaming it, while issues with rate limits and a security proxy have sparked concerns among users of GitHub Copilot. Meanwhile, Anthropic released a report on agentic coding trends, highlighting that developers use AI in about 60% of their work.
- 2h Show HN: Forexfin – Trading calculators, alerts and a journal in one place
- 8h Built a Herdr plugin that finds any Claude Code, Codex, Gemini or OpenCode session by a word you remember and resumes it
- 11h AI coding agents' 0-click RCE flaw could hand attackers keys to the kingdom
- 19h Microsoft agentically ports Copilot runtime to Rust for $120K
- 1d codemap - packs your repository into focused, AI-ready context
867 itemsevent
CoworkIssues with Claude Cowork have been reported, including errors and disruptions for some users on April 16, 2026. Additionally, Google has developed its own desktop Agent to compete with Cowork, while users continue to explore alternatives and troubleshoot bugs in the platform.
- 4h New symbol next to conversation title?
- 4h I built a Mac panel that shows my Claude Code and Cowork sessions, which ones need me and which are done
- 11h I love Cowork and Chat, but I love them SEPARATE and UNEQUAL, does anyone else feel the same?
- 13h How can I disable Cowork again?
- 22h Exporting Cowork chats? No? Really, Anthropic?
Claude Code users: Sonnet vs Opus vs Fable 5.1 — which are you using? (www.reddit.com via reddit)
I've been trying to figure out which model works best for real-world coding with Claude Code. Sonnet — fast and efficient Opus — better for complex reasoning?
- Using Opus 4.6 in Claude Code (plugin) for VS Code (www.reddit.com)
Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the effectiveness of indivi…
How Cooley is accelerating IPO work with ChatGPT (openai.com)
could not extract summary
Show HN: Built a local guardrail layer for Claude Code (github.com via hn)
I built Mati to give me more control over what Claude Code and Codex are allowed to do. Many important product constraints live in an engineer's or product person's head.
- I built a database guardrail for Claude Code (www.reddit.com via reddit)
-
119 items
model roundup
GPT 6On September 14, 2023, OpenAI launched GPT-6 Astra, claiming it represents a significant step towards advanced general intelligence and could revolutionize areas like cybersecurity and software engineering.
687 itemsevent
SecurityOpenAI has released GPT-5.4-Cyber for testing as part of its Trusted Access for Cyber Defense program, aiming to compete with Anthropic's Claude Mythos in the cybersecurity domain. Meanwhile, concerns are rising over the potential risks associated with advanced AI models like Mythos, prompting calls for improved defenses before wider releases.
- 12h When "Review" Becomes Permission: A Prompt Injection Lab
- 13h Liability and Prompt Injection
- 21h How do you handle tool sprawl and untrusted scripts when building with Claude Code and custom agents?
- 1d Show HN: An OSS Python dependency scanner for exploited, unmaintained packages
- 1d PhantomFix: A fake bug to Sentry Seer gets a coding agent to run attacker code
The scaling laws hold that a language model grows more capable with more parameters and more training data, and Mixture-of-Experts (MoE) architectures have ridden these laws to remarkable results, activating only a fraction of an enormous…
Your Agent Aced the Task. Will It Do It Again? (huggingface.co)
Your Agent Aced the Task. Will It Do It Again?
Built a free better file manipulation MCP for better efficiency and security (www.reddit.com via reddit)
https://github.com/HalfLucid/FileInteractionMCP I recently noticed Claude CLI making a lot of bash and perl calls for file edits, so I created some expanded file tooling with Fable to make it more secure and efficient. MIT license so feel…
Ternary Large Language Models (LLM) store every weight as one of three symbols $\{-1,0,+1\}$, so the cost of a ternary model is conventionally referenced to the information-theoretic $\log_2 3 \approx 1.585$ bits per weight. The prevailing…
-
163 items
event
HallucinationClaude Opus 4.6, Anthropic's flagship model, saw its accuracy drop on the BridgeBench hallucination test from 83% to 68%, highlighting a significant regression in handling certain tasks. Meanwhile, biologists are revisiting cases of mushroom-induced hallucinations in China, suggesting ongoing research into natural causes of similar phenomena.
318 itemsevent
HaikuClaude is introducing an advisor strategy with Haiku and Opus, allowing users to consult Opus for mid-task decisions while Haiku handles execution. A nursing student built a large pharmaceutical database using Claude Haiku on the side, highlighting its utility in complex tasks.
We introduce and release ScienceBuddy, an interactive scientific research workspace that brings continually improving scientific agents into researchers' everyday workflows. ScienceBuddy supports researchers in carrying out scientific task…
- Recursive Self Improvement for Coding Agents (cline.ghost.io via hn)
We last highlighted the pacing debate in July when Pacing the Frontier first emerged: And it seems that we’re in for round 2 as Dario, lead author on the original, wrote a rare personal blogpost to spell out how he sees pacing pan out spec…
Leg Type leg claude, leg codex, leg agy or leg grok instead of the bare command. You get the same interactive agent; Leg opens a board next to it, watches the usage limit, keeps a handoff bundle current, and when the limit hits it starts t…
AIEi Paris (Sep 23-24) and AIE NYC (Oct 12-14) is >50% sold out, AIE CODE (Nov 10-12 in SF) and AIEi Shanghai (Nov 5-6) are next on deck before AIEi Sydney (Dec 7-8 alongside NeurIPS) closes the year! It’s very rare that a new startup laun…
-
21 items
model roundup
DeepSeek 4.1DeepSeek has released v4.1 Flash, an updated version of their API with enhanced features like multimodal support and improved speed, set to expire on September 10, 2026. The internal beta testing is now available, with input costing $0.22 and output $0.66 per 1M tokens.
- 18h Ask HN: DeepSeek v4.1 Flash on Local hardware, what tok/s do you see?
- 1d DeepSeek v4.1 Flash avg 102 tps on 4x RTX6000 pro max-q, 2.1x up from v4-flash
- 3d Setting Up Pi with DeepSeek v4.1 Flash on OpenRouter
- 3d DeepSeek v4.1 Flash Is Now Our Best Hacking Model
- 3d Ask HN: What's the most economical approach to the most tokens?
It feels like Opus and Fable are taking the bit in their teeth and just galloping away on things more often than they would even very recently. Not long ago it felt like they'd stay relatively within the bounds of the instructions I gave a…
Capable open-weight models make local coding and reasoning attractive, but their context and execution state strain laptop memory. We present JustFit, an MLX-based inference runtime that combines KVExec for compressed KV execution, PhaseSw…
LLMs respond differently to harmful prompts when AI watermarking is used (arstechnica.com)
In response to a new European Union law, AI platforms are implementing new schemes for watermarking the content they generate. Anthropic recently disclosed its future Claude models will use SynthID-Text, an approach Google created and rele…
Software development follows an implementation-verification loop in which developers or agents iteratively revise an implementation until an evaluator, such as a test suite, accepts it. The evaluator checks the implementation against a set…
hi to all the readers this post is for my recent opensource project called ENZO now answering what is enzo so enzo is an opensource platform where i clubbed all the free available api for anyone use under one hood with more than 2000 model…
could not extract summary
Large language models sometimes behave in ways resembling human emotional responses, and recent work has identified internal representations that may explain this. We ask whether LLMs represent pain distinctly from fear, sadness, and gener…