1. I run Codex and Claude Code in Konsole and realized I don't fully know where all of my tokens are being used, so I started trying to trace them. I learned that the whole conversation gets sent again every turn, so every time a new message…

  2. 17th September 2026 - Link Blog How To Write With An LLM. Thomas Ptacek on using LLMs as copyeditors, not as writing assistants: Rule Number One: You may not use a single word an LLM suggests to you.

  3. Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit discriminatory conte…

  4. It feels like Opus and Fable are taking the bit in their teeth and just galloping away on things more often than they would even very recently. Not long ago it felt like they'd stay relatively within the bounds of the instructions I gave a…

  5. model roundup

    Opus 5
    550 items

    Opus 5, a significant release by Claude AI, is delayed until at least July 24th, according to Polymarket, which predicts an 84% chance of launch on that date. The project aims to enhance AI's ability to think more efficiently, potentially marking a major step forward in artificial intelligence technology.

    event

    Copilot
    694 items

    Microsoft is keeping its Copilot tool for Windows 11 but renaming it, while issues with rate limits and a security proxy have sparked concerns among users of GitHub Copilot. Meanwhile, Anthropic released a report on agentic coding trends, highlighting that developers use AI in about 60% of their work.

  6. could not extract summary

  7. could not extract summary

  8. Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the effectiveness of indivi…

  9. AIEi Paris (Sep 23-24) and AIE NYC (Oct 12-14) is >50% sold out, AIE CODE (Nov 10-12 in SF) and AIEi Shanghai (Nov 5-6) are next on deck before AIEi Sydney (Dec 7-8 alongside NeurIPS) closes the year! It’s very rare that a new startup laun…

  10. event

    Cowork
    867 items

    Issues with Claude Cowork have been reported, including errors and disruptions for some users on April 16, 2026. Additionally, Google has developed its own desktop Agent to compete with Cowork, while users continue to explore alternatives and troubleshoot bugs in the platform.

    model roundup

    GPT 6
    119 items

    On September 14, 2023, OpenAI launched GPT-6 Astra, claiming it represents a significant step towards advanced general intelligence and could revolutionize areas like cybersecurity and software engineering.

  11. This is across the bear type issue. This happens not just with Claude, but with Gemini, AntiGravity, GrokBot, ChatGPT.

  12. The scaling laws hold that a language model grows more capable with more parameters and more training data, and Mixture-of-Experts (MoE) architectures have ridden these laws to remarkable results, activating only a fraction of an enormous…

  13. Hello! Fellow teachers, In recent weeks I’ve been creating resources to teach english here in 🇲🇽.

  14. We last highlighted the pacing debate in July when Pacing the Frontier first emerged: And it seems that we’re in for round 2 as Dario, lead author on the original, wrote a rare personal blogpost to spell out how he sees pacing pan out spec…

  15. event

    Security
    687 items

    OpenAI has released GPT-5.4-Cyber for testing as part of its Trusted Access for Cyber Defense program, aiming to compete with Anthropic's Claude Mythos in the cybersecurity domain. Meanwhile, concerns are rising over the potential risks associated with advanced AI models like Mythos, prompting calls for improved defenses before wider releases.

    163 items

    Claude Opus 4.6, Anthropic's flagship model, saw its accuracy drop on the BridgeBench hallucination test from 83% to 68%, highlighting a significant regression in handling certain tasks. Meanwhile, biologists are revisiting cases of mushroom-induced hallucinations in China, suggesting ongoing research into natural causes of similar phenomena.

  16. Ternary Large Language Models (LLM) store every weight as one of three symbols $\{-1,0,+1\}$, so the cost of a ternary model is conventionally referenced to the information-theoretic $\log_2 3 \approx 1.585$ bits per weight. The prevailing…

  17. Your Agent Aced the Task. Will It Do It Again?

  18. We introduce and release ScienceBuddy, an interactive scientific research workspace that brings continually improving scientific agents into researchers' everyday workflows. ScienceBuddy supports researchers in carrying out scientific task…

  19. AIUC first got our attention with the NFDG backing, and have just announced a $40M series A today, with the most impressive industry advisor list we may have ever seen for an early startup behind AIUC-1, their agent standard backed by real…

  20. model roundup

    DeepSeek 4.1
    21 items

    DeepSeek has released v4.1 Flash, an updated version of their API with enhanced features like multimodal support and improved speed, set to expire on September 10, 2026. The internal beta testing is now available, with input costing $0.22 and output $0.66 per 1M tokens.

  21. I hate passive video calls. You're not presenting, but the camera stays on, so you spend 45 minutes performing attentiveness.

  22. Capable open-weight models make local coding and reasoning attractive, but their context and execution state strain laptop memory. We present JustFit, an MLX-based inference runtime that combines KVExec for compressed KV execution, PhaseSw…

  23. The amount of over engineering in AI right now is getting insane, Is it just me, or has the AI space gotten completely obsessed with over complicating things? Every week I see someone building a node graph setup, spinning up memory layers,…

  24. Software development follows an implementation-verification loop in which developers or agents iteratively revise an implementation until an evaluator, such as a test suite, accepts it. The evaluator checks the implementation against a set…

  25. https://github.com/HalfLucid/FileInteractionMCP I recently noticed Claude CLI making a lot of bash and perl calls for file edits, so I created some expanded file tooling with Fable to make it more secure and efficient. MIT license so feel…

  26. In response to a new European Union law, AI platforms are implementing new schemes for watermarking the content they generate. Anthropic recently disclosed its future Claude models will use SynthID-Text, an approach Google created and rele…

  27. could not extract summary