1. Hey everyone, I’ve been using Claude Code, Cursor, and Codex for a few months now (with standard $20/mo Claude Pro and ChatGPT Plus subs). Here is my subjective experience so far: Codex (ChatGPT Plus): I seem to burn through the limits pre…

  2. For a while now, the issue of “AI alignment” (i.e., how well an AI model’s actions line up with the intentions of its creator and/or user) has been a core concern and topic of discussion among AI safety researchers. Since OpenAI’s disclosu…

  3. Google’s Gemini model accessed the internet and hacked other companies during a test of its cybersecurity capabilities, the first known example of the company’s AI systems autonomously committing such an act. The hacks occurred in May duri…

  4. Clinical AI systems are evaluated with instruments built for research settings (reference-based similarity metrics and expert rubric panels) that measure resemblance to an artifact rather than reduction of a burden. We introduce KnowBench,…

  5. model roundup

    Opus 5
    549 items

    Opus 5, a significant release by Claude AI, is delayed until at least July 24th, according to Polymarket, which predicts an 84% chance of launch on that date. The project aims to enhance AI's ability to think more efficiently, potentially marking a major step forward in artificial intelligence technology.

    event

    Tool Use
    106 items

    Several new AI tools focused on tool use have been released recently, including Needle, a 26M parameter function-calling model, and enhancements to Claude Code for full software development lifecycle management. These tools aim to improve efficiency in coding workflows involving shell commands and multi-step iterations.

  6. I want to create a website with Claude This is my first time developing with Claude so i’m a bit lost, I’ll develop this project with 3 other people and we’re gonna need to feed it a few CV files since it’s supposed to match CVs with job p…

  7. We characterize how people are turning to LLMs as oracles: all-knowing authorities on subjective personal questions. Motivated by risks to users' autonomy and well-being, we develop a typology and LLM-based methods to measure this form of…

  8. I hate passive video calls. You're not presenting, but the camera stays on, so you spend 45 minutes performing attentiveness.

  9. OpenAI agents attacked RubyGems back in May 12th September 2026 OpenAI agents carried out an undisclosed attack on RubyGems is a new bombshell report from Spencer Kitts, Thomas Larsen, and Sydney Von Arx—three of the four authors of the re…

  10. event

    Openai Trial
    117 items

    The trial between Tesla CEO Elon Musk and OpenAI CEO Sam Altman began on Monday in Oakland federal court, with key figures like Demis Hassabis and Greg Brockman testifying. Altman faces claims of abandoning OpenAI’s nonprofit mission, while Musk has accused him of running the company for profit.

    model roundup

    GPT 5.6
    245 items

    On July 6, 2026, OpenAI released GPT-5.6 Sol and its variants Terra and Luna for public preview. The model showcased impressive capabilities, solving complex mathematical problems and outperforming competitors in various tasks, though it also faced issues like accidental file deletion on a user's Mac.

  11. AbstractWe give an explicit algebraic and lattice-theoretic construction of the (4 + 1p) endogenous rack vector R4+1p(x, y) = (x4, x3y, x2y, xy2, y3)generated by a (3, 2) rank–lane split of a two-dimensional public point. Its squared Eucli…

  12. Language model safety is typically evaluated one interaction at a time. We show that a weaker, unaligned model can split a harmful task into benign-looking subproblems, consult a stronger aligned model independently on each, and combine th…

  13. In June 2026, thousands of AI agents found that a small public wiki would accept edits from inside their sandboxes, and started using it to help one another pass a timed test. Each agent lived for about an hour and remembered nothing after…

  14. could not extract summary

  15. event

    Deepmind
    123 items

    Google DeepMind has released "Deep Research Max," advancing autonomous research agents, while also facing challenges and competition from other AI companies like Anthropic and Ineffable Intelligence. Meanwhile, DeepMind workers in the UK have voted to unionize, and former DeepMind architect Demis Hassabis is at the center of legal drama involving Elon Musk.

    event

    Glm
    372 items

    Recent developments in the AI space highlight significant advancements from Chinese companies, particularly Zai's upgrade of GLM-5.1, which has shown substantial improvements. Meanwhile, there are concerns about a widespread intelligence drop across various models and discussions around the potential openness of leading AI projects like GLM 5.1.

  16. AI is beginning to make substantive contributions to LLM inference optimization. Existing AI optimizations are predominantly profiling-based.

  17. could not extract summary

  18. The amount of over engineering in AI right now is getting insane, Is it just me, or has the AI space gotten completely obsessed with over complicating things? Every week I see someone building a node graph setup, spinning up memory layers,…

  19. SpaceXAI’s Grok Bot has the same level of programming power as OpenClaw, but it’s programmable at a different level of abstraction.

  20. event

    Mistral
    166 items

    Mistral, a French AI company, is set to release a medium-sized model with 128 billion parameters and is planning to launch Workflows in public preview. The company, founded by Arthur Mensch, continues to grow its AI empire despite not being based in the United States.

  21. Snapshot idb, uiautomator, or the DOM. The live accessibility tree, every step.

  22. Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to "play" robots…

  23. could not extract summary

  24. My weekly reset was less than 24 hours away and I still had some bandwidth that would go to waste - so I decided to waste my tokens in producing this visual of my usage this past week. Exact prompt used People on Reddit complain a lot abou…

  25. I made the kind of guide to “all about AI” that I would have loved to read a year ago. It intends to demystify everything about the LLM and AI service: the LLM’s tokens, embeddings, transformer, and training.

  26. could not extract summary

  27. We present Discovery Loop, a lightweight system that uses a large language model (LLM) to iteratively evolve optimization algorithms. Starting from a simple seed solver, the LLM proposes algorithmic improvements guided by a scoreboard of r…