event

Arc Agi

54 items · started 2026-04-14 · ongoing (last activity 2026-09-14)

  1. GPT-6 Astra solved twice as many Baba Is You levels as Claude Fable 5.1. It fits a pattern: ARC-AGI-3, the 3D game Portal, MazeBench, and Przemysław “Psyho” Dębiak’s obscure puzzles.

  2. I'm feeling extremely depressed and anxious lately. After seeing Astra saturating ARC-AGI-3 and blowing through many other benchmarks, it's hard to believe we won't have a strong form of AGI by the end of this year, or the next one at the…

  3. GPT-6 Astra OpenAI·Sep 2, 2026·12 harness configurations GPT-6 Astra's best observed ARC-AGI-3 Semi-Private result with the Standard harnessStandard harness enables a model to carry forward notes it chooses to keep with it throughout the e…

  4. GPT-6 Astra, the first Stargate and lightly looped supermodel from OpenAI, launched today, cleanly beating Fable 5.1 on many metrics including completely saturating the hardest versions of FrontierMath (97.6%) and ARC-AGI-3 (99.9%). Lots o…

  5. GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a…

  6. could not extract summary

  7. Claude Fable 5.1 Anthropic·Sep 1, 2026·5 reasoning variants At max effort, Claude Fable 5.1 scores 97.5% on ARC-AGI-1 Semi-Private at $1.40 per task and 90.0% on ARC-AGI-2 Semi-Private at $4.49 per task. ARC-AGI 2 leaderboard Claude Fable…

  8. We solve 20% of ARC-AGI-1 with no pretraining by minimizing the description length during inference time, with the only data involved being the target puzzle itself.

  9. 44% on ARC-AGI-1 in 67 cents I trained a small transformer from scratch in 1.5hrs on a 5090 Beats many LLMs, and scores the same as TRM/HRM This is an upgrade to my previous model Faster, better, cheaper and still open source. Also gets 7%…

  10. AI’s most fervent and optimistic promoters promise a future where AI is innovating its way out of society’s biggest problems. AI systems will conduct scientific research, discover new drugs and materials, run engineering projects, and auto…

  11. SwarmOS pushed GPT-5.6-Sol from a 13.3% baseline to 100% RHAE on ARC-AGI-3 Public, showing how orchestration can multiply long-horizon agent capability.

  12. This is my attempt to explain what the ARC-AGI-3 interactive reasoning benchmark is testing, how the strongest public solutions work, and why I think their dynamic learning loops and harness design matter if you’re building real agents. Co…

  13. 100% on the ARC-AGI-3 public set August 25, 20265 min read Today we're releasing the ARC-AGI-Arcade. An open-source playground for agents to compete on ARC-AGI-3 by learning from each other.

  14. A frontier language model is only one component of an AI agent. The surrounding agent system—often called a harness—determines how the model receives context, uses tools, maintains state, responds to feedback, recovers from failure, and su…

  15. Gemini 3.7 Flash Google·Aug 13, 2026·3 reasoning variants At high effort, Gemini 3.7 Flash scores 95.5% on ARC-AGI-1 Semi-Private at $0.12 per task and 84.6% on ARC-AGI-2 Semi-Private at $0.25 per task. Related models: Gemini 3.6 Flash ARC…

  16. A 129-line skill let a general-purpose coding agent finish all 25 public ARC-AGI-3 games at 100.00 RHAE in 7,645 actions. Replay every action and every prediction.

  17. Abstract A 150M-parameter reasoning model using recurrent latent reasoning and in-context learning achieves a new cost-accuracy frontier on ARC-AGI-1. We introduce BDH-CQ, a reasoning model that combines in-context learning with recurrent…

  18. We introduce ARC-AGI-3, an interactive benchmark for studying agentic intelligence through novel, abstract, turn-based environments in which agents must explore, infer goals, build internal models of environment dynamics, and plan effectiv…

  19. could not extract summary

  20. The Harness Is the Product OpenAI recently reported that GPT‑5.6 Sol scored 13.3% on ARC-AGI-3 using the official harness. With retained reasoning and context compaction enabled, the same model scored 38.3%, while producing six times fewer…

  21. I've been trying to understand the difference between Fable 5 and Claude Opus 5 beyond the benchmark numbers, and I'm curious what other people have observed in real-world use. A few months ago someone explained the idea behind the ARC Pri…

  22. could not extract summary

  23. Anthropic shipped Opus 5 on July 24 at $5/$25 per million tokens (input/output). The benchmarks hold up: 3x ARC-AGI 3 score vs the next best model, beats Fable 5 on OSWorld 2.0 computer use at ~1/3 the cost, +10.2pp on organic chemistry an…

  24. Opus 5 reports 30% on ARC-AGI-3, ~4× the previous best model, ~20× its predecessor Opus 4.8. We tested it on Witness, our held-out suite of ARC-AGI-3-style interactive puzzle games.

  25. It just went live. The headline numbers: Same price as Opus 4.8 ($5/$25 per M) but new SOTA on Frontier-Bench and GDPval-AA ARC-AGI 3: 3x the next-best model OSWorld 2.0: beats Fable 5's best score at ~1/3 the cost Now the default on Max a…

  26. As foundational models approach Artificial Superintelligence (ASI), standard intelligence evaluations like the Massive Multitask Language Understanding (MMLU), Grade School Math 8K (GSM8K), and the Abstraction and Reasoning Corpus (ARC-AGI…

  27. Our previous ARC-AGI-3 agent bundled executable world modeling, scheduled simplification, and exact replay verification, leaving unclear which idea accounted for its performance. We address this attribution question with four nested Codex-…

  28. Inkling from @thinkymachines on ARC-AGI (Verified) - ARC-AGI-2: 36.5%, $0.64/task - ARC-AGI-1: 79.5%, $0.30/task As of today, Inkling is the highest-scoring open-weight model evaluated by ARC Prize on both ARC-AGI-1 and ARC-AGI-2. https://…

  29. Checking your browser Starting verification… This takes a couple of seconds and happens automatically. If you can't pass the test, whitelist your extensions on this website, update your browser and use a modern web browser (Firefox/Chrome/…

  30. Two competing perspectives on fluid intelligence (gf) measures propose that performance is primarily constrained either by working memory capacity or by the ability to induce novel relations. The first perspective is currently dominant in…

  31. We present ARCANA, a collaborative multi agent framework for solving ARC AGI 2 tasks under strict test time and hardware constraints. ARCANA decomposes each task into iterative perception, hypothesis generation, symbolic execution, and ref…

  32. Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionary search, exhaustive sampling, extended chain-of-thought), or benchmark-specific training…

  33. Considered by many to be the best benchmark for abstract reasoning, it's not saturated, and yet it's not included in several benchmarks, let alone this one. I understand they said they were going to do it.

  34. Large language models can produce fluent, internally coherent reasoning traces for abstract reasoning tasks while still being confidently wrong - making selection among candidates, not just generation, the central challenge. I present a so…

  35. TLDR: I built a model that has only one parameter and gets 100% on ARC-AGI-2, the million-dollar benchmark that pushes reasoning models to their limits. Using chaos theory and some deliberate cheating, I crammed every answer into a single…

  36. Anthropic Opus 4.8 is new SOTA on ARC-AGI-3 Score: 1.5%, ~$10K ARC-AGI-3 analysis notes: * Opus 4.8 read the environment an abstraction *above* Opus 4.7, as objects & systems, not pictures * Opus 4.8 succeeded on early levels, but still co…

  37. Ring-2.6-1T made me think about failure placement more than headline strength. It’s a trillion-parameter reasoning model for agent workflows with high and xhigh reasoning-effort modes.

  38. Ring-2.6-1T made me think less about “is this good?” and more about routing. The public profile looks like something I'd at least test for harder agent steps: PinchBench 87.60, AIME 26 95.83, GPQA Diamond 88.27, Tau2-Bench Telecom 95.32, b…

  39. 🔓 13 "Impossible" ARC-AGI-2 Tasks — All Solved These 13 ARC-AGI-2 evaluation tasks have never been solved by any AI system — not GPT-4, not Claude, not Gemini, not NVARC, not MindsAI, not any Kaggle submission. They have a 0% AI solve rate…

  40. ARC-AGI Community Leaderboard ARC-AGI has gained significant popularity over the past two years, and we've been overwhelmed by the number of researchers and builders who want to showcase their work to the community. The ARC-AGI Community L…

  41. The 30-Second Verdict Best for science & reasoning: Gemini 3.1 Pro — leads GPQA Diamond (94.3%) and ARC-AGI-2 (77.1%). Best for coding: ChatGPT (GPT-5.5) — 88.7% on SWE-Bench Verified.

  42. AGI, frontier science, maniacal metaphysics, decentralizationist politics, life and consciousness extension and expansion, psi and psychedelics and etc. etc.

  43. I'm not sure too many people care about the ARC-AGI-2 competition anymore, but still...I thought some might find this interesting. They're running it one last time this year.

  44. GAME ar2511 bp3513 cd8211 dc2210 g50t10 ka5911 lf5210 ls201 m0r029 r11l1 re863 su1510 vc3311 wa3010 VARIANT A0-replay-m0r01 A1-fresh-m0r03 A10-unseen-game3 A11-smaller-model3 A2-no-theory3 A3-no-journal3 A4-no-scratchpads3 A5-no-code-writi…

  45. I was reading the comments to this post and the overall opinion seemed to be that harness makes little/no difference for ARC-AGI-3. Turns out, it makes a huge difference: Hill-climbing ARC-AGI-3 TLDR: if you save game logs - taken actions,…

  46. - GPT-5.5: 0.43% - Opus 4.7: 0.18% ARC-AGI-3 is no joke. I can’t wait to see which models finally crack.

  47. Many people have pointed out that ChatGPT 5.5 appears to be twice as expensive as 5.4 based on API pricing, which makes it look pricier than Opus 4.7. But the comparison is not that simple.

  48. could not extract summary

  49. A few weeks ago, ARC-AGI 3 was released. For those unfamiliar, it’s a benchmark designed to study agentic intelligence through interactive environments.

← all threads