GPT-6 Astra solves puzzles (quesma.com via hn)
event
Arc Agi
-
GPT-6 Astra solved twice as many Baba Is You levels as Claude Fable 5.1. It fits a pattern: ARC-AGI-3, the 3D game Portal, MazeBench, and Przemysław “Psyho” Dębiak’s obscure puzzles.
-
Ask HN: Anyone else feeling uneasy with the latest developments in AI? (news.ycombinator.com)
I'm feeling extremely depressed and anxious lately. After seeing Astra saturating ARC-AGI-3 and blowing through many other benchmarks, it's hard to believe we won't have a strong form of AGI by the end of this year, or the next one at the…
-
OpenAI's GPT-6 Astra on ARC-AGI-3 Scored 99.95% with Provider Adapter Harness (arcprize.org via hn)
GPT-6 Astra OpenAI·Sep 2, 2026·12 harness configurations GPT-6 Astra's best observed ARC-AGI-3 Semi-Private result with the Standard harnessStandard harness enables a model to carry forward notes it chooses to keep with it throughout the e…
-
GPT-6 Astra: an automated AI Engineer you can hire for <$6 an hour (www.latent.space via hn)
GPT-6 Astra, the first Stargate and lightly looped supermodel from OpenAI, launched today, cleanly beating Fable 5.1 on many metrics including completely saturating the hardest versions of FrontierMath (97.6%) and ARC-AGI-3 (99.9%). Lots o…
-
GPT-6 Astra represents a step-function change in interactive reasoning (twitter.com via hn)
GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a…
-
Astra is at 98.6% score on ARC-AGI-3 (cdn.thenewstack.io via hn)
could not extract summary
- GPT-6 Astra Achieves SOTA on ARC-AGI (twitter.com)
-
Claude Fable 5.1 results on ARC-AGI (arcprize.org via hn)
Claude Fable 5.1 Anthropic·Sep 1, 2026·5 reasoning variants At max effort, Claude Fable 5.1 scores 97.5% on ARC-AGI-1 Semi-Private at $1.40 per task and 90.0% on ARC-AGI-2 Semi-Private at $4.49 per task. ARC-AGI 2 leaderboard Claude Fable…
-
ARC-AGI Without Pretraining (2025) (iliao2345.github.io via hn)
We solve 20% of ARC-AGI-1 with no pretraining by minimizing the description length during inference time, with the only data involved being the target puzzle itself.
-
44% on ARC-AGI-1 in 67 cents (mvakde.github.io via hn)
44% on ARC-AGI-1 in 67 cents I trained a small transformer from scratch in 1.5hrs on a 5090 Beats many LLMs, and scores the same as TRM/HRM This is an upgrade to my previous model Faster, better, cheaper and still open source. Also gets 7%…
-
OSS harness took Claude Opus 5 from 30% to 99.95% on ARC-AGI-3 (twitter.com via hn)
AI’s most fervent and optimistic promoters promise a future where AI is innovating its way out of society’s biggest problems. AI systems will conduct scientific research, discover new drugs and materials, run engineering projects, and auto…
-
Pushing GPT-5.6 Luna from 0% to 56% on ARC-AGI-3 Public (int21.ai via hn)
SwarmOS pushed GPT-5.6-Sol from a 13.3% baseline to 100% RHAE on ARC-AGI-3 Public, showing how orchestration can multiply long-horizon agent capability.
-
ARC-AGI-3: A game with no instructions (genaipod.substack.com via hn)
This is my attempt to explain what the ARC-AGI-3 interactive reasoning benchmark is testing, how the strongest public solutions work, and why I think their dynamic learning loops and harness design matter if you’re building real agents. Co…
-
100.00 RHAE on the ARC-AGI-3 public set (www.agno.com via hn)
100% on the ARC-AGI-3 public set August 25, 20265 min read Today we're releasing the ARC-AGI-Arcade. An open-source playground for agents to compete on ARC-AGI-3 by learning from each other.
-
Nvidia AVO achieves 100% in ARC-AGI-3 (developer.nvidia.com via hn)
A frontier language model is only one component of an AI agent. The surrounding agent system—often called a harness—determines how the model receives context, uses tools, maintains state, responds to feedback, recovers from failure, and su…
-
Gemini 3.7 Flash scores on ARC-AGI (arcprize.org via hn)
Gemini 3.7 Flash Google·Aug 13, 2026·3 reasoning variants At high effort, Gemini 3.7 Flash scores 95.5% on ARC-AGI-1 Semi-Private at $0.12 per task and 84.6% on ARC-AGI-2 Semi-Private at $0.25 per task. Related models: Gemini 3.6 Flash ARC…
-
100% RHAE on ARC-AGI-3 public with Claude Code, Opus 5, and one skill (arc-skill.vercel.app via hn)
A 129-line skill let a general-purpose coding agent finish all 25 public ARC-AGI-3 games at 100.00 RHAE in 7,645 actions. Replay every action and every prediction.
-
150M-parameter reasoning model sets new cost-accuracy frontier on ARC-AGI-1 (huggingface.co via hn)
Abstract A 150M-parameter reasoning model using recurrent latent reasoning and in-context learning achieves a new cost-accuracy frontier on ARC-AGI-1. We introduce BDH-CQ, a reasoning model that combines in-context learning with recurrent…
-
ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence (arxiv.org via hn)
We introduce ARC-AGI-3, an interactive benchmark for studying agentic intelligence through novel, abstract, turn-based environments in which agents must explore, infer goals, build internal models of environment dynamics, and plan effectiv…
-
could not extract summary
-
LLM Works. Your Product Probably Doesn't (saito.ai via hn)
The Harness Is the Product OpenAI recently reported that GPT‑5.6 Sol scored 13.3% on ARC-AGI-3 using the official harness. With retained reasoning and context compaction enabled, the same model scored 38.3%, while producing six times fewer…
-
Fable 5 vs Opus 5 according to ARC PRIZE (www.reddit.com via reddit)
I've been trying to understand the difference between Fable 5 and Claude Opus 5 beyond the benchmark numbers, and I'm curious what other people have observed in real-world use. A few months ago someone explained the idea behind the ARC Pri…
-
Enabling two settings tripled our scores on the ARC-AGI-3 benchmark (openai.com via hn)
could not extract summary
-
Anthropic shipped Opus 5 on July 24 at $5/$25 per million tokens (input/output). The benchmarks hold up: 3x ARC-AGI 3 score vs the next best model, beats Fable 5 on OSWorld 2.0 computer use at ~1/3 the cost, +10.2pp on organic chemistry an…
-
Opus 5 ARC-AGI-3 likely benchmaxxed (xcancel.com via hn)
Opus 5 reports 30% on ARC-AGI-3, ~4× the previous best model, ~20× its predecessor Opus 4.8. We tested it on Witness, our held-out suite of ARC-AGI-3-style interactive puzzle games.
-
It just went live. The headline numbers: Same price as Opus 4.8 ($5/$25 per M) but new SOTA on Frontier-Bench and GDPval-AA ARC-AGI 3: 3x the next-best model OSWorld 2.0: beats Fable 5's best score at ~1/3 the cost Now the default on Max a…
-
Beyond Arc and GSM: Langford Coverage as a Benchmark for ASI (zenodo.org via hn)
As foundational models approach Artificial Superintelligence (ASI), standard intelligence evaluations like the Massive Multitask Language Understanding (MMLU), Grade School Math 8K (GSM8K), and the Abstraction and Reasoning Corpus (ARC-AGI…
-
Our previous ARC-AGI-3 agent bundled executable world modeling, scheduled simplification, and exact replay verification, leaving unclear which idea accounted for its performance. We address this attribution question with four nested Codex-…
-
Inkling from @thinkymachines on ARC-AGI (Verified) - ARC-AGI-2: 36.5%, $0.64/task - ARC-AGI-1: 79.5%, $0.30/task As of today, Inkling is the highest-scoring open-weight model evaluated by ARC Prize on both ARC-AGI-1 and ARC-AGI-2. https://…
-
99% on ARC-AGI 3 (public eval, with harness) (xcancel.com via hn)
Checking your browser Starting verification… This takes a couple of seconds and happens automatically. If you can't pass the test, whitelist your extensions on this website, update your browser and use a modern web browser (Firefox/Chrome/…
-
Two competing perspectives on fluid intelligence (gf) measures propose that performance is primarily constrained either by working memory capacity or by the ability to induce novel relations. The first perspective is currently dominant in…
-
We present ARCANA, a collaborative multi agent framework for solving ARC AGI 2 tasks under strict test time and hardware constraints. ARCANA decomposes each task into iterative perception, hypothesis generation, symbolic execution, and ref…
-
Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionary search, exhaustive sampling, extended chain-of-thought), or benchmark-specific training…
-
Fable 5 has been out for a couple of days, why isn't it on ARC-AGI 3? (www.reddit.com via reddit)
Considered by many to be the best benchmark for abstract reasoning, it's not saturated, and yet it's not included in several benchmarks, let alone this one. I understand they said they were going to do it.
-
Large language models can produce fluent, internally coherent reasoning traces for abstract reasoning tasks while still being confidently wrong - making selection among candidates, not just generation, the central challenge. I present a so…
-
-
A one-parameter model that gets 100% on ARC-AGI-2 (eitanturok.github.io via hn)
TLDR: I built a model that has only one parameter and gets 100% on ARC-AGI-2, the million-dollar benchmark that pushes reasoning models to their limits. Using chaos theory and some deliberate cheating, I crammed every answer into a single…
-
Anthropic Opus 4.8 is new SOTA on ARC-AGI-3, Score: 1.5%, –$10K (xcancel.com via hn)
Anthropic Opus 4.8 is new SOTA on ARC-AGI-3 Score: 1.5%, ~$10K ARC-AGI-3 analysis notes: * Opus 4.8 read the environment an abstraction *above* Opus 4.7, as objects & systems, not pictures * Opus 4.8 succeeded on early levels, but still co…
-
Ring-2.6-1T made me think about failure placement more than headline strength. It’s a trillion-parameter reasoning model for agent workflows with high and xhigh reasoning-effort modes.
-
Ring-2.6-1T made me think less about “is this good?” and more about routing. The public profile looks like something I'd at least test for harder agent steps: PinchBench 87.60, AIME 26 95.83, GPQA Diamond 88.27, Tau2-Bench Telecom 95.32, b…
-
🔓 13 "Impossible" ARC-AGI-2 Tasks — All Solved These 13 ARC-AGI-2 evaluation tasks have never been solved by any AI system — not GPT-4, not Claude, not Gemini, not NVARC, not MindsAI, not any Kaggle submission. They have a 0% AI solve rate…
-
OpenClaw leads official ARC-AGI-3 community leaderboard (arcprize.org via hn)
ARC-AGI Community Leaderboard ARC-AGI has gained significant popularity over the past two years, and we've been overwhelmed by the number of researchers and builders who want to showcase their work to the community. The ARC-AGI Community L…
-
Grok vs. ChatGPT vs. Gemini Comparison 2026: Complete Guide (Tested) (aithinkerlab.com via hn)
The 30-Second Verdict Best for science & reasoning: Gemini 3.1 Pro — leads GPQA Diamond (94.3%) and ARC-AGI-2 (77.1%). Best for coding: ChatGPT (GPT-5.5) — 88.7% on SWE-Bench Verified.
-
Playing Around with the ARC-AGI-3 Benchmark (bengoertzel.substack.com via reddit)
AGI, frontier science, maniacal metaphysics, decentralizationist politics, life and consciousness extension and expansion, psi and psychedelics and etc. etc.
-
I'm not sure too many people care about the ARC-AGI-2 competition anymore, but still...I thought some might find this interesting. They're running it one last time this year.
-
Replayable traces of Claude Code runs on ARC-AGI-3 public demo games (arc-agi-runs.web.app via hn)
GAME ar2511 bp3513 cd8211 dc2210 g50t10 ka5911 lf5210 ls201 m0r029 r11l1 re863 su1510 vc3311 wa3010 VARIANT A0-replay-m0r01 A1-fresh-m0r03 A10-unseen-game3 A11-smaller-model3 A2-no-theory3 A3-no-journal3 A4-no-scratchpads3 A5-no-code-writi…
-
I was reading the comments to this post and the overall opinion seemed to be that harness makes little/no difference for ARC-AGI-3. Turns out, it makes a huge difference: Hill-climbing ARC-AGI-3 TLDR: if you save game logs - taken actions,…
-
ARC-AGI-3 Update (GPT-5.5 High and Opus4.7) (www.reddit.com)
- GPT-5.5: 0.43% - Opus 4.7: 0.18% ARC-AGI-3 is no joke. I can’t wait to see which models finally crack.
- Analyzing GPT-5.5 and Opus 4.7 with ARC-AGI-3 (arcprize.org)
-
Common GPT 5.5 pricing misconception. (www.reddit.com)
Many people have pointed out that ChatGPT 5.5 appears to be twice as expensive as 5.4 based on API pricing, which makes it look pricier than Opus 4.7. But the comparison is not that simple.
-
The Human Baseline for ARC-AGI-3 has been updated (www.reddit.com)
could not extract summary
- Measuring Human Performance on ARC-AGI-3 (arcprize.org)
-
GPT vs Claude in a bomberman-style 1v1 game (www.reddit.com)
A few weeks ago, ARC-AGI 3 was released. For those unfamiliar, it’s a benchmark designed to study agentic intelligence through interactive environments.