ARC-AGI-3
An interactive benchmark for studying agentic intelligence through novel, abstract, turn-based environments in which agents must explore, infer goals, build internal models of environment dynamics, and plan effective action sequences without explicit instructions. A 100% score means AI agents can beat every game as efficiently as humans.
Best results
Frontier over time
All results
| # | Model | Score | Conditions | Eval date | Source | Flags |
|---|---|---|---|---|---|---|
| 1 | GPT-5.6 Sol (Max) | 7.80% | — | 09 Jul 2026 | Official leaderboard | Primary |
| 2 | GPT-5.6 Terra (Max) | 0.80% | — | 09 Jul 2026 | Official leaderboard | Primary |
| 3 | GPT-5.6 Luna (Max) | 0.20% | — | 09 Jul 2026 | Official leaderboard | Primary |
