- Type
- AI benchmark series
- Created by
- François Chollet
- First released
- 2019
- Measures
- Fluid intelligence and skill-acquisition efficiency
- Maintained by
- ARC Prize Foundation
- Related
- AI benchmarking, reasoning models, AGI
- Type
- AI benchmark series
- Created by
- François Chollet
- First released
- 2019
- Measures
- Fluid intelligence and skill-acquisition efficiency
- Maintained by
- ARC Prize Foundation
- Related
- AI benchmarking, reasoning models, AGI
History
Chollet proposed ARC-AGI as a critique of how the AI field measured progress. His argument was that raw task skill is a poor proxy for intelligence, because skill can be bought with training data and compute: "the intelligence of a system is a measure of its skill-acquisition efficiency over a scope of tasks, with respect to priors, experience, and generalization difficulty."[1] A system that memorises answers is not demonstrating the ability to learn new ones.
The original ARC-AGI-1 dataset consists of grid-based reasoning puzzles, each showing a few input–output examples from which the test-taker must infer the underlying rule. ARC-AGI-2, released in 2025, expanded the task set and raised difficulty substantially after frontier models began saturating the first version.[6] ARC-AGI-3, introduced subsequently, moved from single-shot puzzles to interactive environments in which an agent must explore, plan and act over multiple turns rather than emit one answer.[4]
The ARC Prize is an annual public competition, run on Kaggle since 2024, carrying a cash prize for the best open-source solution. At the end of 2024 the leading private-evaluation score stood at about 53 percent.[5] In the 2025 competition the winning entry (NVARC) reached 24.03 percent on the ARC-AGI-2 private set, with the runner-up at 16.53 percent — a field in which scores remain far below saturation.[3]
Design Principles
Three ideas define the benchmark's construction:
- Core knowledge priors. Tasks draw only on cognitive building blocks present in early human development — objectness, basic geometry, counting, symmetry — following Elizabeth Spelke's Core Knowledge theory. Language and cultural knowledge are deliberately excluded so that neither humans nor models get an unfair advantage from prior exposure.[2]
- Easy for humans, hard for AI. Puzzles are chosen because people solve them with little effort. This makes the benchmark a direct probe of the gap between human and machine generalisation rather than a test of specialised expertise.[2]
- Skill-acquisition efficiency. Success requires inferring a rule from two or three examples and applying it to a novel instance — the opposite of retrieving a memorised pattern.
Performance of Frontier Models
Reported scores improved sharply between 2024 and 2026. OpenAI's o3 was the first model to draw wide attention to the benchmark, with a high-compute configuration scoring 87.5 percent on ARC-AGI-Pub, the semi-public subset used to slow contamination.[7] On the ARC Prize leaderboard, GPT-5.2 Pro (High) records 90.5 percent on ARC-AGI-1 but 54.2 percent on ARC-AGI-2, while Claude 4.7 (Max) reaches 75.8 percent on ARC-AGI-2.[4] The persistent gap between the two versions illustrates the benchmark's design goal: each generation of harder tasks reopens the measurement.
Critics note that ARC-AGI scores can rise through test-time search and heavy sampling rather than improved one-shot generalisation, which is why the ARC Prize leaderboard reports compute cost per task alongside accuracy.[4]
>Key Takeaways
- ARC-AGI measures how efficiently a system acquires new skills, not how much it already knows.
- Tasks are restricted to core knowledge priors and are designed to be trivially easy for humans.
- Scores have risen from roughly 53 percent (ARC-AGI-1, end of 2024) to over 90 percent on ARC-AGI-1, while ARC-AGI-2 and the interactive ARC-AGI-3 remain far from saturation.
- Compute cost per task is reported alongside accuracy, since test-time search can inflate results.
See Also
🇲🇾 ARC-AGI matters for Malaysia as a measurement question rather than a local industry one. National strategy documents such as the National AI Action Plan and MDEC's AI Nation 2030 push adoption of AI tools, but adoption is not the same as capability: organisations evaluating vendors or models should look for independent, contamination-resistant evaluations rather than vendor-reported benchmark wins. Because ARC-AGI is language-free and culture-neutral, it also travels well — Malaysian students and researchers can compete on it without English-language or Western-knowledge advantages, and it is a useful teaching tool in AI literacy programmes such as AI Untuk Rakyat. Local model teams (Mesolitica's ILMU, Mallam and others) publishing results against international benchmarks help the ecosystem benchmark honestly, while the benchmark's emphasis on reasoning over memorisation aligns with the transparency and reliability principles in the Malaysia AI Governance Framework.[8]
References
- ↑Chollet, F. (2019). On the Measure of Intelligence. arXiv:1911.01547. https://arxiv.org/abs/1911.01547
- ↑ARC Prize Foundation. What is ARC-AGI? https://arcprize.org/arc-agi
- ↑ARC Prize 2025: Technical Report. arXiv:2601.10904. https://arxiv.org/html/2601.10904v1
- ↑ARC Prize Foundation. Leaderboard. https://arcprize.org/leaderboard
- ↑ARC Prize Foundation. History. https://arcprize.org/history
- ↑ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems. arXiv:2505.11831. https://arxiv.org/abs/2505.11831
- ↑ARC Prize Foundation. OpenAI o3 Breakthrough High Score on ARC-AGI-Pub. https://arcprize.org/blog/oai-o3-pub-breakthrough
- ↑AI Malaysia Berhad. National AI Action Plan 2026–2030. https://ai.gov.my/