Research program

ARC-AGI-3: A bounded test of experience-grounded intelligence

ARC-AGI-3 is our first bounded testbed for a broader research question: can an agent enter an unfamiliar world, discover its structure and goals through interaction, and act efficiently without relying on prior task instructions?

Research question

A first-contact learning process

The benchmark brings exploration, world-model formation, goal inference, planning, and execution into a single first-contact learning process.

Unlike a static task with a fully specified objective, an interactive environment requires an agent to learn what matters while it acts. Observations acquire meaning through intervention and consequence; goals must be inferred from feedback; and useful structure must be carried forward across a sequence of decisions.

This makes ARC-AGI-3 a compact setting in which to ask whether experience can organize an agent's representations and control, rather than merely supply additional context to a fixed solver.

What it tests

The minimum coupled problem

Exploration

Choosing informative actions before the environment's rules, objects, or goals are known.

World-model formation

Discovering persistent structure, predicting consequences, and revising beliefs when evidence contradicts them.

Goal and subtask discovery

Inferring what constitutes progress and organizing intermediate objectives without natural-language task instructions.

Planning and execution

Using what has been learned to act efficiently over time while continuing to update the underlying model.

Why bounded

A testbed, not a definition of intelligence

ARC-AGI-3 does not contain the open-ended scale, social structure, language, or cultural inheritance of the human world. It is useful precisely because it isolates a smaller question: whether an agent can turn interaction into structured knowledge and increasingly effective action in a novel environment.

For True Intrinsics, performance is not only a final score. The learning trajectory matters: which actions the agent selects, how quickly it identifies recurring structure, how it revises failed hypotheses, and how much experience it requires before behavior becomes organized.

The benchmark is therefore a minimal experiment within a broader research program. It gives the central thesis a place to fail, be measured, and be revised.

Benchmark definition and resources: ARC Prize — ARC-AGI-3.