Schema: Discovering Unknown Environments
via Agentic Program Induction

1Impossible Research 2UC Berkeley 3Carnegie Mellon University *Project leads
Schema observes a world, induces a program of its mechanisms, and uses the program to plan and act. Benchmark comparisons show gains across ARC-AGI-3, DiG-bench, and MazeBench.

Learning unknown worlds through programs. Schema reaches 99.2% RHAE on ARC-AGI-3, solves all 21 public DiG-bench games, and exceeds the top-50 human median on MazeBench. The plots compare harnesses using the same base models.

LLM Agents in Unfamiliar Environments

In an unfamiliar environment, an agent does not know the rules, the meaning of its observations, or even what counts as success. It must discover these through interaction, then use what it learns to act efficiently.

Humans can make progress by forming a provisional explanation of how the world works, testing it through interaction, and revising it when observations disagree. The same evolving understanding guides both what to investigate and what to do next.

How can we build an agent harness around these same principles?

Schema: Interactive Program Induction

Schema is an agent harness with a persistent program workspace and four operations: hypothesize, certify, plan, and act with verification. The LLM decides what to investigate, when to revise its program, and how to act. The harness makes those programs testable against experience and usable for planning.

The Schema harness connects hypothesis construction, history certification, planning, and verified execution. A prediction mismatch returns a counterexample to the agent.
  • Hypothesize. Build and refine an executable theory of the environment.
  • Certify. Check the theory against all recorded experience.
  • Plan. Use the theory to design experiments and plan actions.
  • Act with Verification. Execute actions and stop when observations contradict predictions.

Experiments

We evaluate Schema on three complementary challenges: action efficiency, hidden-mechanism discovery, and long-horizon learning. Within each benchmark, we compare against baseline harnesses using the same base model.

ARC-AGI-3: Human-level Action Efficiency

Across 25 public visual games, an agent must discover unfamiliar mechanics and apply them to increasingly complex levels. Schema improves results across all four model configurations. With Claude Fable 5, it raises RHAE from 58.7% to 99.2% and clears all 25 games, compared with 14 for Claude Code.

ARC-AGI-3: Schema reaches 86.1, 99.2, 91.8, and 96.7 percent RHAE across Opus 4.8, Fable 5, Sol xhigh, and Sol max.

Consistent gains across base models. Schema improves action efficiency, game completion, and the fraction of games reaching full RHAE. RHAE measures completion and action efficiency relative to first-time humans, not the percentage of games solved.

DiG-bench: Efficient Mechanism Discovery

DiG-bench tests hidden-rule discovery in text-based games with limited lives and action budgets. With GPT-6 Astra, Schema wins 19, 20, and 21 games at medium, high, and maximum reasoning effort, compared with 11, 16, and 19 for the basic harness. At medium effort, it loses only 34 lives rather than 96.

Schema solves more DiG-bench games and loses fewer lives at every reasoning effort, winning all 21 public games at maximum effort.

More games solved, fewer failed attempts. Schema tests competing explanations before risking lives on new answers.

MazeBench: Knowledge Accumulation over Long Horizons

MazeBench is a world of 256 interconnected rooms and 100 gems whose mechanisms must be discovered through interaction. Schema with GPT-6 Astra collects 33 gems and visits 139 rooms in 27,819 actions, exceeding the top-50 human median on both metrics.

MazeBench progress and room coverage: Schema reaches 139 rooms, compared with 113 for Codex using GPT-6 Astra. Colored cells show the rooms visited by each run.

Continued progress and broader coverage. At 20,673 actions, Schema has 30 gems and 133 rooms, compared with 23 gems and 113 rooms for Codex using the same base model. The coverage maps show each run’s recorded endpoint: 139 of 256 rooms (54.3%) for Schema and 113 rooms (44.1%) for Codex. Colored cells mark visited rooms; white circles mark the shared starting room. Dotted lines in the progress plots show the top-50 human median.

What Makes Schema Effective?

Ablations and behavioral traces show how Schema's components support both discovering mechanisms and retaining them over time.

Contribution of Each Components

Replacing the program with prose, or removing certification, planning, or execution-time verification, reduces RHAE from 72.9% to 52.4–62.3% on the ten hardest ARC-AGI-3 games.

Table 1 from the paper. RHAE: Schema 72.9%; prose model 58.8%; without certification 62.3%; without planning 59.2%; without verification 52.4%.

Claude Opus 4.8 · 10 hardest public games · 75 levels · 2,000 actions per game.

Behavioral traces show effects of program representation, certification, planning, and runtime verification.

(a) Representation — CN04. Each marker denotes a completed level, plotted against cumulative actions. Executable programs enable faster progress than prose: level 5 takes 279 actions with Schema versus 1,219 with the prose model.

(b) Certification — AR25. Each revised program is checked offline against recorded transitions before new actions. Without certification, revisions can break previously correct predictions. Schema’s final program reproduces 99.6% of its interaction history.

(c) Planning — LS20. The curves compare level completion with and without search inside the program. On level 5, Schema takes 65 actions versus 233 without planning, even though the latter’s model passes all 252 historical checks at level entry. Gray dashed curves in (a) and (c) show human baselines.

(d) Verification — ten games. Each bar shows the percentage of actions executed after the first prediction mismatch within the same committed plan, when execution-time verification is disabled. Aggregated across these games, such continuations account for 58% of all actions.

Knowledge Compaction and Reuse

By approximately 20,000 MazeBench actions, Schema's program contains 2,012 non-comment lines, compared with 6,133 lines of notes in the prose-model variant. The number of transitions explained per line and function reuse both increase, while forward prediction accuracy remains high.

Schema retains growing experience in a compact program with increasing function reuse and high forward prediction accuracy.

Experience accumulates in shared rules. The program remains compact as the interaction history grows, while preserving predictive value.

Case Studies

Program induction makes beliefs explicit and reusable.

CN04 · Level 5

0 actions0 actions

Join every connector without overlapping the pieces. One piece changes shape as it grows.

The program represents piece bodies and connectors separately, then searches rotations and translations for an arrangement with no body overlap. Schema finishes in 279 actions; its completed board stays visible while the prose model continues through all 1,219 actions.

Certification helps the agent validate hypothesis correctness.

AR25 · Level 3

0 actions0 actions

Move pieces and mirror axes until the reflected shapes cover the targets.

On level 3, a revision without certification breaks earlier predictions: replay of the same 16 transitions from level 1 falls from 16/16 to 1/16. The replay then moves to level 6, where the model confuses two touching pieces. Instead of fixing that recorded mismatch, the agent resets and routes around it with the same model, finishing in 119 actions versus 53 for Schema. Schema’s final certified model reproduces 267/268 transitions (99.6%).

Planning ahead improves action efficiency.

LS20 · Level 5

0 actions0 actions

Match the carried symbol to the goal. Contact with the moving white marker rotates it.

Schema’s final 13-action plan starts with down, then up, timing contact with the moving marker so the carried symbol rotates to match the target. The remaining eleven actions reach the goal. The full level takes 65 actions versus 233, even though the model without planning passes all 252 historical checks at level entry.

Runtime verification further boosts adaptation to new scenarios.

LS20 · Level 2

0 actions0 actions

Reach the goal before energy runs out. Yellow stations refill energy, but each works only once.

Both agents encounter the same missing-station error. Schema stops at the mismatch and updates its model. Without verification, the agent executes another 42 actions, returns to the spent station, and runs out of energy before correcting the rule. The complete runs, including retries, take 59 and 143 actions.

Gameplay Videos

Explore all 25 ARC-AGI-3 runs, all 21 DiG-bench games, and the full MazeBench replay. Select a game below to watch Schema learn and act, with observed game states and progress from the recorded runs.

Schema · Claude Fable 5

LS20: discover movement, symbol transformations, and energy constraints, then reuse the program to plan through all seven levels.

BibTeX

@misc{zeng2026schema,
  title={Schema: Discovering Unknown Environments via Agentic Program Induction},
  author={Guanning Zeng and Jiani Wang and Wenjie Ma and Shaofeng Yin and Chenyang Wang and Shichen Liu and
          Angjoo Kanazawa and Wode Ni and Xiuyu Li and Andrea Zanette and Haiwen Feng},
  year={2026},
  eprint={2609.39140},
  archivePrefix={arXiv},
  primaryClass={cs.AI},
  url={https://arxiv.org/abs/2609.39140}
}

@misc{schema2026blog,
  title={{[schema]: Frontier Models with the Right Harness Achieve $\sim$99\% on ARC-AGI-3 Public}},
  author={Zeng, Guanning and Wang, Jiani and Ma, Wenjie and Yin, Shaofeng and Wang, Chenyang and Liu, Shichen and
          Kanazawa, Angjoo and Ni, Wode and Li, Xiuyu and Zanette, Andrea and Feng, Haiwen},
  year={2026},
  howpublished={Impossible Research. \url{https://schema-harness.github.io/}},
  url={https://schema-harness.github.io/},
}