OPINE-World: Programmatic World Modeling with Ontology-error-Prioritized Interactive Exploration
Quick Answer
OPINE-World is an LLM agent that learns object-centric programmatic world models through interaction, achieving an action-efficiency score of 78.4 on the ARC-AGI-3 benchmark, solving 20 out of 25 games without per-game training.
Key Points
- OPINE-World uses a loop of hypothesis and testing with two cooperating agents.
- The model employs Bayesian measures of object-type adequacy termed ontology error.
- It demonstrates data efficiency and reusability compared to traditional deep network models.
- The benchmark -3 tests skill-acquisition efficiency with withheld object vocabulary.
- OPINE-World's performance surpasses human baseline in action efficiency.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Learning how an environment behaves from interaction is central to building agents that adapt to unfamiliar tasks. World models learned with deep networks are flexible but data-hungry and transfer poorly beyond their training distribution. Program-synthesized world models, written as source code by LLMs and refined through counterexample-guided inductive synthesis (CEGIS), are instead data-efficient and reusable, yet they have been demonstrated mainly on structured-state worlds with a given object vocabulary, and a single program search does not scale to pixel-rendered environments whose object structure must be hypothesized flexibly. We introduce OPINE-World, an LLM agent that learns an object-centric programmatic world model online from interaction. OPINE-World couples two cooperating agents in a loop of hypothesis and test, one acting in the environment and one synthesizing the model in code with replay verification and model-based planning, and it steers exploration with a Bayesian measure of object-type adequacy we call ontology error. We evaluate OPINE-World on ARC-AGI-3, a benchmark for skill-acquisition efficiency in which the object vocabulary, the goal, and the action semantics are withheld. OPINE-World solves 20 of 25 games without per-game training and reaches an action-efficiency score of 78.4 against the human baseline.
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.01531 [cs.AI] |
| (or arXiv:2607.01531v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.01531 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: David Courtis [view email]
[v1]
Wed, 1 Jul 2026 23:04:47 UTC (983 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.