
Unicorn, pelican, Middle-earth: OpenAI co-founder Karpathy is looking for the next AI vibe test
Quick Answer
Andrej Karpathy utilized Claude Opus 5 to transform a paragraph from 'Lord of the Rings' into a 3D browser scene, demonstrating the model's potential for on-demand game worlds.
Quick Take
The project, which ran for two hours with a budget of one million tokens, highlights the limitations of existing benchmarks like the pelican test, as it produces vibe checks rather than hard metrics.
Key Points
- Karpathy's project rendered a 3D scene using Claude Opus 5 and Three.js.
- The model operated for two hours with a budget of approximately $10.
- Existing benchmarks like the pelican test are seen as nearing their limits.
- Karpathy envisions on-demand game worlds for users as characters.
- The model's reliance on screenshots led to output errors.
DeepSignal Analysis
What happened
Andrej Karpathy used Claude Opus 5 to convert a paragraph from 'Lord of the Rings' into a 3D browser scene, generating 5,500 lines of code. The project ran for two hours with a budget of one million tokens, showcasing the model's capabilities for creating game worlds. However, existing benchmarks like the pelican test are seen as insufficient for measuring performance.
Key evidence
- Karpathy's project with Claude Opus 5 produced a 3D scene from Tolkien's text, resulting in 5,500 lines of code.
- The model operated for about two hours and utilized a budget of one million tokens, approximately costing $10.
- Karpathy argues that the pelican test is nearing its limits, as it primarily provides vibe checks rather than concrete metrics.
Why it matters
This project illustrates the evolving capabilities of AI models in generating complex visual outputs from textual prompts. Karpathy's work highlights the potential for on-demand game creation, which could change how users interact with digital environments. However, the reliance on subjective assessments like vibe checks raises questions about the effectiveness of current evaluation benchmarks in the AI field.
📖 Reader Mode
~2 min readOne paragraph of "Lord of the Rings" in, 5,500 lines of code out. Andrej Karpathy had Claude Opus 5 turn Tolkien's opening into a 3D browser scene.
The model worked for about two hours, had a budget of one million tokens (roughly $10), and rendered everything in Three.js, a graphics library for 3D in the browser. Karpathy posted the source code online so others can play back and modify the scene.
Karpathy's argument is that the well-known pelican test is close to maxed out. Simon Willison's fun benchmark asks a model to draw an SVG of a pelican riding a bicycle, a prompt chosen to be absurd enough that the subject wouldn't already exist in training data. The idea goes back further. In 2023, Microsoft had GPT-4 draw a unicorn in TikZ as part of its Sparks of AGI paper and pointed to it as evidence of spatial reasoning. GPT-4 at the time was trained on text data only, unlike the multimodal models that now power every major chatbot.
Karpathy calls the result rough but fun. Eleven Labs provided the audio, while the model placed and animated objects on its own. Nobody would build worlds like this by hand, Karpathy says, but with a model it's nearly free. He sees potential for on-demand game worlds where users could be dropped in as side characters or protagonists.
One problem remains. The model can't watch its own video output and had to rely on screenshots instead, which led to errors. Other users are having Opus 5 write entire browser games, from first-person shooters to Minecraft clones. None of this produces hard benchmarks. They're vibe checks.
Karpathy co-founded OpenAI and served as head of AI at Tesla. He later ran the education startup Eureka Labs and has been working on pre-training research at Anthropic since May 2026.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

