
Claude Opus 5 pushes prompt-to-game AI from rough color blocks to full 3D prototypes with physics and music
Quick Answer
Claude Opus 5 enables the creation of full 3D game prototypes directly from prompts, outperforming previous models like Claude 4.
Quick Take
Users have generated diverse games, including a Call of Duty-style shooter and a Minecraft clone, all with procedural assets and physics running in the browser. This marks a significant leap in AI-driven game development, showcasing improved visual outputs and interactivity.
Key Points
- Opus 5 generates complete 3D games with procedural assets, eliminating the need for external assets.
- Demos include a shooter, submarine game, and a Minecraft clone, all created from single prompts.
- Visual outputs from Opus 5 are significantly improved compared to Claude 4, with realistic physics.
- Community tests include recreating Minecraft, showcasing Opus 5's capability in procedural generation.
- Opus 5's code-based output allows for easy editing and reuse, unlike video-based models.
DeepSignal Analysis
What happened
Claude Opus 5 allows users to create full 3D game prototypes directly from prompts, showcasing significant advancements over Claude 4. Users have generated various games, including a first-person shooter and a Minecraft clone, all with procedural assets and physics running in the browser.
Key evidence
- Matt Shumer created a Call of Duty-style shooter using only Opus 5, with no external assets involved.
- Alex Ermolov reported that Opus 5 produced a snowboard demo with no visual glitches and realistic sliding physics.
- Pankaj Kumar's Minecraft clone built with Opus 5 features an infinite procedural world and generated all graphic assets at runtime.
Why it matters
The advancements in Claude Opus 5 represent a shift in AI-driven game development, moving from traditional asset pipelines to procedural generation. This could streamline the development process and democratize game creation, allowing users without extensive programming skills to produce complex games. The ability to generate complete game prototypes directly from prompts may also encourage innovation in game design.
What to watch
📖 Reader Mode
~4 min readA first-person shooter, a submarine game, a kart racer, a Minecraft clone: over the past few days, people have generated all of these from a single prompt each. The models write the geometry, the textures, and sometimes even the music as code that runs straight in the browser.
Writing a scene from scratch takes more than clean syntax because the model has to know how the world actually looks and moves, how a snowboard slides over snow, or how wind travels through a field of grass. None of that is in the prompt; it has to come from the model itself.
Six demos built without a single external asset
Matt Shumer posted a Call of Duty-style shooter where everything on screen, right down to the textures, is generated code. "Not a single external asset was used," he wrote on X. The prompt and source code are up on GitHub under the name Claude-of-Duty.

Alex Ermolov put out a snowboard demo and said Opus 5 is on par with Fable and ahead of every other model he's tested. His first run came back with no visual glitches, he says, and the "sliding physics feel right."
Chetaslua's ballista makes the gap to other models pretty stark. The Roman siege weapon sits in front of a castle, has a complete control panel, and feels closer to a finished game than a prototype. Chetaslua had already run the same prompt through GPT-5.6 Sol and Kimi K3. Both managed a recognizable ballista, but with fewer working parts and far less going on around it.
User Chris ran a comparison that shows how quickly this has improved. He put year-old output from Claude 4 Opus next to what Opus 5 produces today. For the new run, he asked for a car on a dirt road, explicitly without textures. Opus 5 came back with mud tracks, vegetation, and layered lighting. A year ago, the same kind of request produced flat blocks of color.

More demos show the range of what's possible.Lentils packed a scenic landscape into one HTML file, with millions of grass blades bending to simulated wind. Pietro Schirano built a submarine game where the model generated the 3D objects, textures, and music itself. Ryan Campbell put his kart racer online as a playable website.
Procedural code replaces traditional asset pipelines
In traditional 3D development, developers assemble assets from libraries, load textures as image files, and handle physics through an engine. The Opus 5 demos work differently: Geometry is procedural, built by code that computes points and surfaces at runtime. Textures exist as shaders, which is to say more code. The model writes the physics and controls along with everything else, and since it all runs in the browser on Three.js, one HTML file is enough.
Anthropic says in its Opus 5 announcement that the model produces "much stronger visual outputs" than earlier versions. One testing partner quoted there calls the results the best animations, games, and 3D work he's seen from any Opus model.
It's a different approach from Google's Genie 3 or the video models sometimes described as world models. Those systems compute every frame and hand back an image sequence. The output only exists as a video stream, so there's nothing to edit or reuse afterward. What Opus 5 produces is code you can open and change.
Minecraft clones are becoming the community's go-to stress test
Beyond the one-off demos, the community has settled on an informal test of its own: rebuild Minecraft from a single prompt. Pankaj Kumar's Opus 5 version has an infinite procedural world, 15 biomes, and both survival and creative modes. Every graphic asset comes from code, he writes, with textures and sounds generated at runtime. The build ate 25 million tokens.
Other head-to-head tests follow the same script. Harshith had Opus 5, Fable 5, GPT-5.6 Sol, and several other models generate a 3D Airbus H145 in Three.js and lined up the results. atomic.chat went a step further and had all four models build three physics scenes, among them a tornado tearing across a field and a wrecking ball swinging into an apartment block. By atomic.chat's account, only Opus 5 handled all three convincingly.

Prompts like these work as a rough first gauge, since anyone can recreate them and judge the results on screen. Programming skills aren't required, though they can help with the prompting. A benchmark in any methodical sense, however, this is not. There are no standardized tasks, no rating scale, and no control over how many attempts each model gets.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

