Every new AI model, tested live. Every demo, yours to play.
I'm Vito Ruocco. Every week I put the newest AI models to work on real games and apps, film the results, and publish everything here so you can try it yourself in the browser.
AI demos you can play
Games and apps written by the newest models in one prompt. Same prompts, side by side.
VOXARA: one game prompt, five AI models
A multiplayer voxel sandbox written by Opus 5.5, GPT-6.1 Sol, Luna, Haiku 5.5 and Step-5. Play them and vote.
Claude Haiku 5.5 vs GPT-6 Luna vs Opus, Sol
Haiku 5.5 and GPT-6 Luna, same price per token, on the prompts Opus 5.5, Sol and Sonnet 5.5 got. 19 demos in your browser.
Mistral Large 4 vs Opus 5.5 vs GPT-6.1 Sol
Mistral's 1-trillion-parameter open model on the same four prompts as Opus 5.5 and GPT-6.1 Sol. All 12 demos run in your browser.
GPT-6.1 Sol vs Claude Opus 5.5: the $200 plan test
The same four prompts to the models behind ChatGPT Pro and Claude Max: a racing game, a working 3D kitchen, a voxel pagoda and a card machine. Use them in your browser.
Kolibri-1 vs Qwen, Nemotron and Mistral: same prompts
Germany's open Kolibri-1 against the models it says it beats: one game and one app, same prompts, ten demos you can use in the browser.
MiniMax M3.1 Flash? Space Bunny vs MiniMax M3
Demos written by Space Bunny Alpha (the free model the community links to MiniMax M3.1 Flash) and by MiniMax M3, on the same prompts: a voxel island, a wave shooter and a desktop OS. Play them all.
Cloudline: one prompt, five models
The same tram-game prompt given to GPT-6 Astra, an early Gemini 4 checkpoint, Claude Sonnet 5.5, GPT-6.1 Sol and Gemini 3.8 Flash. Play all five.
Claude Sonnet 5.5 Demos
Five apps and games written in one prompt each by Claude Sonnet 5.5, Anthropic's mid-size model. Open them and try them live.
GPT-6.1 Sol Demos
Four apps and games written in one prompt each by OpenAI's GPT-6.1 Sol, plus a head-to-head against Claude Sonnet 5.5 on the exact same prompts. Open them and try them live.
Space Bunny Alpha Demos
Six apps and games written in one prompt each by Space Bunny Alpha, the free mystery AI model. Open them and try them live.
Fable 5.1 vs Ling 3.1 Flash (free): the Fable 5.5 tests
Four tests people ran on the leaked Claude Fable 5.5, given to Fable 5.1 and to the free Ling 3.1 Flash: an art-history film, a forest jeep, a rope puzzle and a controller designer. Use them all in your browser.
Get the exact prompts behind every demo
- Every prompt word for word, with model, settings, time and cost
- The full source code of every demo, as a ZIP
- Step-by-step notes on how each video was made
- Members-only videos and uncut gameplay
€9 / month
Cancel anytime
All articles
251 articles on AI models, leaks, benchmarks and gaming. Browse them all.

5 AI Models Built a Multiplayer Voxel Game: VOXARA Tested
Opus 5.5, GPT-6.1 Sol, Luna, Haiku 5.5 and Step-5 got one prompt: a multiplayer voxel game. Play all five and vote.…

Claude Haiku 5.5 vs GPT-6 Luna: Same Price, One Winner
Claude Haiku 5.5 and GPT-6 Luna cost the same per token. I gave both the prompts Opus 5.5, Sol and Sonnet 5.5 got. 19 demos, live.…

Google Playground, Capcom’s AI Engine Evolution, Sony’s IP Revival & Xbox’s Bold Comeback — The Biggest Gaming Stories of October 2026
Introduction: A Week of Transformation in Gaming If there's one word to describe the gaming industry this first week of October 2026, it's transformation. From …

October 8, 2026: The Day AI Went Geopolitical — From Mistral’s Trillion-Parameter Beast to Trump’s Golden Age Summit
If you blinked this week, you missed a seismic shift in the technology landscape. On October 7th and 8th, 2026, three monumental stories converged to reshape th…

GPT-7 Bel’s 722 Math Papers + Mistral Large 4 Tested
OpenAI released 722 math papers from an unreleased model. I gave Mistral Large 4 the prompts Opus 5.5 and GPT-6.1 Sol got.…

GPT-Next Leaked: GPT-6.1 Sol vs Claude Opus 5.5, Tested
Claude Max gives 5x the usage of ChatGPT Pro, says SemiAnalysis. I gave both $200 models the same four prompts.…

October 2026: The Month Gaming Went Supernova — Gears E-Day, GTA 6, Steam’s £11bn Year, and the Industry at a Crossroads
From Gears of War: E-Day's triumphant return to GTA 6 looming on the horizon, Steam's £11 billion year, PlayStation's disc-less future, and blockbuster sales — …

Beam 501B: Reflection AI Unleashes the Open-Source AI Colossus That Changes Everything
Introduction: The Sound of a Paradigm Shift On a quiet Monday in early October, the AI world woke up to a thunderclap. Reflection AI — a relatively young lab th…

Kolibri-1 Tested: Germany’s Free AI vs Qwen and Mistral
Aleph Alpha's free Kolibri-1 beats Qwen on paper. I gave it and four rivals the same prompts. Here is what really happened.…
More from VitoMind
New videos every few days on YouTube.

Claude Haiku 5.5 vs GPT-6 Luna: Same Price, One Clear Winner

GPT-7 Bel Wrote 722 Math Papers, So I Tested Mistral Large 4 vs Opus 5.5

GPT Next Leaked + Claude Max vs ChatGPT Pro I Tested Both $200 Models

Germany's FREE AI Model Beats Qwen I Tested It Against 4 Rivals

