GPT-Next Leaked: GPT-6.1 Sol vs Claude Opus 5.5, Tested

Published: October 6, 2026 | By Vito Ruocco

OpenAI has a new model that nobody outside the company was supposed to see. It is called GPT-Next, and one of the first things people saw it build was a voxel pagoda floating in the clouds. The same weekend, SemiAnalysis published a report claiming that Anthropic’s $200 Claude Max plan gives you about five times more than OpenAI’s $200 ChatGPT Pro plan. Those two stories meet in one question: if you are going to spend $200 a month on AI, which model actually builds better things? So I took four prompts in the style of the GPT-Next leaks and gave them, word for word, to the two models behind the two plans: GPT-6.1 Sol and Claude Opus 5.5.

This is the written version of my latest VitoMind video. It covers what we really know about GPT-Next and why the number people are excited about is its price, OpenAI’s 28-day shipping pledge (day two landed this morning), the GPT-6.1 Astra model OpenAI pulled for safety reasons and the political reaction in Washington, Reflection AI’s 501-billion-parameter open model, Anthropic’s report on GLM-5.3, everything new from Google, OpenAI’s text watermark for Europe, two useful Claude updates, and the SemiAnalysis report. Then comes the part that matters most: the same four prompts on Sol and Opus 5.5, all eight results published on my site, and my honest verdict, with the costs.

If you prefer reading, everything from the video is below, with the sources, the exact numbers and a few details that did not fit into eight minutes.

The Short Version

  • GPT-Next is an unreleased OpenAI model spotted on October 5. The early outputs look a lot like GPT-6.1 Sol; the interesting part is an early estimate that it is very cheap to run.
  • OpenAI’s Codex team pledged to ship one clear improvement or a full usage reset every day for 28 days. Day one made GPT-6 Astra and GPT-6.1 Sol about 50% faster; day two made Auto-review free.
  • GPT-6.1 Astra was pulled after internal testing found alignment problems, the FTC says it will compel AI executives to testify, and the White House created a “Super Intelligence Force”.
  • Reflection AI’s Beam is a 501B-parameter open-weight model with 23B active, with weights promised this month under Apache 2.0. Anthropic, meanwhile, warns that GLM-5.3’s safeguards fall to simple tricks.
  • Google: new Gemini 4 Argon checkpoints, Nano Banana 2.1 is out, and a mysterious “Gemini RSI Model” placeholder.
  • SemiAnalysis: Claude Max delivers roughly five times the API-equivalent usage of ChatGPT Pro at the same $200.
  • My test: GPT-6.1 Sol won the pagoda and the racing game, Opus 5.5 won the kitchen, the card machine was a draw. Sol did all four for about $0.55; Opus cost about $2.98, plus about $3.85 lost on runs that returned nothing.

GPT-Next: what the leak actually shows

GPT-Next surfaced on Monday, October 5, when a well-known leaker on X, Fandu (@mrfanduu), posted a short screen recording of a voxel pagoda with the caption “Voxel pagoda by gpt-next model”. It is a small page called “A garden outside time”: a pagoda on an isometric island, cherry trees, a little day and night switch, all of it drawn with cubes and written as code. Credit for spotting the model itself went to another tester, @DavidSZD1.

The second line of that post is the one people got excited about. Fandu shared an early estimate: about 40,000 credits produced roughly 5 billion tokens worth of work, cache included. Nobody outside OpenAI can check that number, and it depends heavily on how much of the work was served from cache. But if it is even roughly right, GPT-Next is a very cheap model to run, and that would matter more than a small jump in quality.

A French leaker, Mirochill, posted more results the same afternoon, including a full 3D kitchen and a voxel island. His verdict was refreshingly honest: GPT-Next looks a lot like GPT-6.1 Sol, with no notable improvement to expect, and the only real difference could be the price, which is itself uncertain. I did not show his voxel island in the video because it is full of characters from a well-known game franchise; the kitchen is enough to see the style.

Keep that comparison in mind, because GPT-6.1 Sol is one of the two models I tested at the end. If GPT-Next really is “Sol, but cheaper”, then testing Sol today tells you a lot about what GPT-Next will feel like.

OpenAI’s 28-day pledge: day one and day two

GPT-Next might arrive sooner than you think, because of a promise. On Sunday, October 4, Tibo (Thibault Sottiaux), who leads Codex at OpenAI, posted that the team is locking in: the only things being worked on are simplifications, more efficiency for more usage, groundbreaking features, or new models. A few hours later came the challenge: “Over the next 28 days, each day we’ll either ship one thing that is a clear improvement and relevant for most codex/work users or ship a full reset.”

Day one arrived on Monday. GPT-6 Astra and GPT-6.1 Sol were made about 50% faster by default through the subscription, across OpenAI’s own products and the partners that use Sign in with ChatGPT, including OpenCode, Pi, Amp and Devin. In Tibo’s words, that means reaching about 50 tokens per second instead of 30. No settings to change on your side.

Day two landed on Tuesday morning: Auto-review is now free for everyone signed in with a ChatGPT account. Auto-review is a second agent that watches every action the main agent takes and blocks the risky ones, so you can let a long task run without approving every single step, and without the decision fatigue of the default sandbox mode that asks you about everything. Tibo’s diagram shows ten thousand actions, of which seven hundred and twenty were sent to auto-review and seven were denied. It no longer draws from your plan’s usage. You will find it under Settings, Permissions, Auto-review.

One more detail from the same weekend. Tibo hinted that “6.1” is coming, and many people assumed he meant the next Astra. He did not: he clarified that it is 6.1 Sol Ultrafast. The fast tier, in other words, not a new flagship.

Why not Astra? OpenAI pulled it

The reason there is no GPT-6.1 Astra is simple: OpenAI cancelled it. The Wall Street Journal reported on September 28 that the October release had been abandoned, and OpenAI confirmed it. CSO Online’s summary of the internal findings is blunt: the model “could evade oversight, misrepresent its actions and operate beyond its authorized scope, while attempting to use external tools it knew were unsafe.”

It is rare for a major lab to withhold a finished model, and it is happening at a moment when Washington is paying close attention. The FTC has confirmed it is investigating OpenAI, Anthropic and other labs over the harm their agents may cause to consumers, and it is drafting civil investigative demands to force testimony from the companies’ executives. An FTC official told Electronics Weekly that the agency plans to compel those executives to testify “about the dangers they allege their products may have to consumers, to Americans.”

Then, on Sunday, President Trump announced a “Super Intelligence Force”, led by the Director of National Intelligence, Jay Clayton, who effectively becomes the White House’s AI czar. NPR reported the appointment; according to Quartz and other outlets, the task force has 120 days to report on the risks and opportunities of AI. Whatever you think of the name, it means that the people building these models will be answering questions in public much more often.

Reflection AI’s Beam: a 501-billion-parameter open model

While some labs hold models back, a new American lab is going the other way. Reflection AI introduced Beam, its first open-weight model, on October 5.

Beam is a sparse mixture-of-experts model with 501 billion total parameters, of which only 23 billion are active for each token. That is what makes it efficient for its size: you pay for the compute of a 23-billion-parameter model on every token, while the full 501 billion parameters hold the knowledge. Reflection says it pretrained the model on 23.8 trillion curated tokens and ran its reinforcement learning on 10,500 NVIDIA GB300 GPUs over four weeks. It was trained from scratch, not built on top of another open model, and the company promises the weights, a technical report and a model card later this month, under the Apache 2.0 license.

On Reflection’s own charts, Beam trades blows with GLM-5.2 and Qwen 3.8 Max on coding and agent benchmarks such as DeepSWE, Terminal Bench and SWE-Bench Pro. As always with launch charts, wait for independent tests before you believe the bars. But a large, efficient, truly open model from a US lab is something the Western open-source world has been missing.

Anthropic on GLM-5.3: open weights cut both ways

Open weights also have a cost, and Anthropic’s Frontier Red Team just put numbers on it. Its research post on GLM-5.3, the latest model from Zhipu AI (Z.ai), says the model “has strong capabilities for autonomously building end-to-end cyber exploits”, at a level comparable to Anthropic’s own restricted Claude Mythos Preview. In Anthropic’s tests, GLM-5.3 built working end-to-end exploits in about 12% of attempts, against 14% for Mythos Preview, while earlier models were near zero.

The difference is what happens around the model. Claude’s cyber capabilities ship with safeguards, and the versions with reduced safeguards are limited to vetted defenders. GLM-5.3, Anthropic writes, was released without meaningful safeguards: “attackers can bypass GLM-5.3’s safeguards between 64% and 100% of the time with simple techniques.” Anyone can download it. Anthropic’s recommendation is to give vetted defenders more access to strong cyber tools and to have governments safety-test capable models, including GLM-5.3’s successors.

Google: Argon checkpoints, Nano Banana 2.1 and an “RSI” placeholder

Gemini 4 Argon is still only available to trusted testers, but new internal names keep showing up. The leaker Lyra (@lyraxana) posted a fresh batch of model placeholders: an “LSP Teamfood Slot”, “barium-b” with a tool leakage fix, and “barium-b” with summaries at two different medium-effort settings. A flurry of checkpoints like this, tested internally (“teamfood”) one after another, usually means a wider release is getting close.

Testers keep comparing it, too. One of the most shared tests this week was an SVG of a minimal isometric card-swiping machine: Gemini 3.5 Flash, a newer Flash, 3.5 Pro and Claude Fable 5 side by side. Mirochill says Gemini 4 Argon beat Fable 5 on this one. Remember that machine, because I gave the same idea to Sol and Opus.

Nano Banana 2.1 is now out on Gemini web and in the Gemini app, after showing up in Google Flow under the codename “beluga”. The Hype ran a fair comparison against Tencent’s HY Image 3.5: the same five prompts, one shot each, no edits, all in a 3D style. Nano Banana kept the characters closer to the reference image; HY Image won on light and mood; and Nano Banana ignored a “no text” instruction once. HY Image costs about $0.024 per 2K image through OpenRouter, while Nano Banana in Flow runs on subscription credits. On anime, Mirochill says the new model is better than before but still not great.

And then there is this. A leaker posted a placeholder from Google’s backend called “Gemini RSI Model rev latest”. RSI stands for recursive self-improvement: a model that helps improve itself. Some people called it proof. It is not. A name in a log proves nothing about what is behind it, so keep both feet on the ground with this one.

textGrain: OpenAI watermarks ChatGPT text in the EU

This one affects you directly if you live in Europe. OpenAI announced that over the coming weeks it will add an invisible watermark to eligible ChatGPT and Codex text output in the European Union, to comply with the transparency rules of the EU AI Act.

The technology is called textGrain. Nothing is hidden in the characters themselves: in OpenAI’s words, textGrain “adds an invisible statistical signal to the model’s word choices”, and a detector holding a secret key can look for that signal. OpenAI says that in its evaluations textGrain “matched or exceeded the performance of other approaches we tested, including SynthID for text”, with no meaningful loss in quality, and it plans to release the technology as open source. Outside the EU it is not a default: API customers anywhere can opt in for select models, and it stays off unless they do.

It has a weak spot, and OpenAI is open about it. As TechCrunch reports, replacing 10% of the words with synonyms dropped detection from about 92% to 66%, and short passages, math answers and translated text are harder to detect. A missing watermark, OpenAI adds, does not prove that a human wrote something. Angel (@Angaisb_) asked Opus 5.5 to make a short explainer video about how it works, which is a nice example in itself: “Which one of these was written by an AI?”

Two useful Claude updates

The first is a Claude Code mod called Cache Control, recommended by Daniel San (@dani_avila7). It draws a bar for Claude’s five-minute prompt cache and warns you before it expires. Keep the cache warm and Claude reuses context it has already processed instead of reading the same tokens again, so long sessions burn far fewer tokens and last longer before you hit your limits. If you missed it, I covered what Claude Code mods are in my previous video and article.

The second: Claude Projects can now connect to a folder on your computer. The session still runs in the cloud, but when a task needs your files, Claude asks to use one folder you approve and reads and edits the files in place, without you uploading anything. It started rolling out on October 5.

SemiAnalysis: Claude Max vs ChatGPT Pro

Now the big one. SemiAnalysis tested the usage limits of every major AI subscription plan, from Anthropic and OpenAI to Meta, SpaceX’s xAI, MiniMax, Moonshot, Z.ai, Cursor and Cognition. The headline is brutal: “Anthropic Subscriptions Offer 5x+ More Value Than OpenAI”.

The method is simple to explain. Subscriptions do not show you token prices, only a usage meter. So SemiAnalysis measured how much work you can do on each plan and calculated what the same work would cost if you paid for it through the API. On the $200 plans, Claude Max gives you about twelve thousand dollars’ worth of Opus 5.5 usage, while ChatGPT Pro gives you about two thousand dollars’ worth of GPT-6.1 Sol. Chetaslua’s summary of the numbers: about $12,529 of Sonnet, $11,726 of Opus and $2,485 of Fable on Claude, against $2,084 of Sol and $2,897 of GPT-6 Astra on ChatGPT. SemiAnalysis calls Anthropic “an overwhelmingly better deal, offering ~5x the API-equivalent value across the board” at the mid tier, and notes that OpenAI recently halved the tokens you get on each tier.

There are two caveats, and SemiAnalysis raises the first one itself. Sol is much cheaper per token than Opus, so “API-equivalent value” partly measures how expensive each model is, not only how generous the plan is. The second caveat is mine: more usage only matters if the output is good. Raw volume does not tell you which model builds the better thing. So I tested that.

The test: the same four prompts on GPT-6.1 Sol and Claude Opus 5.5

I wrote four prompts in the style of the GPT-Next leaks: a five-storey voxel pagoda on a floating island with a time-of-day slider, a fully interactive 3D kitchen where everything works when you click it, a polished hover-racing game with five AI rivals, and an isometric SVG card machine like the one testers gave Gemini and Fable. I gave each prompt, word for word, to OpenAI’s GPT-6.1 Sol and to Anthropic’s Claude Opus 5.5, through the API on OpenRouter: one request each, with one added line asking for a single self-contained HTML file.

In the video I use all eight results live, starting from the home page of this site: AI Demos, today’s card, the gallery, then each demo, the way any visitor would. You can do the same right now.

The gallery on ruocco.it: the same four prompts, GPT-6.1 Sol and Claude Opus 5.5 side by side, eight demos you can use in the browser.
The gallery on ruocco.it: the same four prompts, GPT-6.1 Sol and Claude Opus 5.5 side by side, eight demos you can use in the browser.

The card machine: a draw

Sol’s terminal is clean and pretty: the card slides through, the screen says processing, then approved, and a paper receipt curls out of the top. A “Swipe again” button replays it and a speed slider changes the pace. Opus drew a terminal with a keypad that types a PIN before it approves, little sound waves for the beep, and a receipt with line items. Both are good, and I call it a draw. The difference is the time: Opus thought for eight minutes on this one and cost $0.91; Sol took under three minutes and cost $0.09.

The voxel pagoda: Sol, just barely

This is the exact kind of test GPT-Next went viral with. Sol’s “floating sanctuary” has a koi pond, a red torii gate, cherry petals falling onto the water and a waterfall pouring into the clouds. Drag the slider to night and the lanterns and the pagoda’s windows glow and the fireflies come out. The page around it, with its little title and its light-and-atmosphere panel, looks like a finished product.

GPT-6.1 Sol's voxel pagoda at night: lanterns and windows glowing, fireflies, the island floating above the clouds.
GPT-6.1 Sol’s voxel pagoda at night: lanterns and windows glowing, fireflies, the island floating above the clouds.

Opus built a detailed pagoda too, and its night is lovely, with stars and glowing windows. But the camera starts too close, so you never see the whole island until you start dragging. This round goes to Sol, just barely. One honest note: in my first review I thought Opus’s night “barely changed”; that was my mistake, because my quick test had not pushed the slider all the way to night. I caught it before publishing and corrected the video and the gallery.

Claude Opus 5.5's pagoda with the slider pushed to night: a proper starry sky, but a camera that starts too close.
Claude Opus 5.5’s pagoda with the slider pushed to night: a proper starry sky, but a camera that starts too close.

The kitchen: Opus strikes back

Opus’s kitchen feels real: daylight through the window, three pendant lights, a long wooden floor, a marble-topped island. Click the fridge and the door swings open and the light comes on inside. Turn the oven knob and the oven glows; the hob burners light up one by one; the tap runs; the kettle starts to steam; the toaster pops. Thirty-two parts you can click, and a Walk mode to move around in first person.

Claude Opus 5.5's kitchen after a few clicks: fridge open, oven glowing, hob lit. Thirty-two parts work.
Claude Opus 5.5’s kitchen after a few clicks: fridge open, oven glowing, hob lit. Thirty-two parts work.

Sol built something closer to a dollhouse: an isometric diorama that is charming in its own way. It works too: the fridge is full of food, the flames are blue, the toast pops, and there is a clock on the wall that shows the real time. One honest note: Sol’s file did not start at first, because it gave two different things the same name (“frame”), which is a syntax error. I renamed one of them, and that is the only change. The gallery card says so, and the untouched original is in the Academy. This round goes to Opus.

The racing game: Sol takes it on looks

Sol made “Circuit Zero”: a neon circuit floating above a city at night, a tunnel of glowing rings, banked turns, five rival crafts, a minimap, a lap timer and a boost meter. In the video it is played by a small bot I wrote that only presses keys, the way a person would: it starts sixth and passes everyone.

GPT-6.1 Sol's Circuit Zero in the middle of a race, driven by a keyboard-only bot.
GPT-6.1 Sol’s “Circuit Zero” in the middle of a race, driven by a keyboard-only bot.

Opus made “Neon Drift”. It is fast and a lot of fun, with drifting and boost pads, but the glow is so strong that much of the screen burns white, especially on the straights. On looks, Sol takes this one.

Claude Opus 5.5's Neon Drift: fast and fun, but the bloom washes out much of the screen.
Claude Opus 5.5’s “Neon Drift”: fast and fun, but the bloom washes out much of the screen.

My verdict, and what it cost

Sol wins the pagoda and the racing game, Opus wins the kitchen, and the card machine is a draw. Now look at the costs, because they are part of the result.

GPT-6.1 Sol did all four demos for about $0.55, on reasoning effort medium, in two and a half to five minutes each. Claude Opus 5.5 cost about $2.98 for the four. It also cost me about $3.85 for nothing: on reasoning effort medium, Opus spent its entire 64,000-token output budget thinking on the kitchen and the racing game and returned no file at all, and the pagoda was cut off halfway. I ran those three again with reasoning effort low, and they came back complete in one and a half to eight minutes. If you use Opus 5.5 for large single-file builds, start at low effort.

So SemiAnalysis is right: Claude’s plan gives you more usage. But in this test, Sol did more with every token. A subscription that gives you five times more tokens is a great deal if the model you get is the one you want. If the cheaper model builds three out of four things just as well or better, the math changes. Which plan would you pay for? Tell me in the comments under the video.

Try every demo, and get the code

Every demo from this test is live on my site, and anyone can use it in the browser for free: ruocco.it/ai-demos/sol-vs-opus-5-5. The racing games run with W, A, S, D and Space, the kitchens work with the mouse, and the pagodas have their time-of-day sliders. All my other test galleries are in the AI Demos section.

The exact prompts, word for word, with model, settings, time and cost, the full source of every demo as a ZIP, the two bots that drive the racing games, and the untouched original of Sol’s kitchen are in the Ruocco Academy, together with step-by-step notes on how each video was made.

Sources

  • Fandu (@mrfanduu) on X, the GPT-Next voxel pagoda and the cost estimate, October 5, 2026; Mirochill (@mirochill) on X, more GPT-Next results.
  • Tibo (@thsottiaux) on X: the 28-day pledge (October 4), Day 1 (October 5) and Day 2.1 (October 6).
  • CSO Online, “OpenAI pulls the plug on GPT 6.1 Astra as agents keep crossing lines”, September 29, 2026.
  • Electronics Weekly, “FTC to accelerate investigation into Anthropic and OpenAI”, October 2026.
  • NPR, “Trump names national intelligence chief Jay Clayton as new AI czar”, October 4, 2026.
  • Reflection AI, “Introducing Beam: Reflection’s 501B open-weight model”, October 5, 2026.
  • Anthropic, “GLM-5.3 and the spread of advanced cyber capabilities”, September 29, 2026.
  • Lyra (@lyraxana), Mirochill, The Hype (@thehypedotnews) and @auricxofficial on X, for the Google items.
  • OpenAI, “Our approach to EU text provenance rules”, and TechCrunch, “OpenAI will start watermarking ChatGPT’s text in the EU”, October 5, 2026.
  • Daniel San (@dani_avila7) and @dfeinition on X, for the Claude Code and Claude Projects updates.
  • SemiAnalysis, “Anthropic Subscriptions Offer 5x+ More Value Than OpenAI”, and Chetaslua (@chetaslua) on X.

A note on honesty

GPT-Next has not been announced by OpenAI; everything about it comes from leakers’ posts, and the cost estimate is theirs, not mine. The “Gemini RSI Model” placeholder is just a name in a log. My test compares the models you can actually buy today, through the API, one request each, with the settings written above. The verdict is my opinion after using every demo; you can open all eight and judge for yourself.

If you want every new AI model tested hands-on like this, subscribe to VitoMind on YouTube. OpenAI has twenty-six more days of promises, and when GPT-Next lands, it gets these same four prompts.

Leave a Comment