GPT-7 Bel’s 722 Math Papers + Mistral Large 4 Tested

Published: October 7, 2026 | By Vito Ruocco

OpenAI just released 722 mathematical manuscripts. Not written by mathematicians: written by an internal model that OpenAI has not released, has not named, and that the internet is already calling GPT-7 “Bel”. On the same day Mistral launched the biggest model it has ever built, Mistral Large 4, a one-trillion-parameter open-weight model the company itself nicknames “Le Chonk”. So I did what I always do: I took the four prompts I had given Claude Opus 5.5 and GPT-6.1 Sol the day before, gave Mistral the exact same words, published all twelve results on my site, and opened every one of them live. The price gap is huge. The quality gap is bigger.

This is the written version of my latest VitoMind video. Before the test there is a lot of news: someone rebuilt Photoshop in Rust with Claude Opus 5.5 and you can run it in a browser tab (I did, on camera); Anthropic is opening Claude Mythos 5.1 to security professionals, in the same week Mythos found a bug that attackers exploited within a day; Claude now lives inside Google Docs, Sheets and Slides; Claude Code got a new effort setting for sub-agents; Google released Nano Banana 2.1; OpenAI shipped four things on day two of its 28-day pledge and then reset everyone’s limits anyway; and Meta published six math papers of its own. Then comes the part that matters most: Mistral Large 4 against Opus 5.5 and Sol, on my site, with every cost written down.

If you prefer reading, everything from the video is below, with the sources, the exact numbers, and a few details that did not fit into ten minutes.

The Short Version

  • OpenAI’s math release: a GitHub repository called math with 722 manuscripts in 372 families of results, produced by an unreleased internal model that was pointed at about 4,000 problems. Each result used about three hours of ChatGPT Pro thinking compute on average. Many come with Lean proofs; OpenAI also writes that some unformalized results “could have issues”.
  • PhotoCraft is a clean-room rebuild of Adobe Photoshop in pure Rust, open source, about a week old, with more than eleven thousand stars on GitHub. Forty of its last hundred commits are co-authored by Claude Opus 5.5. It compiles to WebAssembly, so it runs in a browser.
  • Claude Mythos 5.1 can now be requested by verified security professionals through Anthropic’s expanded Cyber Verification Program, in three tiers: Defense, Red Team and Specialized.
  • Mythos found CVE-2026-61500 in the Rejetto HFS file server by chaining two harmless-looking weaknesses. Attackers were probing for it within a day of disclosure.
  • Mistral Large 4: 1 trillion parameters, 49 billion active, natively multimodal, built and served in Europe, open weights at the end of October. Very strong at cyber security, with reasoning off by default.
  • My test: on reasoning high and low, Mistral Large 4 mostly thought until it ran out of output and returned no file. With reasoning off it answered in minutes for a few cents, but its 3D results are broken. Opus 5.5 and GPT-6.1 Sol win all four.

PhotoCraft: Photoshop, Rebuilt in Rust, With Claude

Let’s start with the one that blew my mind. PhotoCraft lives on GitHub at storytold/photocraft, and its own description is direct: an open-source, clean-room reimplementation of Adobe Photoshop in pure Rust. It is dual-licensed under MIT or Apache 2.0, it runs on macOS, Windows, Linux and FreeBSD, and the repository was created at the end of September. When I recorded the video it had more than eleven thousand stars; a few hours later it was past twelve thousand.

The part that matters for an AI channel is in the commit history. Open almost any recent commit and at the bottom of the message you find the same line: Co-Authored-By: Claude Opus 5.5. I counted the last hundred commits through the GitHub API: forty of them carry that exact trailer. The project’s own guide for contributors is literally addressed to “AI agents and contributors”. This is software built the way a lot of software will be built from now on: a small team directing a model that writes, tests and fixes a huge codebase.

A PhotoCraft commit on GitHub: the message ends with Co-Authored-By: Claude Opus 5.5, like 40 of the last 100.
A PhotoCraft commit on GitHub: the message ends with “Co-Authored-By: Claude Opus 5.5”, like 40 of the last 100.

And PhotoCraft is not alone. The same ArtCraft family is rebuilding six more Adobe apps in the same way: VectorCraft for vector illustration (Illustrator), FilmCraft for video editing (Premiere), LightCraft for photo libraries and raw development (Lightroom), PrintCraft for PDFs (Acrobat), EffectCraft for motion graphics (After Effects) and DesignCraft for page layout (InDesign).

I used it live, in a browser tab

The best part is that the whole engine and interface compile to WebAssembly, and the official release ships a web build. So in the video I simply used it, the way you would. I clicked Open and loaded Hokusai’s Great Wave (public domain, which is why I picked it). I took the rectangle marquee and selected the crest of the wave. Then Filter, Distort, Twirl: the preview runs live on the canvas, only inside the selection. A little more angle, OK. The menus are where your hands expect them, with the same names and the same shortcuts as Photoshop.

PhotoCraft's web build: a Twirl filter applied only inside a marquee selection on Hokusai's Great Wave.
PhotoCraft’s web build: a Twirl filter applied only inside a marquee selection on Hokusai’s Great Wave.

Then an adjustment layer: Hue and Saturation. I dragged the hue and the whole print changed colour without touching a single original pixel; hiding the layer brings the original back, showing it applies the change again. Finally the type tool: a headline typed straight onto the canvas, as an editable type layer.

A Hue/Saturation adjustment layer and an editable type layer, in a browser tab.
A Hue/Saturation adjustment layer and an editable type layer, in a browser tab.

A technical note for anyone who wants to try it on a server: PhotoCraft prefers WebGPU and falls back to WebGL2. In my headless recording browser the WebGPU device could not start, so I hid WebGPU and it ran on WebGL2 on the same graphics card without any problem.

Let me be fair about what it is. The developers say it clearly in the README: it is an early alpha, much of Photoshop’s feature surface exists in some form, but it is not a Photoshop replacement for daily professional work yet, and there are gaps in AI features, typography and plugins. Some coverage of the project has claimed it already recreates 90% of Photoshop; the project itself does not claim that. But a free Photoshop in Rust, written with an AI, running in a browser tab? A year ago that sounded like science fiction.

Claude Mythos 5.1 Opens to Security Professionals

Now Anthropic. Mythos is the model family Anthropic kept locked away because it is too good at hacking. Mythos 5.1 shares its weights with the public Claude Fable 5.1, but with much looser cyber and biology guardrails.

Since October 6, verified security professionals can apply for it through the expanded Cyber Verification Program, which now merges the old program and Project Glasswing into three tiers:

  • Defense Access is for defensive work: security operations, incident response, reverse-engineering malware, analyzing and validating vulnerabilities. Security teams at companies, nonprofits, universities, government bodies, critical-infrastructure operators, small security firms, open-source maintainers, and individual researchers with a track record of reported vulnerabilities can apply. Reviews take days.
  • Red Team Access adds authorized penetration testing and red-teaming. This tier is for organizations only; individual researchers are not eligible, and reviews take weeks.
  • Specialized Access has the fewest cyber blocks and is reserved for a small set of verified organizations authorized to test safety-critical systems such as flight operating systems and power grids, reviewed in collaboration with the US government.

The program covers Claude Opus 5.5, Claude Sonnet 5.5, Claude Mythos 5.1 and future models. Anthropic also gives a striking number: between April and July 2026, Project Glasswing partners uncovered at least 129,000 verified software vulnerabilities, and more than 33,000 of the vulnerabilities found were rated critical or high severity. Data retention is required for monitoring misuse, with a zero-data-retention route for some customers.

Why they are being careful: the Rejetto HFS bug

A real example landed the same week. Mythos found a critical authentication bypass in Rejetto HFS, a popular open-source file server, now tracked as CVE-2026-61500. What makes it interesting is how. Mythos noticed two things that look harmless on their own: the key that signs session cookies is derived from JavaScript’s Math.random(), which is not designed for security, and an unauthenticated endpoint leaks raw Math.random() outputs. Chain them, and the exploit published by Horizon3 needs about twelve requests to rebuild the random generator’s internal state, recover the signing key and forge an administrator session cookie, which leads to remote code execution.

According to Security Affairs and VulnCheck’s canary data, probes for the vulnerability started within twenty-four hours of public disclosure. That is the whole debate about cyber-capable AI in one story: the same reasoning that helps a defender find a bug before attackers do is exactly why the model is not handed to everyone.

Two Quick Claude Updates

Claude now works inside Google Docs, Sheets and Slides. In Google Workspace it sits in a sidebar next to your file, reads what you have open, and edits it in place, and you approve each edit before it lands. It also works the other way: paste a Google file link in Claude, or ask for a new doc, sheet or deck, and it opens next to the chat so you and Claude can edit it together.

In Claude Code, version 2.1.292 adds an effort parameter to the Agent tool, so each sub-agent runs at the effort level you ask for. A simple job runs at low effort and burns fewer tokens, the hard part gets high effort, and the cost becomes predictable.

Google: Nano Banana 2.1

Google released Nano Banana 2.1, its new image generation and editing model, with improvements in visual design, mask-based editing, subject consistency across edits, and more natural-looking images. It is available in the Gemini app, Google AI Studio and other Google surfaces.

The early numbers that caught my eye came from AI/ML API, which ran the same prompt five times on Nano Banana 2.1 and on GPT Image 2.5: about 4.4 cents and 11 seconds per image for Nano Banana, against almost 7 cents and 50 seconds for GPT Image. Cheaper, and roughly five times faster. Image quality is a matter of taste, and GPT Image still has its fans, but for anything at scale that speed difference matters.

OpenAI’s 28-Day Pledge: Day Two, and the Reset

On October 5, Tibo, who leads Codex at OpenAI, made a public promise: for 28 days, every day, the team either ships one clear improvement that matters to most Codex and ChatGPT users, or it gives everyone a full usage reset. Day one made GPT-6 Astra and GPT-6.1 Sol about 50% faster by default.

Day two was packed, with four announcements:

  • Auto-review is free. The “Approve for me” permission mode uses a second agent to review risky actions, so you can let long tasks run without approving every step. It no longer counts against your plan, where it could take between 2 and 10 percent.
  • Simpler API tiers. Five paid tiers became three, Build, Launch and Grow, and the top tier now needs $500 of total spend instead of $1,000.
  • A Meetings plugin for the ChatGPT Mac app (beta, Pro and Business) takes notes and summaries, and you can pass that context to Codex.
  • The Decisions API, a real-time classifier that picks the next model, tool or action, up to ten times faster than doing it with GPT-6 Luna.

Then Tibo asked the community a simple question: four updates, or a reset? The community voted reset. In his words, the team shipped “four things that were deemed good to great and some math proofs, but the vote is clear and the community demands a reset”, and he admitted the game “seems” rigged in the reset’s favour. Rules are rules: everyone got the reset, and nobody lost the four updates. Good marketing, and a good day to be a Codex user.

GPT-7 “Bel”? OpenAI’s 722 Math Manuscripts

Those “math proofs” deserve their own section, because they may be the biggest story of the week. On October 6, almost as a side note, OpenAI published a GitHub repository called math with “a broad range of new mathematical results produced by an internal frontier model”.

The numbers from the repository itself: the catalogue contains 722 manuscripts organized into 372 families, where a family groups a principal result with companion arguments, consequences or alternative proofs. The vast majority came from the same procedure with an unreleased internal OpenAI model; over the evaluation, the model was posed approximately 4,000 problems; on average, each result used three hours of ChatGPT Pro thinking compute. OpenAI says it consulted the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study on how to release the material.

OpenAI has not named the model. The community calls it “Bel” and guesses GPT-6.5 or GPT-7, but that is a guess, not an announcement.

There is serious material in there. The repository highlights work on a zero-free region for the Riemann zeta function and a proof of the Hodge Conjecture for CM abelian varieties (a special case, not the general conjecture), plus results in operator algebras and many other areas. Many manuscripts come with Lean formalizations, which means a computer can check the proof instead of anyone having to trust what the model wrote. But OpenAI is also honest about the limits: not all results have Lean proofs yet, and “some of the unformalized results could have issues”, which they promise to fix quickly.

Two results, explained by people outside OpenAI

A physicist, Favio Vázquez, picked one result and explained it in 46 seconds: a puzzle from 1950 about colouring the plane so that no two points exactly one unit apart share a colour. OpenAI’s preprint claims five colours never work, and it ships a Lean proof a computer can check. He also points out that nobody outside has reported running it yet, and that six versus seven colours is still open.

Another creator asked Claude to explain OpenAI’s 23-page proof on the free group factor problem and got back a full 3D animation. The claim is that L(F2) and L(F3), two algebras built from free groups, are the same algebra, which would settle a question open for more than fifty years. The animation itself is labelled “claimed, not yet independently verified”, which is exactly the right attitude.

Be careful with the hype

One viral post ranked the results against a list of famous open problems and announced that the model had fully cracked things like the Hadwiger conjecture, “in two days”. Nobody outside OpenAI has checked that, and OpenAI does not make that claim. It is a claim, not a result. If even a fraction of these 722 manuscripts survive serious review, it is still a huge moment: the first time a lab has pointed thousands of research tasks at a single model and published the output at this scale.

Meta is doing the same

OpenAI is not alone. On October 2, Meta published six math papers written by mathematicians together with Muse Spark, using the normal meta.ai chat interface in Thinking Mode, with no custom research setup. Five of them present answers to previously open research questions, in probability, differential equations, group theory, optimization and algebra. A second group of mathematicians reviewed the work, and every paper clearly marks which passages were drafted by researchers and which by AI. Math is becoming the new benchmark, and the labs are now competing on results rather than on test scores.

Mistral Large 4, “Le Chonk”

Now the model I tested for you. Mistral Large 4 is a 1-trillion-parameter, natively multimodal Mixture-of-Experts model with 49 billion active parameters. Mistral calls it the best open-weight model developed in the US or Europe on aggregated benchmarks, trained on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own European datacenters and deployable end to end from Europe. The preview API is live in Mistral Studio, the weights arrive at the end of October, and the official price is $1.36 per million input tokens and $4.18 per million output tokens (it was half that on OpenRouter during launch week).

Its strongest area is cyber security. In a capture-the-flag speedrun of 19 challenges, with real tool calls and solve times, it solved 18. Given an unknown binary, it reverse-engineered it end to end, concluded it was Cobalt Strike, extracted the indicators and the malware configuration, and wrote a report with a YARA rule to catch future incidents. Mistral says that is a day of work for an analyst, done in twelve minutes. It also reports 82% on vulnerability patching and 93% on Cybench.

For creative coding, the first independent tests were mixed. Dominique Silvestre gave the same 3D prompt to eight models, four open (Mistral Large 4, Kimi K3, GLM 5.3, DeepSeek V4 Pro) and four closed (Claude Opus 5.5, Sonnet 5.5, GPT-6 Astra, GPT-6.1 Sol). On the Eiffel Tower, Mistral was the best open model, but GPT-6 Astra won overall. On the aquarium, Mistral was near the bottom and Opus won. The price story was the headline: his three scenes cost $0.87 on Mistral and $13.83 on Opus.

He also left an important warning: Mistral Large 4 has reasoning off by default. With it off, his Eiffel Tower came out as a black screen. With reasoning on high it became a different model, but slow: 66 minutes for the three scenes. Another tester compared it with Kimi K3 on a rocket scene and called it disappointing. So I tested it myself.

The Test: Mistral Large 4 vs Claude Opus 5.5 vs GPT-6.1 Sol

On October 6 I gave Claude Opus 5.5 and GPT-6.1 Sol four prompts: a voxel pagoda on a floating island with a day-to-night slider, an isometric card payment machine animated in SVG, a 3D kitchen where everything important works when you click it, and a neon hover-racing game with five AI rivals. On October 7 Mistral Large 4 got the exact same words, byte for byte, through OpenRouter. All twelve results are published on my site, and in the video I opened them the way you would: ruocco.it home, All demos, today’s card, the gallery, each demo.

The gallery on ruocco.it: all twelve demos side by side, with what each run cost.
The gallery on ruocco.it: all twelve demos side by side, with what each run cost.

What happened behind the scenes

This matters as much as the demos. With reasoning on high, Mistral Large 4 spent 100,000 output tokens thinking about the card machine and never wrote the file. I stopped the other three high runs after 28 minutes with no answer. On reasoning low, three of the four prompts ran out of a 64,000-token output the same way. When I read its reasoning, the reason was clear: it writes the whole program inside its thinking, then starts over, and runs out of room before it ever answers. Only the pagoda came back on low. One more try of the racing game with a 160,000-token budget ended with a provider error after 100,000 tokens of thinking.

So for the card machine, the kitchen and the racing game I used Mistral’s default: reasoning off. Those answered in one to five minutes, for two to four cents each. (One reasoning-off kitchen attempt stopped after a single sentence because the model tried to call a tool that did not exist; I sent the prompt again.) None of Mistral’s files were edited by hand.

The voxel pagoda

Mistral’s pagoda, on reasoning low, took 13.5 minutes and cost five cents. Cherry trees, a red torii gate, lanterns, clouds under the floating island, a time-of-day panel. At sunset it is honestly nice.

Mistral Large 4's pagoda at sunset: cherry trees, a torii gate, clouds under the island.
Mistral Large 4’s pagoda at sunset: cherry trees, a torii gate, clouds under the island.

But drag the slider to night and the scene goes almost black. The prompt asked for lantern light, glowing windows, fireflies and stars; you get stars and a dark silhouette. Sol’s pagoda from the same prompt glows at night, with lit windows, lanterns and fireflies. That is the gap.

The card machine

Reasoning off: 73 seconds, under two cents. The card slides into the reader, the screen shows Processing, a green tick, and a receipt comes out. It works. But the texts on the little screen overlap, and next to Opus’s version, with a keypad that types a PIN and a receipt with real line items, it looks like a sketch.

The kitchen

Under five minutes, about four cents. On paper it is the most ambitious of Mistral’s files: thirty parts registered as interactive, from the fridge to the drawers, exactly as the prompt asked. Then you look at the room: walls in the wrong places, beams going through the floor. I clicked the fridge, the oven and the tap, and nothing happened. The appliances exist in the code, but they are hidden behind a wall, so every click hits the wall.

Mistral Large 4's kitchen: walls and beams in the wrong places hide the appliances, so clicks do nothing.
Mistral Large 4’s kitchen: walls and beams in the wrong places hide the appliances, so clicks do nothing.

Opus 5.5’s kitchen from the same prompt is a different world: a realistic room with daylight and pendant lights, a fridge that opens with its light on, blue flames on the hob, a kettle that boils.

Claude Opus 5.5's kitchen from the same prompt: 32 parts that work when you click them.
Claude Opus 5.5’s kitchen from the same prompt: 32 parts that work when you click them.

The racing game

Mistral called it Neon Velocity: a title screen, a countdown, a minimap and a speedometer, all as requested. Start the race and the track tips on its side, the camera rolls with it, and the lap counter never moves. The boost cannot work either: the code waits for a key named “space”, but the browser reports the Space key as a blank character, so pressing it does nothing (the arrow keys have the same bug). In the video the game is driven by a small bot that only presses W, A and D, like a player would.

Neon Velocity by Mistral Large 4: the track tips on its side and the lap counter never moves.
Neon Velocity by Mistral Large 4: the track tips on its side and the lap counter never moves.

Here is GPT-6.1 Sol’s Circuit Zero from the same words: banked turns, a tunnel of rings, five rivals racing a real line.

GPT-6.1 Sol's Circuit Zero from the same prompt.
GPT-6.1 Sol’s Circuit Zero from the same prompt.

My verdict, and the cost

The four Mistral files that came back cost fourteen cents together. On October 6, Opus 5.5 cost about $2.98 for its four and Sol about $0.55. But the Mistral runs that thought and returned nothing cost me another dollar, for nothing, so the real bill was $1.15. For creative coding, Opus 5.5 and GPT-6.1 Sol win every single one of the four.

That does not make Mistral Large 4 a bad model. It is a cyber security model with open weights on the way, built and hosted in Europe, and that is where I would use it: security work, data that has to stay in Europe, and anything where you will run it yourself. For one-shot games and 3D scenes, today, it is not in the same league, and its reasoning budget is something you have to manage carefully or you pay for thinking that never becomes an answer.

Try Every Demo, and Get the Code

Every demo from this video is on my site and free to use in your browser: ruocco.it/ai-demos/mistral-large-4-vs-opus-sol. All twelve, side by side, with the cost and time of every run and an honest note on each card about what works and what does not. You can find every other test in the AI demos menu.

The exact prompts, the full source code of every demo as a ZIP, the bot that drives the racing games, and Mistral’s reasoning text from the runs that never became a file are in the Ruocco Academy, together with step-by-step notes on how each video is made.

Sources

An Honesty Note

“GPT-7” and “Bel” are community names; OpenAI has not named the model behind the math release. The mathematical results are OpenAI’s claims at different stages of verification, and the rankings circulating on social media are not OpenAI’s. The PhotoCraft demo in the video uses the project’s official WebAssembly release, run locally; PhotoCraft is not hosted on my site. Mistral’s numbers come from my own runs on October 7, 2026 through OpenRouter, in launch week, and a model this new can change quickly. Every demo on my site is exactly what the model wrote, except where a card says otherwise.

If you want every new model tested hands-on like this, subscribe to @VitoMind on YouTube. It is free, and it really helps a small channel grow.

Leave a Comment