GPT 6.1 Sol vs Sonnet 5.5: Same Prompt, a Third of the Cost — Tested Live

OpenAI answered Sonnet 5.5 in just seven days with GPT 6.1 Sol — “near Astra intelligence for a fifth of the price.” I tested it live, head-to-head against Sonnet 5.5 on the exact same prompts. Here’s what happened.

What Is GPT 6.1 Sol?

GPT 6 Sol came out one week ago and it has already been replaced. The headline from OpenAI’s DevDay says it all: near Astra intelligence for a fifth of the price. Pricing is $2 per million input tokens and $10 per million output — exactly what Sonnet 5.5 costs. Cached input is just $0.10, a 95% discount.

Benchmark Results

On SWE-bench (long software tasks in real codebases), Sol matches GPT 6 Astra at about a fifth of the cost. On business workflow automation, Sol is two points above Opus 5.5 at medium effort for roughly a third of the cost. On Terminal Bench Science at max effort, Sol costs $5.47 per task — Opus 5.5 costs $23.21 and Astra $23.80. Astra still holds the top score at 68.1%.

Artificial Analysis gives Sol 52 points on its intelligence index. Astra has 53 — just one point apart. Cost per task at max effort: 72 cents for Sol vs $3.26 for Astra, less than a quarter. The catch is that it writes 10-30% more tokens than the old Sol, and the sweet spot is low and medium effort. Every demo I ran was on medium.

Community Tests

Within hours of release, the community was testing it. Jazzy gave Sol and Astra the same prompt at max reasoning: Sol took 10 minutes and $1.50; Astra took 25 minutes and $11. Against Opus 5.5: Sol finished in 45 minutes for $7.15; Opus took 1 hour 32 minutes at $14.84.

Atomic Chat asked Sol and Sonnet 5.5 for a 3D keyboard in Three.js — Sol was 30 times cheaper. Harshit compared old Sol vs new Sol side by side in Codex with the same prompt (“Animate a cyclist”) — the difference is striking. Not everything goes Sol’s way: Onyx found Opus 5.5 still wins for turning images into web pages, and Lucky Faraday’s Minecraft test was unimpressed.

My Live Tests

Four prompts, one shot each, medium effort. Everything is live on my site. You can open every demo yourself:

Demo 1 — Mechanical Keyboard: A 3D keyboard assembled piece by piece. Aluminum case, circuit board lighting up, stabilizers, plate, 87 switches row by row, every keycap with its own legend. Jump to any step on a timeline, switch to exploded view, change the lighting effect. Four minutes of generation, 13 cents.

Demo 2 — Ashen Vigil (Dark Fantasy Boss Fight): An original dark fantasy boss fight in the spirit of souls-like games. Chests with increased capacity, soldier guards, lock-on combat, light combos and jump attacks. The boss — the Cinder Warden — has a health bar and shows its next move. Parries, ripostes, ground slams, leaps, and a floor fire in phase two. 60,000 characters of code, five and a half minutes, 21 cents. Ran first time with zero errors at 60 FPS.

Head-to-Head: Sol vs Sonnet 5.5

Yesterday I gave Claude Sonnet 5.5 these prompts. Today I gave GPT 6.1 Sol the exact same words.

Round 1 — 360° Smartwatch Viewer: Sonnet’s result looks more premium with a big watch, clean numbered hotspots, and a flying camera. Sol’s has the same five hotspots, five finishes, three watch faces, and an exploded view. Sonnet took 11 minutes and 92 cents. Sol took under five minutes and 17 cents.

Round 2 — New York Skyline (2,000+ lines): Sonnet built a beautiful sunset scene with boats and the Brooklyn Bridge but hit the output limit and needed a second request. Sol did it in one request — over 3,600 lines with the Empire State Building, Chrysler Building, Statue of Liberty, ferry, yellow cabs, helicopter, and a sky that transitions from dawn to night. Seven minutes and 45 cents against 16 minutes and $1.48. Sol wins this round.

My Verdict

Sonnet 5.5 still has better taste, but Sol gets you 90% of the way in half the time for a third of the price. For everyday work, that’s the one I’d pick.

Other DevDay Announcements

DOTS: Always-on agents powered by GPT-6 Astra, each with its own cloud computer. You give a DOT a responsibility, not a prompt. It figures out next steps, keeps working between conversations, and reaches more than 4,000 apps via integrations in ChatGPT, Slack, or Teams.

Pricing Changes: The $200 Pro plan is open for new signups again but gives roughly half the usage it used to. A new Pro 500 tier offers 25x the Plus allowance with UltraFast access — up to 300 tokens/second in Codex, 8x faster (6x in the API).

Code Security: Scans whole GitHub repos, keeps checking every new commit, and prepares fixes for review — even with your laptop closed. The timing is relevant: Microsoft just disclosed Jade Buffer, an AI agent attacker that hit Azure and wiped 100+ storage accounts in seven minutes, using credentials posted in plain text in a public GitHub issue.

Other Updates: Codex cloud environments with ready-to-go dependencies, Codex CLI redesign with analytics/themes/diagrams/math in terminal, and the Decisions API on GPT-6 Luna for fast, cheap routing in multi-agent systems.

Try the Demos Yourself

Every demo shown in this video is available on my site. Open them, tweak the code, and see Sol in action:

→ GPT 6.1 Sol — Live Demos & Code

Which side are you on — Sol or Sonnet? Drop your answer in the comments. Hit like, subscribe, and I’ll see you tomorrow.

Leave a Comment