Claude Haiku 5.5 vs GPT-6 Luna: Same Price, One Winner

Published: October 8, 2026 | By Vito Ruocco

Anthropic’s smallest model just built a racing game, a voxel temple and a working kitchen for me, and each one cost about five cents. Claude Haiku 5.5 came out on October 7, and it is ninety percent cheaper per token than Haiku 4.5. That puts it at exactly the same list price as OpenAI’s small model, GPT-6 Luna: ten cents per million input tokens and fifty cents per million output tokens. So I ran the obvious test. I gave both models the exact prompts that Claude Opus 5.5, GPT-6.1 Sol and Claude Sonnet 5.5 had already received in my earlier videos, published all nineteen results on my site, and used every one of them live. One of the two cheap models clearly comes out ahead on quality. The other one is much cheaper in practice. And there is a catch in Haiku’s pricing that almost nobody talked about on launch day.

This is the written version of my latest VitoMind video. Before the test there is plenty to cover: what people built with Haiku 5.5 on its first night, the official benchmarks, a monthly API credit that many Claude Max and Team subscribers do not know they are getting, the independent numbers from Artificial Analysis, and the tokenizer detail in the migration guide that can quietly multiply your bill by five. Then comes the part that matters most: Haiku 5.5 against GPT-6 Luna, on my site, with every cost written down.

If you prefer reading, everything from the video is below, with the sources, the exact numbers, and several details that did not fit into seven minutes.

The Short Version

  • Claude Haiku 5.5 (model id claude-haiku-5-5) was released on October 7, 2026. Anthropic calls it “the cheapest, fastest, and most capable small model we’ve ever released”, built for high-volume work and for sub-agent jobs under Opus 5.5 and Sonnet 5.5.
  • Price: $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens; $0.50 and $2.50 above that line. Haiku 4.5 cost $1 and $5.
  • Benchmarks: OSWorld computer use from 15.7% to 72.4%, Humanity’s Last Exam (no tools) from 10.2% to 45.9%, Terminal-Bench 4.0 from 0% to 39.2%. It beats GPT-6 Luna on every test where Anthropic lists both.
  • First Haiku with an effort setting: low, medium, high, xhigh and max. Medium is the default.
  • Launch-day extras: Claude Max 5x subscribers get a $100 monthly API credit, Max 20x get $200, Team plans up to $500. Sonnet 5.5 cache reads are now half price, and the Python and TypeScript SDKs add computer use and browser use in beta.
  • Artificial Analysis gives Haiku 5.5 a 43 on its Intelligence Index, ahead of Gemini 3.8 Flash (41) and GPT-6 Luna (38), but at max effort it uses about 162,000 output tokens per task, roughly three times Luna.
  • The catch: a new tokenizer makes the same text about 30% more tokens, so a 90,000-token Haiku 4.5 prompt becomes about 117,000 tokens and crosses into the five-times-higher price tier.
  • My test: Haiku wins the voxel pagoda and the racing game, Luna wins the kitchen, the card machine is a draw, and both lose the zombie game to Sonnet 5.5. Haiku used about eight times more tokens than Luna, so on my prompts it cost about nine times more, and still about fifteen times less than Opus 5.5.

What People Built With Haiku 5.5 on Launch Night

The first thing I do with any new model is look at what the community made in the first few hours, because that tells you more than any chart. With Haiku 5.5, the first night was surprisingly strong.

An isometric dungeon game in one hour

Leon Lin (@LexnLin) posted an isometric action game made with Haiku 5.5: a hero in a red cape, a forest path, a sunken court full of little crab monsters, five Lightseeds to collect and a boss called the Gatewarden. It looks like a small, polished level from a real indie game, with soft lighting and fog. His comment was simple: “impressive as hell for Haiku 5.5 with one hour of work”, and it used about one percent of his five-hour Claude limit. That second number is the interesting one. A small model is not only cheaper per token; on a subscription it also eats much less of your usage window.

Motion design in fifteen minutes

@cherry_mx_reds shared a short motion-design reel built with Haiku: a blob that morphs into a star, kinetic typography with a variable-weight font, a wall of tiles flipping in perspective, and a punchy blue title card at the end. Fifteen minutes, according to the author. This is the kind of job where a small, fast model makes sense: lots of small iterations, each one cheap.

Haiku 5.5 against Grok 4.7

@notjazii gave the same prompt to Haiku 5.5 and to xAI’s Grok 4.7, both at the highest reasoning available. Haiku took 27 minutes and cost about 80 cents in Claude Code; Grok took 35 minutes and cost about $4.10 in Grok Build. The two runs happened in different tools, so it is not a clean model-to-model comparison: the harness matters a lot for agent work. But the results are worth a look. Haiku’s pixel-art night town has a flying dragon, koi in the river and a red bridge; Grok’s version is calmer and simpler.

The zebra and the rocket: Haiku 5.5 vs GPT-6 Luna

Two more posts compared Haiku directly with GPT-6 Luna, OpenAI’s small model. @AI_Screening asked both for a zebra running across the savanna in Three.js. Haiku’s zebra has properly mapped stripes, a convincing gallop and some depth in the landscape behind it; Luna’s is flatter.

@Bhavani_00007 asked both for a rocket launch scene at max effort: Haiku 5.5 cost $1.30, GPT-6 Luna $1.50, about two minutes each. Here I disagree with the author’s “Haiku cooked Luna”: to my eye Luna’s rocket is sharper, with a real launch tower and a cleaner flame. That is exactly why I wanted to run my own test instead of trusting highlight reels. It is not a one-sided story.

What Anthropic Actually Announced

Anthropic’s announcement page positions Haiku 5.5 very clearly. It is designed for high-volume, cost-sensitive tasks: summaries, compaction, database queries and classification. It “pairs well with Opus 5.5 and Sonnet 5.5 as a subagent on coding work”, and since it is Anthropic’s fastest model at standard speed, it is also aimed at live customer support and browser use. In other words: the big model plans, Haiku does the many small jobs.

The benchmarks

The jump from Haiku 4.5 is enormous. These are Anthropic’s own numbers:

  • OSWorld 2.1 (computer use, offline subset): 72.4%, up from 15.7%. GPT-6 Luna: 48.9%. Sonnet 5.5: 83.9%.
  • Humanity’s Last Exam: 45.9% without tools (from 10.2%) and 57.4% with tools (from 18.7%).
  • Terminal-Bench 4.0: 39.2%, from 0.0%. GPT-6 Luna: 16.4%. Sonnet 5.5: 70.6%.
  • FrontierCode 1.1: 46.4% against 42.4% for Luna.
  • Chartography (visual reasoning): 46.4%, from 6.4%. Luna: 29.1%.
  • GDPval-AA and AA-Briefcase (knowledge work): 1620 and 1578, against 735 and 614 for Haiku 4.5 and 1437 and 1336 for Luna.

On every row where Anthropic lists both small models, Haiku 5.5 is ahead of GPT-6 Luna. Anthropic is also honest about the ceiling: its own Terminal-Bench chart shows that Sonnet 5.5 and Opus 5.5 “remain better choices for complex agentic coding tasks”, while Haiku suits narrowly scoped tasks that used to be too expensive.

The first Haiku with an effort setting

Haiku 5.5 is the first Haiku-class model with adjustable effort: low, medium, high, xhigh and max, with medium as the default. You choose between cheaper and smarter per request, and at lower settings the model can skip thinking entirely on simple requests. Keep this in mind, because in my own test the effort setting turned out to matter more than anything else.

Pricing, and the 100K line

For prompts up to 100,000 tokens, Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens. Cache reads are $0.01 per million and five-minute cache writes $0.125. Above 100,000 tokens everything goes up five times: $0.50 input, $2.50 output, $0.05 cache reads. For comparison, Haiku 4.5 cost $1 and $5, and Sonnet 5.5 costs $2 and $10. Anthropic says about 90% of requests to its previous Haiku model were under the 100,000-token line, and that on average Haiku 5.5 costs around 75% less to run than Haiku 4.5, a figure that already accounts for the tokenizer change I will come back to below.

Anthropic's Haiku 5.5 page: the new monthly API credit for Max and Team subscribers, highlighted live in the video
Anthropic’s Haiku 5.5 page: the new monthly API credit for Max and Team subscribers, highlighted live in the video

The Part Almost Nobody Talked About: Free API Money for Max and Team

Further down the same page there is an announcement that got very little attention on launch day. Anthropic is rolling out a monthly API credit for Claude subscribers, for use on the Claude Platform: Max 5x users get $100 a month, Max 20x users get $200, and Team subscribers get up to $500, pooled across the team. The credits can be used on any model, and Anthropic says they are meant for experimenting with tools, apps and agents that call the API.

If you already pay for Max, this is real money for the API that you were not getting before, and with Haiku 5.5 at ten cents per million input tokens, $100 goes a very long way.

Two more changes landed on the same day. Sonnet 5.5’s cache reads are now 50% cheaper, $0.10 per million tokens instead of $0.20; since cache reads are a large share of what agents consume, Anthropic estimates this makes Sonnet 5.5 around 20% cheaper on most agentic work. And the Claude Python and TypeScript SDKs now support computer use and browser use in beta, which fits Haiku’s speed and price.

What customers measured

Anthropic’s page also quotes early customers. These are vendor-selected quotes, so read them as such, but two numbers stand out. HubSpot says Haiku 5.5 got the best score it has ever seen on its CRM evaluation suite, 92.8% averaged over three runs, and was the fastest model to finish an audit task with the highest hit rate. Asana measured over a 30% reduction in latency for task completions and up to 2.5 times faster inference per agent turn compared with the model it uses today. AlphaSense, Box, Rogo and Cognition report similar gains over Haiku 4.5.

The Independent Numbers: Smart, but Token Hungry

Artificial Analysis tested Haiku 5.5 on launch day. At max effort it scores 43 on the Artificial Analysis Intelligence Index, up 26 points from the last Haiku a year ago. That places it slightly ahead of GLM-5.3 Flash (42), Gemini 3.8 Flash (41) and GPT-6 Luna (38), comparable to Kimi K3 (44), a 2.8-trillion-parameter open-weights model, and 13 points behind Claude Sonnet 5.5 (56).

Now the warning. At max effort, Haiku 5.5 uses about 162,000 output tokens per Intelligence Index task, roughly three times GPT-6 Luna at max (about 50,000), and more than Opus 5.5 at max. Moving from xhigh to max adds two points for about 1.8 times the tokens. At similar intelligence it is also hungrier: Haiku 5.5 on high scores 38 with about 55,000 tokens per task, the same score Luna reaches with about 50,000, and the gap widens at lower effort settings. So yes, Haiku is cheap per token, but it spends a lot of them.

There is good news in the same report. Haiku 5.5 has lower factual knowledge than bigger models, as you would expect, but it is more willing to admit when it does not know: its hallucination rate on AA-Omniscience is 40%, against 55% for Gemini 3.8 Flash and 77% for GPT-6 Luna. Artificial Analysis also notes that its AutomationBench score (35%) is probably understated, because a safety issue made the pre-release model over-refuse.

The Hidden Catch: The 100K Cliff

The detail that can actually hurt your bill is in the migration guide, and Developers Digest explained it well. Haiku 5.5 uses a new tokenizer, the same family as Sonnet 5.5 and Opus 5.5, and “the same input text produces approximately 30% more tokens” than on Haiku 4.5. The 100,000-token price line is counted in the new tokens.

So a prompt that was 90,000 tokens on Haiku 4.5 becomes roughly 117,000 tokens on Haiku 5.5. It crosses the line, and the whole request is billed at the higher tier, five times the price. It is still cheaper than Haiku 4.5 per token, but the “90% cheaper” headline only holds if you stay under the line. If you replay long conversations or big documents into a small model, trimming the context and using prompt caching matter more than the sticker price.

The migration guide has other breaking changes worth knowing before you switch: the model id has no date suffix; fixed thinking budgets return an error and are replaced by adaptive thinking with an effort setting; non-default temperature, top_p and top_k are rejected; assistant prefill is no longer accepted; and the computer-use tool moves to a new toolset version. Recount your tokens and re-test your prompts before moving production traffic.

My Test: Haiku 5.5 vs GPT-6 Luna, Live on ruocco.it

Benchmarks are useful, but I wanted to see what these two cheap models actually build. Everything is on my site, and in the video I use it exactly like a visitor would: the ruocco.it home page, “All demos”, today’s gallery at the top of the list, and then every demo, one by one.

The gallery on ruocco.it: 19 demos, with the cost of every run at the top
The gallery on ruocco.it: 19 demos, with the cost of every run at the top

The setup

Claude Haiku 5.5 and GPT-6 Luna received the exact words that Claude Opus 5.5 and GPT-6.1 Sol got in my October 6 test (a voxel pagoda, a racing game, a working 3D kitchen and an SVG card machine), plus the zombie survival game prompt that Claude Sonnet 5.5 got on September 29. One request each through OpenRouter, nothing fixed by hand. The only line I add to every prompt asks for one self-contained HTML file with Three.js loaded from a CDN.

What it cost

  • Claude Haiku 5.5: $0.19 for the four Opus/Sol prompts, $0.08 for the zombie game.
  • GPT-6 Luna: about $0.02 for the four prompts, under a cent for the zombie game.
  • GPT-6.1 Sol: $0.55 for the same four (October 6).
  • Claude Opus 5.5: $2.98 for the same four (October 6).
  • Claude Sonnet 5.5: $0.93 for the zombie game (September 29).

There is a story behind Haiku’s number. On its default effort, medium, three of my five runs ran out of output. Two of them, the kitchen and the zombie game, spent all 128,000 output tokens thinking and returned no file at all; the racer ran out in the middle of the file. That was $0.19 for nothing. I ran those three again on effort low, and they came back complete, between six and nine minutes each. The zombie game still hit the 128,000-token limit on low, so a second request asked the model to continue exactly from the last character, and it finished the file for two more cents. Luna, on medium, answered every prompt in one to two and a half minutes using 8,000 to 17,000 tokens. This matches what Artificial Analysis measured: Haiku thinks a lot.

The voxel pagoda

Claude Haiku 5.5's Sakura Pagoda at night: lit windows, fireflies and stars
Claude Haiku 5.5’s Sakura Pagoda at night: lit windows, fireflies and stars

Haiku’s pagoda is really good for five cents. Cherry trees drop petals, there is a red torii gate and a waterfall pouring off the edge of the floating island. When I drag the time slider to night, the pagoda’s windows light up, fireflies come out and the sky fills with stars. The only weak point is the camera, which starts too close: you need to zoom out to see the whole island.

GPT-6 Luna's Moonlit Pagoda: tidy, but smaller and simpler
GPT-6 Luna’s Moonlit Pagoda: tidy, but smaller and simpler

Luna’s pagoda, from the same prompt, is tidy: torii, stone steps, a lantern, falling petals. But it is smaller, simpler and has less going on. The pagoda goes to Haiku.

The racing game

Claude Haiku 5.5's Neon Apex: a bright neon city, boost meter, minimap and five rivals
Claude Haiku 5.5’s Neon Apex: a bright neon city, boost meter, minimap and five rivals

Haiku called its game Neon Apex. I press Start race, and you get a bright neon city, glowing arches over the track, a boost meter, a minimap and five rivals that actually race and overtake. In the video the car is driven by a small bot that only sends keyboard events, so the game itself is untouched. There is one visible bug: the “GO!” sign from the countdown never leaves the screen.

GPT-6 Luna's Neon//Apex: a great title screen, and a game that never starts
GPT-6 Luna’s Neon//Apex: a great title screen, and a game that never starts

Funny detail: both models picked the same name. Luna’s Neon//Apex has a great title screen, and then nothing happens. The script declares a function called start twice in the same module, so the browser refuses to run the game at all. Racing goes to Haiku.

The kitchen

Claude Haiku 5.5's kitchen: the fridge opens, but most clicks never arrive
Claude Haiku 5.5’s kitchen: the fridge opens, but most clicks never arrive

Haiku’s kitchen is the better-looking one: warm light, wooden cabinets, pendant lamps, a plant, a fruit bowl. I click the fridge and it swings open, light on, food inside. I click the kettle: nothing. The tap: nothing. The hob: nothing. The strange part is that the animations are all in the code. Haiku registered 27 working parts, and calling the page’s own toggle function from the console makes the kettle steam and the tap run. But most of the objects never catch the mouse click, so for a normal visitor they simply do not work.

GPT-6 Luna's kitchen: flatter light, but the cabinets open, the hob lights and the kettle steams
GPT-6 Luna’s kitchen: flatter light, but the cabinets open, the hob lights and the kettle steams

Luna’s kitchen looks flatter, with less interesting lighting, but the clicks work: a cabinet swings open, a hob burner lights up, the kettle steams. For a prompt whose whole point is “everything works when you click it”, that wins. The kitchen goes to Luna.

The card machine

Both models produced a working isometric SVG card terminal: the card goes in, a “Processing” state, an approved tick and a paper receipt that curls out, on a loop, with a replay button and a speed slider. Haiku’s has keys that press as the PIN is typed; Luna’s has neat sound-wave rings for the beep. Neither is better in a meaningful way. A draw.

The zombie game, where the small models hit their limit

The zombie prompt is the hardest one: a round-based first-person survival game in a boarded-up farmhouse, with windows the zombies tear open, boards you rebuild for points, wall weapons, a mystery crate, perk machines and an upgrade forge. In Haiku’s version the core loop works. Zombies tear the boards off the windows, I shoot them through the gap and the points come in. But the starting pistol is never drawn on screen (the code only builds the weapon model when you get a new gun), and zombies walk right into the camera. Luna’s version is darker and has the same problems: no weapon model and zombies glued to the camera.

Claude Sonnet 5.5's Last Light from the same prompt: a gun in your hands, a proper house, a real game
Claude Sonnet 5.5’s Last Light from the same prompt: a gun in your hands, a proper house, a real game

Here is Sonnet 5.5 on the same prompt from September 29: a gun in your hands, a proper house with lighting, a glowing mystery crate, perks. A real game. This is why the big models still exist, and why Anthropic itself says Haiku is for narrow tasks, not for building a whole game in one shot.

My Verdict

On quality, Haiku 5.5 wins: it takes the pagoda and the racer, Luna takes the kitchen, the card machine is a draw, and on the big game both lose to Sonnet 5.5. Luna’s racer not running at all weighs heavily.

On cost, the picture flips. The two models have the same price per token, but Haiku used about eight times more tokens, so on my prompts it cost about nine times more than Luna. It is still about fifteen times cheaper than Claude Opus 5.5 on the same four prompts, and its pagoda in particular is not far from what Opus produced.

So here is how I would use it. For sub-agents, summaries, classification, extraction and quick front-end jobs, Haiku 5.5 is a great deal, especially if you have the new Max credit to spend. Keep it on low or medium effort, watch the 100,000-token line, and do not expect it to replace Sonnet or Opus on big, multi-part builds. If raw cost per task is all that matters and you can live with more bugs, GPT-6 Luna is remarkably cheap.

Try Every Demo, and Get the Code

All nineteen demos are live and free to use in your browser: Haiku 5.5 and GPT-6 Luna side by side with Claude Opus 5.5, GPT-6.1 Sol and Claude Sonnet 5.5 on the same prompts. Each card shows how long the model took and what it cost, and notes what works and what does not.

Open the gallery: Claude Haiku 5.5 vs GPT-6 Luna

The exact prompts, word for word, with model, settings, time and cost, the full source of every demo as a ZIP, and the bots I used to play the racer and the zombie game are in the Ruocco Academy. All the other galleries are on the AI demos page.

Sources

A Note on Honesty

Every demo in my gallery is exactly what the model wrote, with no hand edits; the bugs described above are left in on purpose so you can see them. The racer and zombie gameplay in the video is driven by small bots that only send keyboard and mouse events; for the zombie game the bot runs on a local copy that adds one read-only line exposing the game state, and god mode is on. The community demos belong to their authors and were made in their own tools, so they are not directly comparable to each other or to my runs. Customer quotes come from Anthropic’s page and were selected by Anthropic, and the Max and Team API credit is rolling out during the launch week according to Anthropic. Prices and scores are those published on October 7 and 8, 2026.

If you want every new AI model tested hands-on like this, subscribe to @VitoMind on YouTube.

Leave a Comment