4Geeks chosen to deliver AI education in the Bahamas alongside Harvard, Oxford, and Columbia.See more
8 min read

Claude Haiku vs Sonnet: When the Small Model Is the Right Call

Claude Haiku 5.5 vs Sonnet 5.5 for API builders: prices side by side, one job priced on each model, the effort setting and when Sonnet or Opus wins.

Claude Haiku vs Sonnet comes down to the shape of the job. Claude Haiku 5.5 is Anthropic's small, fast model for narrow tasks you run thousands of times, like sorting tickets or pulling one field out of a document. Claude Sonnet 5.5 is the everyday workhorse for longer tasks that need more judgment, like fixing a bug across several files. Pick Haiku when the task is short and repeated. Pick Sonnet when it's long and open ended.

The gap is mostly price. On Anthropic's list prices, Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. Sonnet 5.5 costs $2 and $10. That's 20 times cheaper on short prompts. Haiku 5.5 launched on October 7, 2026. It's also the first Haiku with an effort setting, so you can ask the cheap model to think harder when a task needs it.

Prices, model names and dates verified on anthropic.com and the Claude Platform docs on October 8, 2026.

The short answer: which model to call

Say you're building with the Claude API, the way your code talks to Claude. Here's a starting point based on what Anthropic recommends:

Your task looks likeStart withWhy
Classifying, routing, extracting fields, short summariesHaiku 5.5Fast, cheap, and the work is narrow
Many small subtasks handed out by a bigger modelHaiku 5.5Anthropic built it to work as a helper model
Live chat, support replies, browser and computer use where speed mattersHaiku 5.5Anthropic's fastest model at standard speed
Everyday coding, data analysis, writing, agents that use toolsSonnet 5.5Better judgment at a moderate price
Long, open ended agentic coding that runs for a long timeOpus 5.5Anthropic says Sonnet or Opus is the better choice here

A token is a small piece of text, often part of a word. You pay for the tokens you send (input) and the tokens the model writes back (output). That's why the shape of the job matters so much: a narrow task with a short answer barely costs anything on Haiku.

What changed with Haiku 5.5

Anthropic released Claude Haiku 5.5 on October 7, 2026. It joined Opus 5.5 (September 22) and Sonnet 5.5 (September 28) in the Claude 5.5 family. Here's what's new:

  • Price. Anthropic says it costs around 75% less to run than Haiku 4.5 on average. Per request, it's priced 90% lower for prompts up to 100,000 tokens and 50% lower above that.
  • Effort control. It's the first Haiku where you choose how hard the model thinks. More on that below.
  • Same size limits as the bigger models. The docs list a 1M token context window and 128K max output for Haiku 5.5, Sonnet 5.5 and Opus 5.5 alike.
  • Where you can use it. On the Claude Platform as claude-haiku-5-5, plus Amazon Web Services, Google Cloud and Microsoft Foundry. It's also in Claude Code, and Free, Pro, Max, Team and Enterprise users can pick it in the Claude apps on web, iOS and Android.

One small catch sits in Anthropic's fine print. Haiku 5.5 has a new tokenizer, the part that splits text into tokens, so it uses slightly more tokens per task than Haiku 4.5. The 75% figure already accounts for that.

Haiku 5.5 vs Sonnet 5.5 prices side by side

These are Anthropic's published prices per 1 million tokens, as listed on its Haiku 5.5 launch page and pricing page on October 8, 2026.

Price per 1M tokensHaiku 5.5 (prompts up to / over 100K)Haiku 4.5Sonnet 5.5
Input tokens$0.10 / $0.50$1.00$2.00
Output tokens$0.50 / $2.50$5.00$10.00
Cache reads$0.01 / $0.05$0.10$0.10
Cache writes$0.125 / $0.625$1.25$2.50

Two things stand out. First, the gap shrinks on long prompts. Over 100,000 tokens, Haiku 5.5 is 4 times cheaper than Sonnet 5.5, not 20. Second, Sonnet got cheaper too. On the same day, Anthropic cut Sonnet 5.5 cache reads from $0.20 to $0.10 per million tokens. It says that makes Sonnet 5.5 around 20% cheaper on most agentic work. Cache reads are what you pay when the model reuses text it has already seen, like a long system prompt.

For reference, Opus 5.5 lists at $4 input and $20 output.

One job, priced on each model

Say you run a support inbox and want every incoming ticket labeled as billing, bug or other. You get 100,000 tickets a month. Each request sends about 2,000 tokens (your instructions plus the ticket) and gets back about 100 tokens.

That's 200 million input tokens and 10 million output tokens a month. At list prices:

ModelInput costOutput costMonthly total
Haiku 5.5$20$5$25
Haiku 4.5$200$50$250
Sonnet 5.5$400$100$500
Opus 5.5$800$200$1,000

This is an illustration, not a quote. Your real bill depends on your prompts, and the model's thinking counts toward its output tokens, so a higher effort setting costs more. Anthropic also offers a 50% discount through its Batch API for work that doesn't need an instant answer, and prompt caching can cut the input side further.

Still, the shape is clear. For a narrow labeling job, paying 20 times more for Sonnet only makes sense if Haiku gets the labels wrong often enough to hurt. Test that on a few hundred real tickets before you decide.

Where Haiku falls short

Anthropic is direct about this. Its launch post says Sonnet 5.5 and Opus 5.5 "remain better choices for complex agentic coding tasks." Haiku 5.5, it says, "is best suited to more narrowly scoped tasks."

Its own benchmark numbers show the gap. These are Anthropic's reported results:

  • Terminal-Bench 4.0 (multistep work in a command line): Haiku 5.5 scored 39.2%, Sonnet 5.5 scored 70.6%.
  • OSWorld 2.1 (operating a real computer, offline subset): Haiku 5.5 scored 72.4%, Sonnet 5.5 scored 83.9%. Haiku 4.5 scored 15.7% on the same test.

So Haiku 5.5 is a huge jump over the old Haiku, but it isn't a Sonnet replacement for long coding sessions. If a task needs the model to plan, try something, notice it failed and change course many times, start with Sonnet or Opus.

The effort setting, explained

Effort works like telling a coworker how much time to spend. "Give me a quick answer" versus "take your time and get it right." Haiku 5.5 supports five levels: low, medium, high, xhigh and max.

Here's how Anthropic's docs suggest using it on Haiku 5.5:

  • Medium is the default. Start there for most work, including agentic coding.
  • Low is the cheapest and fastest. Use it for chat, short tool tasks and simple, high volume requests. In long agent prompts, the model is more likely to skip a check at low.
  • High is for knowledge work, longer agent tasks and strict instruction following.
  • Xhigh and max only where your tests show a quality gain. At that point, compare against Sonnet 5.5 on quality, cost and speed.

Sonnet 5.5 defaults to high. That's one reason to set effort on purpose instead of comparing the two models at their defaults.

A simple routing pattern: the big model plans, Haiku does the small calls

You don't have to pick one model for everything. Anthropic's docs describe two common ways to combine them:

  1. Orchestrator and workers. A stronger model, like Opus or Fable, breaks a big job into pieces. Haiku handles the many small pieces in parallel. Most of your tokens get billed at Haiku's rate.
  2. Executor and advisor. Haiku does the routine work and hands off only the hard decisions to a bigger model.

Here's what that looks like for a research task. The big model reads the request and decides it needs ten documents summarized and one figure pulled from each. Ten Haiku calls do the summaries and the extraction. The big model then writes the final answer from those ten short notes. You paid the expensive rate only for the planning and the final write up.

Start simple. Use one model, measure where it fails, and add routing only where it saves real money or fixes a real problem.

Go deeper

Want to build your first AI agent without guessing?

AI Flex lets you learn AI at your own pace, with guided paths on agents and automation, an AI tutor and mentors. No fixed cohort, and you can cancel anytime.

Frequently Asked Questions