HomeLearn

Claude Haiku 5.5 Pricing: API Rates, Costs & Usage

Claude Haiku 5.5 pricing starts at $0.10 per 1M input tokens and $0.50 per 1M output tokens. See cache, long-context, batch, and real-world costs.

Bhavyadeep
Written by
Bhavyadeep
Priyanka Singh
Reviewed by
Priyanka Singh
Last updated: 
October 9, 2026
0
 min read
Select Emergent as your Preferred news source
Table of Contents

TL;DR

  • Claude Haiku 5.5 pricing starts at $0.10 per 1M input tokens and $0.50 per 1M output tokens for requests up to 100K tokens.
  • Requests exceeding 100K tokens cost $0.50 per 1M input tokens and $2.50 per 1M output tokens.
  • Cache reads cost $0.01 per 1M tokens, and 5-minute cache writes cost $0.125 per 1M tokens for requests with 100K tokens.
  • The Claude Haiku 5.5 model offers a 1M tokens context window, can generate up to 128K output tokens, and uses adaptive thinking with an effort parameter.
  • Haiku 5.5 is about 75% cheaper to operate than Haiku 4.5 on average, after accounting for pricing and tokenizer differences.
  • Batch requests offer a 50% discount on input and output tokens. This makes Haiku 5.5 particularly attractive for high-volume workloads that don't require immediate responses.

‍

Claude Haiku 5.5 is positioned as Anthropic’s smallest, fastest, and most cost-effective model for high-volume AI tasks. Its use cases include classification, extraction, summarization, routing, browsing, customer support, and subagents.

While the advertised price is low, the real costs depend on how long your prompt is, how much output there is, caching, and the type of processing you choose, standard or batch. Here’s what you should know before you build on Claude Haiku 5.5.

Claude Haiku 5.5 Pricing At a Glance

Claude Haiku 5.5 offers two price levels based on prompt length. The lower prices apply to prompts up to 100K tokens, while those over 100K tokens cost more.

Token type Prompts up to 100K Prompts over 100K
Input $0.10 / 1M $0.50 / 1M
Cache reads $0.01 / 1M $0.05 / 1M
5-minute cache writes $0.125 / 1M $0.625 / 1M
1-hour cache writes $0.20 / 1M $1.00 / 1M
Output $0.50 / 1M $2.50 / 1M

About 90 percent of input requests for Haiku 4.5 were in the low prompt-size bucket. With improvements to its tokenizer, Anthropic estimates that Haiku 5.5 is about 75% cheaper to operate than Haiku 4.5 on average, after accounting for pricing and tokenizer differences.

Thus, Haiku 5.5 is especially appealing for use cases involving tens or even millions of relatively small input requests.

How Prompt Caching Changes Claude Haiku 5.5 Costs

Prompt caching can reduce the cost of repeatedly sending the same context.

For example, an application might repeatedly send:

  • System instructions
  • Product documentation
  • Knowledge bases
  • Policies
  • Reference material
  • Agent instructions

Instead of paying the full input price every time, cached content can be read at a substantially lower rate.

For prompts up to 100K tokens:

  • Input: $0.10 per 1M tokens
  • 5-minute cache write: $0.125 per 1M tokens
  • 1-hour cache write: $0.20 per 1M tokens
  • Cache read: $0.01 per 1M tokens

That means a cached token read costs only 10% of the standard input price. For applications that repeatedly reuse large blocks of context, caching can therefore have a meaningful effect on the final bill.

The 100K Prompt Pricing Threshold

One of the most important details in Claude Haiku 5.5 pricing is the 100K-token threshold.

Requests with prompts up to 100K tokens receive the lowest rates. Once the prompt exceeds 100K tokens, the higher rate applies to the request. Anthropic lists the higher rates as $0.50 per 1M input tokens and $2.50 per 1M output tokens.

Prompt size Input price Output price
Up to 100K tokens $0.10 / 1M $0.50 / 1M
Over 100K tokens $0.50 / 1M $2.50 / 1M

This distinction matters when you're building applications that send very large documents, codebases, conversations, or knowledge bases to the model.

If you don't need the entire context on every request, summarizing or retrieving only the relevant information can help keep usage within the lower pricing tier.

Batch Processing: How to Cut Costs Further

Claude Haiku 5.5 supports Batch processing with a 50% discount on input and output tokens.

That makes Batch particularly useful for work where the user doesn't need an immediate response.

Examples include:

  • Classifying large datasets
  • Summarizing documents
  • Processing support tickets
  • Extracting information from files
  • Generating metadata
  • Running background AI workflows
  • Processing large queues of requests

For example, the standard input rate for prompts up to 100K tokens is $0.10 per 1M tokens. With the 50% Batch discount, that becomes effectively $0.05 per 1M input tokens.

The same calculation applies to output: $0.50 per 1M tokens becomes $0.25 per 1M output tokens.

What Does Claude Haiku 5.5 Actually Cost Per Task?

Token pricing tells you the rate, but it doesn't tell you what an individual application will cost.

Your actual cost depends on:

  • Number of input tokens
  • Number of output tokens
  • Prompt length
  • Cached context
  • Thinking effort
  • Number of model calls
  • Whether you're using Batch processing

Consider a simple application that processes 1,000 requests per month.

Assume each request contains:

  • 5,000 input tokens
  • 1,000 output tokens
  • No caching
  • Standard processing

At the lower Haiku 5.5 rates:

Input cost

5,000 × 1,000 = 5 million input tokens

5 × $0.10 = $0.50

Output cost

1,000 × 1,000 = 1 million output tokens

1 × $0.50 = $0.50

Estimated monthly API cost: $1.00

This illustrates why Haiku 5.5 is aimed at high-volume applications. At relatively modest token volumes, the raw model cost can be very low.

Actual costs will vary depending on token usage, caching, tool calls, and the application architecture.

Claude Haiku 5.5 vs Haiku 4.5 Pricing

The difference becomes particularly clear when comparing the two models.

Model Input, up to 100K Output, up to 100K Cache read
Claude Haiku 4.5 $1.00 $5.00 $0.10
Claude Haiku 5.5 $0.10 $0.50 $0.01

Haiku 5.5 is therefore substantially cheaper on listed rates for the lower prompt-size tier. Anthropic also says it is its fastest model at standard speed and is designed for high-volume workloads.

The comparison isn't purely about price, though. Haiku 5.5 also adds a 1M-token context window, up to 128K output tokens, and adaptive thinking with an adjustable effort setting.

Claude Haiku 5.5 vs Sonnet 5.5 Pricing

Haiku 5.5 sits considerably below Sonnet 5.5 pricing.

Model Input Output Cache reads
Claude Haiku 5.5 $0.10 $0.50 $0.01
Claude Sonnet 5.5 $2.00 $10.00 $0.10

This represents the usual charge for prompts below the relevant price ceiling. Anthropic classifies Sonnet 5.5 as capable of more complex tasks, whereas Haiku 5.5 is geared toward high-volume tasks.

In practice, Haiku 5.5 suits applications that need to make large numbers of model calls at a low cost, while Sonnet 5.5 has greater capacity to justify the extra expense. Take a deep look at the difference between the Sonnet and Haiku here.

claude haiku 55 vs sonnet 55 pricing

Source

Claude Haiku 5.5 Capabilities

Price is only one part of the model's value.

Claude Haiku 5.5 provides:

  • 1M-token context window
  • 128K maximum output
  • Adaptive thinking
  • Adjustable effort
  • Text and image input
  • Text output
  • Fast latency
  • June 2026 knowledge cutoff

The model is particularly suited to classification, extraction, routing, summarization, compaction, customer support, browser use, and subagent workloads.

Anthropic also specifically positions Haiku 5.5 as a useful subagent alongside larger Claude models such as Sonnet 5.5 and Opus 5.5.

Where You Can Use Claude Haiku 5.5

Claude Haiku 5.5 is available through the Claude API and major cloud platforms, including:

  • Claude Platform
  • Amazon Bedrock
  • Google Cloud
  • Microsoft Azure
  • Claude Code

If looking to access one can access it through the model ID:

claude-haiku-5-5

If you're using a cloud provider, pricing and deployment terms can vary, so check the provider's pricing before estimating production costs.

Is Claude Haiku 5.5 Cheap Enough for High-Volume AI Apps?

For many high-volume workloads, yes. The combination of $0.10 input pricing, $0.50 output pricing, inexpensive cache reads, and 50% Batch discounts makes Haiku 5.5 particularly attractive for applications that generate many relatively short requests.

It is especially compelling for workloads where a larger model would be unnecessarily expensive, such as:

  • Classification
  • Routing
  • Data extraction
  • Summarization
  • Document processing
  • Customer support
  • Content transformation
  • Subagent tasks

The model isn't intended to replace Anthropic's larger models for every workload. However, Anthropic clearly says that Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding tasks.

How to Keep Your Claude Haiku 5.5 Costs Low

Four strategies can make a meaningful difference:

1. Keep prompts below 100K tokens when practical

The lower pricing tier is substantially cheaper.

2. Cache repeated context

If your application repeatedly sends the same instructions or reference material, prompt caching can reduce input costs.

3. Use Batch for non-urgent work

Batch processing cuts input and output prices by 50%.

4. Use Haiku for the tasks it is designed for

Don't use an expensive frontier model for simple classification or extraction if Haiku 5.5 can handle the task reliably.

Build AI Apps With Claude Haiku 5.5

A low API price makes experimentation easier, but turning a model into a production application still requires an application layer around it. That means an interface, workflows, and a backend.

That’s where Emergent comes in. With Emergent, a vibe coding tool, you describe an application in plain language and it builds the full-stack product around AI models. Through its Universal LLM Key, Emergent gives you access to models from providers including Anthropic, OpenAI, and Google in one place, so you can wire AI into a complete product, internal tool, or workflow without starting from scratch.

So once you’ve decided which model fits your use case, Emergent is where you turn that choice into a working app.

Start Building on Emergent.

Was this article helpful?
About the writer

Bhavyadeepsinh Rathod is SEO Content Manager at Emergent.sh, where he covers the tools, frameworks, and workflows driving the next era of vibe coding. With 8+ years in tech content marketing, he brings a sharp SEO lens to complex subjects, making Emergent's ecosystem of AI builder tools discoverable for the builders, creators, and teams that need them most. He specializes in making complex topics feel simple, relevant, and easy to act on.

Cta image

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.

  • Production-ready apps
  • Web & mobile apps
  • Deploy in minutes
Try For Free
Share this article:

Frequently Asked Questions

Your Questions, Answered

How much does Claude Haiku 5.5 cost?
Claude Haiku 5.5 costs $0.10 per 1M input tokens and $0.50 per 1M output tokens for prompts up to 100K tokens. For prompts over 100K tokens, the rates increase to $0.50 and $2.50 respectively.
Is Claude Haiku 5.5 cheaper than Haiku 4.5?
Yes. For prompts up to 100K tokens, Haiku 5.5's input/output rates are 90% lower than Haiku 4.5's. Estimated running cost savings are 75%.
How much does Claude Haiku 5.5 cost per 1M tokens?
For prompts up to 100K tokens, Claude Haiku 5.5 costs $0.10 per 1M input tokens and $0.50 per 1M output tokens. Cache reads cost $0.01 per 1M tokens.
Does Claude Haiku 5.5 have a free API?
Anthropic's published Haiku 5.5 documentation provides usage-based API pricing rather than listing a free API tier. Access through Claude.ai subscriptions is a separate product offering from API usage.
Does Claude Haiku 5.5 support prompt caching?
Yes. Cache reads cost $0.01 per 1M tokens for prompts up to 100K tokens. Anthropic also offers 5-minute and 1-hour cache-write options.
Does Claude Haiku 5.5 support Batch processing?
Yes. Anthropic offers a 50% discount on input and output tokens when using the Batch API.
What is the context window for Claude Haiku 5.5?
Claude Haiku 5.5 has a 1M-token context window and supports up to 128K output tokens.
Is Claude Haiku 5.5 good for coding?
Haiku 5.5 can be used for coding and works particularly well as a subagent for larger Claude models. However, Anthropic recommends Sonnet 5.5 and Opus 5.5 for more complex agentic coding tasks.
Is Claude Haiku 5.5 worth the price?
High-volume, time-critical tasks will benefit greatly from Claude Haiku 5.5 because of its strong price-to-capability ratio. The inexpensive token prices, caching capabilities, Batch discount, 1M context window, and thinking abilities render the model quite ideal for high-volume tasks.
Start Building
on Emergent today
Try Emergent

https://api.linear.app/graphql