Small models do the unglamorous work in most AI systems: sorting messages, summarizing calls, pulling one field out of a document, answering the same customer question for the hundredth time. That work runs at volume, so the per-token price matters more than anywhere else. Anthropic’s new Haiku 5.5 cuts that price sharply for ordinary-length prompts and adds an effort dial, which is good news for Alberta businesses running chatbots and lead-intake agents.
What Anthropic released
“Claude Haiku 5.5 is designed for high-volume, cost-sensitive tasks.”Anthropic, Introducing Claude Haiku 5.5, October 7, 2026
Anthropic lists quick, repetitive workloads such as summaries, compactions, database queries and classification. It also says Haiku 5.5 pairs well with Opus 5.5 and Sonnet 5.5 as a subagent on coding work, and calls it its fastest model to date, suited to live customer support and browser use. Anthropic is also clear about the limit: it says Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding tasks.
The benchmark figures are Anthropic’s. On OSWorld 2.1 (offline subset), a computer-use test, Anthropic reports 72.4% for Haiku 5.5, against 15.7% for Haiku 4.5 and 48.9% for OpenAI’s GPT-6 Luna. Anthropic says Haiku 5.5 costs around 75% less to run than Haiku 4.5 on average. Its footnote explains that the per-token price is 90% lower for requests up to 100,000 tokens and 50% lower above that, and that Haiku 5.5’s updated tokenizer uses slightly more tokens per task.
Prices and the 100K-token threshold
| Token type | Prompts up to 100K | Prompts over 100K |
|---|---|---|
| Input | $0.10 | $0.50 |
| Output | $0.50 | $2.50 |
| Cache reads | $0.01 | $0.05 |
| Cache writes | $0.125 | $0.625 |
| Batch input / output | $0.05 / $0.25 | $0.25 / $1.25 |
The threshold is the detail to plan around. Anthropic’s pricing docs say a request whose prompt is over 100,000 tokens pays the higher prices, even when part of the prompt is a cache hit. The docs also say that specifying US-only inference adds a 1.1x multiplier on Claude 4.6 and later models, while global routing, the default, uses standard pricing. Anthropic says prompts up to 100,000 tokens made up around 90% of requests to its previous Haiku model.
Cheaper Sonnet 5.5 caching and new API credits
On the same day, Anthropic halved the price of Sonnet 5.5 cache reads to $0.10 per million tokens and says this makes Sonnet 5.5 around 20% cheaper on most agentic tasks. Anthropic also says it is rolling out monthly API credits this week: $100 for Max 5x subscribers, $200 for Max 20x and up to $500 for Team, pooled across users. If your heavier agents already run on Sonnet 5.5 with prompt caching, re-check your bill for that change alone.
What Alberta teams should do (our advice)
What follows is our advice, not Anthropic’s. Start with a price check on your highest-volume work: website chatbots, after-hours lead intake, call summaries and ticket triage. As an illustration from the listed rates, a chatbot conversation with 5,000 input tokens and 500 output tokens would cost about $0.00075 on Haiku 5.5, or about $7.50 for 10,000 conversations. That’s our arithmetic, in US dollars, and excludes retries, tools and taxes. Your real token counts will differ, so measure a week of traffic.
Next, test effort levels on your own examples. Lower effort may be enough for routing and classification, while lead qualification may need more. Score answers, not just cost. Keep prompts under 100,000 tokens where you can, and keep your monthly cost cap and alerts in place: a cheaper model makes it easier to scale a loop that shouldn’t be running. Our AI receptionist Alberta page covers call and chat intake, and AI automation for Alberta trades covers the lead-to-quote workflow.
