Haiku 5.5
Haiku 5.5 benchmarks
Read Haiku 5.5 benchmarks with their test conditions, separate published scores from local results, and design a comparison for your own work.
Haiku 5.5 benchmarks help you choose what to test. They cannot tell you whether a particular answer is safe to ship. A model can perform well on a coding evaluation and still omit the one condition in your refund policy that matters. Start with the published evidence, then translate it into a small evaluation whose failures resemble the failures your users would actually notice.
This page separates manufacturer results, independent evaluations and your own saved comparisons. We have not run a public benchmark suite for this website. The worksheet below is an evaluation design, not a claim that Haiku has achieved a particular result on our infrastructure.
Published Haiku 5.5 benchmarks
Anthropic's October 7, 2026 launch report includes the following results. These are selected manufacturer-reported measurements, with their original benchmark labels. They are not interchangeable percentages of general intelligence.
| Evaluation | Haiku 5.5 | What to preserve when citing it |
|---|---|---|
| GDPval-AA v2.1 | 1620 | Rating scale and evaluation version |
| OSWorld 2.1 | 72.4% | Offline subset |
| Terminal-Bench 4.0 | 39.2% | Agentic command-line task setting |
| Humanity's Last Exam | 45.9% | No-tools condition |
Before reusing a number, open the report and its linked system card. Record the model setting and test harness. A tool-enabled result answers a different question from a result without tools. A later benchmark version may also change the task set, scoring or contamination controls. Putting both numbers in one column without those details creates a comparison the original authors did not make.
Three evidence layers for Haiku 5.5 benchmarks
Manufacturer reports establish what the developer says it measured. Independent evaluators supply another test design, which may use different prompts, budgets and endpoints. Your own evaluation establishes whether a model helps with your application. Agreement across these layers is useful, but disagreement is not automatically evidence that somebody is wrong.
For an independent view of Haiku 5.5 benchmarks, Artificial Analysis maintains a model page with its methodology and configuration. Inspect the named effort setting before comparing an index value with another model page. This site links to the original rather than copying a changing leaderboard and making a stale ranking look current.
A private comparison in our workspace belongs to the third layer. It contains your input and the returned outputs for the selected settings. A handful of saved tasks is useful evidence for you, but it is not a representative sample of every developer's workload. Do not publish a win rate unless you can also explain the sample and scoring procedure.
Define acceptance before looking at answers
The most useful local Haiku 5.5 benchmarks start with a decision that can change. For invoice extraction, that might be whether unattended processing is allowed. For email drafting, it might be whether an editor spends less time correcting the draft than writing from scratch. Those decisions need different scoring rules.
Write a short acceptance checklist before generating anything. An extraction task might require the correct invoice number, currency and total, with missing purchase-order references returned as missing. A fluent paragraph is irrelevant to that score. An email task might instead require an accurate deadline, an appropriate recipient and no unauthorized promise. Grammar alone would miss the expensive mistakes.
Keep critical failures separate from ordinary imperfections. Inventing an amount can disqualify an extraction result even when every other field is correct. A slightly awkward phrase should not carry the same weight. A single average score can otherwise make a model look usable by hiding a small number of unacceptable outputs.
Build a sample that can prove you wrong
Include ordinary cases, difficult cases and cases that should produce an abstention. For a support workflow, use a straightforward eligibility question, an ambiguous request and a request that requires information you have not supplied. Deliberately include a plausible but unsupported customer claim. The system should not treat that claim as company policy.
Keep a held-out portion untouched while tuning prompts. If you repeatedly rewrite the prompt after seeing every test answer, you are training the prompt to your worksheet. An improvement on that worksheet may disappear on the next real ticket. Save the prompt revision alongside the result so another reviewer can reconstruct the test.
Synthetic cases are appropriate for an initial check because you control the expected answer and avoid exposing customer data. Label them as synthetic. Later, use authorized, minimized production examples to check whether your initial assumptions survive realistic vocabulary, missing context and uneven formatting. Do not confuse the convenience of synthetic data with evidence of production coverage.
Compare effort and cost on the same task
Haiku 5.5 benchmarks at one effort setting do not establish performance at every setting. Hold the task and acceptance criteria constant while you change effort. Record whether the additional reasoning changes the answer enough to matter. A longer response may be more complete, or it may simply require more review.
For each run, record input usage, billable output usage and whether the answer was accepted. Where your application retries failures, include the retries. A cheap first attempt followed by repeated repairs may cost more per accepted item than a more expensive first attempt that passes. Keep manufacturer API estimates distinct from this site's credit usage and from a gateway's invoice.
Latency also needs a named measurement. Time to first visible text matters for chat, while completion time matters for a batch of records. Queueing, network delay and application validation can change what the user experiences. Do not label an end-to-end stopwatch measurement as raw model generation speed.
A scorecard for local Haiku 5.5 benchmarks
Use one row per input and model setting. The following fields are a worksheet specification, not measured results from Haiku 5.5 benchmarks on this site.
| Field | What to record |
|---|---|
| Case ID | Stable identifier that survives prompt revisions |
| Expected outcome | Required facts, allowed labels or correct abstention |
| Configuration | Model, provider, effort and output budget |
| Result | Accepted, repairable or rejected, with a reason |
| Usage | Reported input and billable output, including retries |
| Review | Reviewer identity or initials and any disagreement |
If two reviewers disagree, resolve the acceptance rule before counting a winner. Their disagreement may reveal an ambiguous task rather than an ambiguous model result. Preserve the original judgments so that later changes to the rubric do not silently rewrite history.
Common pitfalls in Haiku 5.5 benchmarks
One model receives a detailed system prompt while another receives only the user question. A runner truncates one answer earlier. A tool-enabled agent is compared with a plain text completion. A result with an incorrect answer is counted as fast because it finished quickly. These problems all survive a polished chart.
Another failure is selecting only memorable examples. One excellent answer and one disappointing answer can both be real while saying little about the overall distribution. Keep failures and refusals in the record. If an input was invalid because of your request construction, label that integration failure separately rather than quietly dropping the row.
For recurring Haiku 5.5 benchmarks, store the exact test-set version and rerun after changing the provider route, prompt or parser. You do not need a giant evaluation platform to start. A versioned file of cases and a disciplined review sheet are enough to expose many regressions, provided the same rules are applied to every candidate.
Turn the evidence into a routing decision
The result may be a division of work rather than one universal winner. Straightforward records can go to the lower-cost setting; ambiguous records can go to a reviewer or another model. Define the escalation signal using observable failures, such as missing required evidence, rather than asking the model to declare itself confident.
Start a same-input comparison for a text task your team understands. Keep the official price calculator nearby, but make acceptance the first decision. Haiku 5.5 benchmarks become useful when they change a workflow you can measure, not when they supply a headline with no connection to the work.