Price the consequence of a wrong answer
A token-price gap is easy to see. The cost of an unsupported promise in a customer reply is harder to put in a table. Before comparing Haiku with Opus, identify the mistakes that would require a human correction or make the output unusable.
Use Opus as a comparison baseline, not as an unquestioned judge. Both outputs need the same evidence checks. For high-stakes decisions, a model-to-model comparison does not replace qualified human review.
Build a constraint-heavy test
A useful case is a customer message plus a short policy with an exception. Ask for a reply that follows the policy, recognizes the exception only when the facts support it, and asks for missing information before promising a resolution.
Change one fact in a second row so the exception no longer applies. This paired variation makes an important failure visible: repeating an answer that looked right on the first row but ignores the changed condition.
- Every promise is supported by the supplied policy.
- The exception changes the answer only when its condition holds.
- Missing evidence produces a question rather than an invented fact.
- The response still meets the requested length and format.
Separate ordinary work from the boundary cases
Keep routine and difficult cases labeled in the input IDs. A high acceptance rate dominated by easy rows can hide the exact cases that matter. Review those rows individually instead of treating one aggregate number as a deployment decision.
Start with enough examples to expose the failure you care about, then add new cases when real work reveals another pattern. The batch tool does not impose a business row-count cap, although request-size and available-credit limits still apply.
What this comparison cannot establish
This workspace sends text and records responses. It does not run a repository, browse websites, invoke external tools, or reproduce a long-running agent benchmark. Manufacturer claims about those settings need their own evaluation setup.
Your saved pair can support a narrower decision: whether these outputs meet this task's standard at these settings. Keep that conclusion narrow. A result on policy drafting is not evidence about code execution or autonomous work.
Manufacturer sources
- Claude Haiku 5.5 model reference · Checked 2026-10-09
- Anthropic API pricing · Checked 2026-10-09
These sources describe the manufacturers' products. This page does not publish a site benchmark or reproduce a third-party ranking. Your saved outputs are private observations of your own inputs and settings.