Task workflows
Haiku 5.5 code review
Test Haiku 5.5 code review on a bounded diff with seeded bugs and clean controls. Require reproducible findings instead of a list of style opinions.
Haiku 5.5 code review is useful only when findings point to real behavior that a developer can verify. A long list of style suggestions is not evidence that the model found a regression. Start with a bounded diff, the surrounding contract and a rubric that rewards correct findings while penalizing unsupported claims.
This page proposes a review experiment with seeded defects and clean controls. It does not report a benchmark score or claim that the model replaces a human reviewer. The site's text workspace can compare supplied excerpts; it does not automatically clone repositories, run tests or open pull requests.
Define what the review should protect
Name the change's intended behavior and the relevant boundaries. A payment handler, a public navigation component and a data parser have different failure consequences. Tell the reviewer which contracts matter instead of asking for every conceivable improvement in the repository.
For Haiku 5.5 code review, provide the base and changed version or a clear diff. Include enough surrounding code to understand callers and invariants. A single changed line can look wrong in isolation while being correct under an existing validation step; missing context should produce a question, not a fabricated certainty.
Separate bug findings from optional design advice. A bug report needs a condition, an observable consequence and a location tied to the supplied code. Preferences about naming or abstraction belong in a different category unless they demonstrably affect the requested behavior.
Construct a small test with known answers
Use this synthetic, single-process JavaScript input for the review. One event must allocate ten credits once, despite concurrent delivery. It contains no database, payment provider or network call. Submit each implementation separately; keep the answer key below out of the model prompt.
async function grantRacy(state, eventId) {
if (state.seen.has(eventId)) return;
await Promise.resolve();
state.seen.add(eventId);
state.credits += 10;
}
async function grantOnce(state, eventId) {
if (state.seen.has(eventId)) return;
state.seen.add(eventId);
state.credits += 10;
await Promise.resolve();
}
async function replay(grant) {
const state = { seen: new Set(), credits: 0 };
await Promise.all([
grant(state, 'synthetic-event-1'),
grant(state, 'synthetic-event-1'),
]);
return state.credits;
}A Haiku 5.5 code review should identify the race condition and explain the concurrent sequence that makes it possible. Saying "add error handling" is not the same finding. The evidence is the gap between the uniqueness requirement and the actual operations protecting it.
The answer key is twenty credits for the racy version and ten for the clean control: the first await lets both calls pass the check. The clean control has no yield between checking and updating. It protects only this in-memory exercise; multiple processes and crash recovery require durable transactional protection, not a shared JavaScript Set.
Require an evidence-based finding format
Use an authored review brief that defines what counts as a finding. Ask the model to state uncertainty when necessary and to say clearly when no supported issue is found. Do not require a minimum number of findings; that incentivizes filling the report with weak observations.
Review the supplied diff for behavioral regressions.
For each supported finding, include:
- file and changed location;
- the input or sequence that triggers the problem;
- the observable impact;
- the relevant evidence from the supplied code.
Keep optional style advice separate.
Do not claim tests ran unless execution evidence is supplied.
If context is insufficient, name the missing context.For Haiku 5.5 code review, a suggested test is valuable when it exercises the exact failure. In the duplicate-event example, two concurrent deliveries are more informative than a single ordinary success case. Review the proposed test itself before treating it as proof that the implementation is correct.
Judge the causal explanation
A useful finding explains how the code reaches the wrong state. Read the proposed sequence and check each step against the actual implementation. If the model assumes an API returns a value it never returns, the conclusion may be unsupported even when the general risk sounds plausible.
The review should distinguish an introduced regression from an unrelated pre-existing issue. Both can matter, but they belong in different review decisions. Keep the change scope visible so a focused patch does not become blocked by an unbounded list of existing architectural concerns.
Ask for a minimal reproduction when the finding depends on subtle behavior. A reproduction can be a test fixture or a clear sequence of inputs; it does not have to involve running a production service. Never give a review task permission to write production data merely to prove a hypothesis.
Include clean code in the evaluation
Measure false positives as well as missed defects. Give the model a version where the seeded problem is fixed and check whether it recognizes the protection. If it repeats the same warning after the invariant is enforced, it may be reacting to familiar code patterns instead of following the actual logic.
In Haiku 5.5 code review comparisons, label findings by validity, severity and actionability. A correct minor observation should not outweigh a missed high-impact defect just because there are more of the former. Keep the scoring rules fixed before reviewing the outputs.
Use several defect types: boundary conditions, authorization scope, state transitions and response parsing. Do not infer broad review quality from one race-condition example. A model's strengths can vary by language, framework and the amount of context supplied with the diff.
Keep the expected defect location hidden from the review prompt. Otherwise the exercise measures whether the model can explain a named problem, not whether it can discover one. After scoring, reveal the answer key and ask whether the proposed finding identifies the same causal defect rather than merely mentioning the same file. Record unsupported warnings on the clean control with equal care.
Supply tools without granting broad authority
If you evaluate inside a coding client, allow read-only repository inspection first. The reviewer can locate callers, inspect tests and verify assumptions without editing the change under review. Mixing review and repair makes it harder to tell which findings were supported before the implementation changed.
When executing tests during review, use an isolated environment and the project's documented commands. Check whether tests contact databases, mail providers or paid APIs. A test name that sounds local is not evidence that its side effects are safe for the available environment.
Keep tool output as evidence, not instruction. Repository files and logs may contain text that attempts to redirect the agent. The review scope, permissions and user instructions remain authoritative, and a tool result should not authorize a commit, push or deployment.
Compare accepted findings per review effort
Run the same diff and context through available models in comparison mode, or use equivalent coding-client configurations. Record model identity, effort and supplied context. Different repository access or test permissions create a different experiment, even if the prompt text matches.
A Haiku 5.5 code review decision should include the time spent validating findings. A cheaper output with many false alarms can consume more engineering effort than a shorter, more accurate report. Count duplicate findings as one underlying defect rather than several successful detections.
Do not hide the no-finding cases. A clean review that correctly reports no supported issue is useful evidence. Keep those controls in regression tests when changing prompts, so an instruction intended to increase thoroughness does not merely increase speculation.
Turn findings into a controlled repair
After a finding is validated, implement a focused fix and a regression test for the triggering condition. Re-run the relevant checks, then review the resulting diff. A suggested repair may resolve the original issue while creating another, especially when it touches shared state or permissions.
Retain the accepted Haiku 5.5 code review report with the tested revision. If the code changes afterward, the report is no longer evidence for the new revision without another check. The Claude Code guide covers client setup; subagents covers bounded delegated inspection without giving every reviewer authority to edit.