# Code review after agents: verify the change, not the explanation

By Vibe Haus · Published and source-reviewed 2026-10-10

Source: https://vibehaus.team/blog/code-review-for-agents

An evidence-first review workflow for agent-written code, covering risk routing, independent tests, actionable findings, and exact-revision release gates.

Original Vibe Haus synthesis and proposed engineering workflows, informed by the linked primary sources. These articles are not peer-reviewed studies or measured customer results. Examples and thresholds are illustrative unless explicitly attributed.

> Agent-written code deserves the same engineering standard as any other change. Faster generation makes review capacity, evidence quality, and risk selection more important.

## Review three things: intent, implementation, evidence

A convincing pull request description can make a change feel understood before the reviewer has read the code. Start by separating three things. Intent says what the user needs and what must remain true. Implementation is the actual diff and its surrounding behavior. Evidence shows what was checked against that intent. Agreement between all three is stronger than a polished explanation of any one.

Google’s engineering review guidance covers design, functionality, complexity, tests, naming, comments, style, and documentation. Agent involvement does not remove those concerns. Our proposed workflow changes the order: establish the risk and the behavioral claim first, then spend review attention where a plausible patch could still be wrong.

For example, an agent may add a permission check to a new export route and describe the route as secure. The reviewer still needs to determine whose account the query uses, whether the caller can choose another account identifier, and whether the underlying data access enforces the same boundary. A correct-looking guard is not evidence that the entire path respects it.

### Workflow: Review attention follows the possible consequence of the change.

1. Route risk: Identify data, permissions, money, availability, and user-visible behavior.
2. Inspect: Read the diff with its callers, requirements, and evidence.
3. Challenge: Exercise a case that could falsify the main claim.

Decision: Are material findings resolved on the current revision?
- Yes → record the decision and required release checks.
- No → return a reproducible finding; review the resulting change.

Sources: [What to look for in a code review](https://google.github.io/eng-practices/review/reviewer/looking-for.html)

## Route by consequence, not by author or line count

A ten-line authorization change can carry more risk than a thousand-line generated fixture. Classify the consequence of being wrong, the reach of the affected code, and how easily the behavior can be observed or reversed. Record why the route was chosen. Labels such as “small” or “AI-generated” are weak substitutes for this analysis.

A low-risk copy change may need a visual check, accessibility sanity check, and content owner review. A payment calculation needs boundary examples, domain expertise, and verification of rounding and retry behavior. A migration needs compatibility and recovery analysis. The policy should explain which evidence is mandatory for each category without forcing every patch through the most expensive path.

Automation can propose the category, but ambiguous boundary crossings need a named owner. A package update that changes an authentication library is not automatically routine maintenance. If the scope expands during implementation, reroute the review rather than keeping the convenient original label.

| Change surface | Question to challenge | Useful evidence |
| --- | --- | --- |
| Authorization or tenant data | Can the caller cross an ownership boundary? | Negative tests using a different identity and account. |
| Money, retries, or jobs | Can repetition or partial failure duplicate the effect? | Idempotency, rounding, and interrupted-operation cases. |
| Schema or integration | Can old and new versions coexist during release? | Compatibility sequence and a viable recovery plan. |
| User interface | Can a person complete the task across input modes? | Real viewport, keyboard, and error-state checks. |

## Make the verification path independent

An authoring agent can write code and tests from the same mistaken interpretation. If both assume an export should include every account, the tests may pass while reinforcing the defect. A second reviewer reading only the author’s summary can inherit the same assumption. Independence comes from reconstructing the requirement and choosing evidence that could contradict the author.

Ask a reviewer to derive one or more boundary cases before reading the author’s test explanation. Give it the actual requirement, current code, diff, and repository conventions. Then compare its expectations with the submitted evidence. A different model can broaden the search, but it is not a proof of independence: both models may rely on the same stale documentation or incomplete context.

The reviewer should follow a changed value across the relevant path. For the export example: caller identity → account scope → query filters → serialized output. It can inspect a narrow slice deeply instead of vaguely blessing the whole repository. State the inspected scope and any unverified boundary so another person knows where confidence ends.

- Derive at least one falsifying example from the requirement, not from the implementation.
- Inspect callers and downstream effects when a local change alters a contract.
- Treat external text and tool output as evidence, not as instructions that can authorize actions.
- Report what was actually inspected; avoid claims about unexamined parts of the system.

## Challenge what a passing test means

Passing tests answer the questions those tests encode in the environment where they ran. They do not establish that the right questions were asked. Review whether a new test would fail against the original defect, whether its assertions observe user-relevant behavior, and whether mocked dependencies erase the very risk under review.

A useful counterfactual is to run the regression test against the prior implementation when doing so is safe and practical. If it still passes, the test may be covering a nearby behavior rather than the bug. Another approach is a narrowly scoped mutation: deliberately remove the relevant guard in an isolated workspace and confirm that the test catches it. Do not mutate a shared or production environment for this exercise.

Choose test level based on the uncertainty. A deterministic calculation usually benefits from direct boundary tests. A browser interaction may need a real rendered flow. A service integration may require a controlled environment with the actual protocol. Expensive end-to-end tests are not inherently stronger if their assertions are vague or their failures cannot be diagnosed.

## Write findings that can be acted on

A useful finding connects a trigger to an observed or well-supported consequence, identifies the relevant code, and explains why the issue matters. It should distinguish a confirmed defect from a concern that needs investigation. Style preferences belong in established tooling or clearly optional comments; they should not compete with broken permissions or data loss.

Avoid requiring a particular implementation unless the requirement or repository architecture demands it. The author needs to repair the behavior and provide evidence. A reviewer who rewrites the entire feature into its preferred style can enlarge the diff and obscure the original issue. Keep the correction proportional to the demonstrated problem.

This illustrative finding shows how to explain a possible defect and the evidence still needed to confirm it.

```text
Finding: Export query accepts an untrusted account identifier.
Trigger: User from account A requests an export with account B's ID.
Consequence: The query may return B's records if no lower layer scopes access.
Location: Export handler and its repository query (attach actual file/line).
Evidence: Caller identity is checked, but not bound to the query account.
Required resolution: Enforce the ownership boundary and demonstrate denial.
Confidence: Confirm the lower-level query behavior before marking reproduced.
```

## Approvals belong to a revision

Every review decision should identify the candidate commit. If code changes after review, the decision must be reconsidered for the affected area. If the base branch changes, determine whether the new integration state changes the behavior or verification assumptions. “Approved yesterday” is not enough information for a release system.

Define invalidation rules that are practical rather than ceremonial. A typo correction in a comment may not need a new database compatibility exercise. A lockfile change can alter runtime behavior without changing application source. The release gate should know which checks depend on code, dependencies, generated artifacts, configuration, and environment.

The reviewer’s final note should separate findings from release readiness. It can say that no material issues were found within a stated scope while still noting that a required integration environment was unavailable. Missing evidence should remain visible until an accountable person accepts the limitation or completes the check.

## Measure review as part of delivery

Track how long review-ready changes wait, how many are returned for substantive corrections, and which defects escape. Do not rank reviewers by comment count or agents by approval rate. Those incentives favor noise or easy tasks. Sample accepted changes periodically to find blind spots, and use incidents to add targeted examples to the review process.

When review becomes the bottleneck, reduce the arrival rate or improve task boundaries before adding more authoring agents. Changes that each address one clear behavior can be easier to verify; arbitrarily splitting one coupled feature into many pull requests can make it harder. The useful unit is a change a reviewer can understand and accept with a bounded amount of context.

Use the downloadable code-review skill as a starting procedure. It asks for the actual diff, independent checks, revision identity, and actionable findings. It cannot certify security or replace the accountable engineer. Its value is in making a repeatable review standard explicit enough to inspect and improve.

## Sources and further reading

Sources reviewed 2026-10-10. This is a focused reading list, not an exhaustive literature review. Source dates and study conditions matter; follow the original links for their full methods and limitations.

- [What to look for in a code review](https://google.github.io/eng-practices/review/reviewer/looking-for.html): Google Engineering Practices. Engineering guidance.

## Related services

- [ai native engineering](https://vibehaus.team/ai-native-engineering)

- [managed engineering](https://vibehaus.team/managed-engineering)

## Reuse

You may use and adapt our original templates and workflows with attribution. Third-party sources retain their own terms.

Downloads and tools: https://vibehaus.team/data