AI Code Review Receipts for Teams

By Rogier Muller10.09.26
AI Code Review Receipts for Teams

Do not let an AI review become another opaque opinion in the pull request. For engineering teams using coding agents, the useful pattern is to make AI code review produce a receipt: what it inspected, what it claims, what evidence supports the claim, and what still needs a human decision.

Most teams do not need more code review tools before they need a better review workflow. The tool can be an IDE agent, a CLI agent, a PR bot, or an MCP-connected reviewer. The operating rule is the same: the review comment should be traceable back to files, tests, logs, and repository rules.

Make the review produce evidence

An LLM code review is useful when it can point to the exact diff, the rule it applied, and the check it wants a human to trust. It is weak when it writes a confident paragraph with no file path, no line range, and no test result.

I put this in the Review step of our methodology: the agent may help find risk, but the team still owns the decision. That keeps the workflow useful for developer productivity without turning review into rubber-stamping.

In a normal GitHub pull request workflow, ask the agent to inspect changed files first, then read the relevant tests, then read CI output. If the PR changes packages/api/src/billing/checkout.ts, the receipt should say whether it inspected the billing tests, the route guard, and the migration or fixture files that the change depends on.

Keep the first integration read-only

Start with read-only access to the PR diff, CI logs, repository rules, and issue context. Do not give the review agent write access, merge access, or permission to dismiss review comments until the team likes the receipts it produces.

MCP matters here because it gives coding agents a common way to reach outside the editor or terminal. The Model Context Protocol specification describes capabilities such as tools, resources, and prompts, which are exactly the surfaces teams use to connect agents to GitHub, docs, test output, and private knowledge bases.

That does not mean every code review AI path should start with a custom MCP server. Use the smallest integration that gives the agent the evidence it needs. For a narrower example of packaging agent judgments behind typed tool boundaries, see typesafe-mcp Adds Typed Agent Judgments.

Use this review receipt procedure

Use the same procedure whether the reviewer runs in an editor, a terminal, or a PR automation. The point is to make the agent leave enough evidence that another engineer can verify the review without replaying the whole chat.

  1. Ask the agent to list the changed files and classify the change as behavior, test-only, refactor, configuration, dependency, or documentation.
  2. Ask it to name the repository rule or convention it will apply before it comments on the code.
  3. Ask it to inspect the nearest tests and call out missing coverage only when it can name the missing behavior.
  4. Ask it to read available CI output and separate failed checks from checks it did not run or could not access.
  5. Ask for findings in severity order, with a file path, line range, claim, and suggested verification step.
  6. Ask for open questions that require product, security, or maintainer judgment instead of letting the model guess.
  7. Require a final receipt before the human reviewer approves, requests changes, or ignores the AI findings.

This is slower than asking for one summary paragraph. It is also much easier to audit when a review comment is wrong.

Paste this receipt into the pull request

Use a small artifact, not a long prompt library. This version is enough for most team skills and review guardrails.


## AI review receipt

PR:
Reviewer tool or agent:
Date:

## Scope inspected
- Changed files inspected:
- Related tests inspected:
- CI or local checks inspected:
- Repository rules or docs inspected:

## Checks requested or run
- Typecheck:
- Unit tests:
- Integration or end-to-end tests:
- Lint or formatting:
- Security or dependency checks:

## Findings
| Severity | File and line | Claim | Evidence | Suggested next step |
| --- | --- | --- | --- | --- |
| Blocker / High / Medium / Low |  |  |  |  |

## Human decisions still needed
- Product or UX judgment:
- Security or privacy judgment:
- Release or migration judgment:
- Maintainer ownership:

## Reviewer decision
- Approve / request changes / ignore AI finding / ask for more evidence:
- Human reviewer:

Know where the tool should stop

AI review can catch missing tests, inconsistent error handling, accidental public API changes, and a route that lost an authorization guard. It is less reliable at product intent, rollout risk, legal constraints, and whether the team wants the behavior at all.

Do not let a code review LLM approve code it just generated unless a separate human reviewer reads the receipt and the diff. Agentic coding works better when delegation and review are separate jobs.

There is also a cost tradeoff. A detailed review receipt uses more tokens and more tool calls than a shallow PR summary. Use it on risky changes first: auth, billing, migrations, permissions, concurrency, public APIs, and generated code that touches production paths.

Further reading

Try it on one pull request

Pick one active PR with real risk, paste the receipt template into the review, and require the agent to fill it before anyone treats its comments as review input.

Where does your team stand?

Each team member completes the proficiency matrix individually. You receive a PDF with the team baseline and a recommended next step.

Assess your team