AI Code Review With Receipts

By Rogier Muller10.07.26
AI Code Review With Receipts

This research library uses AI-assisted source research and drafting. Linked sources support product claims; analysis and proposed exercises are our interpretation. Unless an article documents a test and its results, do not read it as a hands-on review or an independently verified benchmark.

Do not let an agent approve code silently. AI code review is useful when it produces a small, reviewable receipt: what changed, what it inspected, what it could not prove, and what a human still owns.

Most code review tools can comment on a pull request. That is not the hard part. The hard part for an engineering team is knowing whether the model actually checked the risky parts of the change, especially when coding agents can call external tools through MCP and pull context from GitHub, issue trackers, docs, and internal systems.

In AI coding training, I treat this as the Review step in the methodology: the agent may inspect and explain, but the team still owns the merge decision.

Require a review receipt before merge

A review receipt is a short record attached to the pull request. It should be boring enough that people use it, and specific enough that a reviewer can spot gaps.

The receipt should name the files reviewed, the commands or checks run, the assumptions made, and the unresolved risks. If an LLM code review says “looks good” but cannot show which tests ran or which migration path it inspected, the reviewer should treat that as a comment, not a review.

Use the receipt for both agent comments and human follow-up. The point is not to make the model sound confident. The point is to make its work auditable.

Keep the agent close to the diff

Start with the pull request diff, changed files, tests, and relevant local rules. Only then add wider context.

A normal GitHub PR workflow is enough to see the shape. If a Rails app changes a background job, a migration, and a billing model, the agent should first summarize those touched files, then check the migration safety, job retry behavior, and billing edge cases. It should not wander into a full architecture review unless the reviewer asks for it.

MCP matters here because it gives agents a standard way to call tools and read resources outside the editor. That can make a code review LLM more useful, but it also widens the blast radius. Give the review agent read-only access first, and prefer narrow tools such as “fetch PR diff”, “list changed files”, and “read linked issue” before broader repo or database access.

For a related pattern, see how jev-mcp packages agent judgments as tools. The useful pattern is not the specific implementation. It is the idea that judgments should be explicit tool outputs, not hidden chat residue.

Run the review in a fixed order

Use one procedure across code review tools. The exact product matters less than the order of operations.

  1. Read the PR title, description, linked issue, and changed files.
  2. Identify the riskiest change category: data, auth, money, concurrency, performance, public API, or dependency behavior.
  3. Ask the agent to inspect only that risk category first.
  4. Run or name the relevant checks, including tests, linters, type checks, migration checks, or static analysis.
  5. Ask for a receipt with evidence and unknowns.
  6. Have a human reviewer accept, reject, or amend each material finding.
  7. Keep the receipt with the PR so future reviewers can see what was checked.

This procedure is slow enough to catch mistakes and fast enough to use on normal work. It also prevents a common failure mode: the agent produces a polished general review while missing the one thing the PR could break.

Paste this review receipt into the PR


## AI review receipt

PR:
Reviewer:
Agent or tool used:
Date:

## Review scope
- Changed files inspected:
- Linked issue or spec inspected:
- Areas intentionally skipped:

## Risk category
- Primary risk: data | auth | money | concurrency | performance | public API | dependency | other
- Why this risk matters:

## Checks run or verified
- Tests:
- Type checks or lint:
- Migration or schema checks:
- Security or permission checks:
- Manual commands:

## Findings
| Finding | Evidence | Human decision | Follow-up |
|---|---|---|---|
|  |  | accept / reject / defer |  |

## Unknowns
- What the agent could not verify:
- What still needs human judgment:

## Merge decision
- Reviewer decision:
- Required changes before merge:

Keep this short in real use. If it grows into a report, people will stop filling it in.

Set limits before adding integrations

Do not connect every system on day one. Start with the repo, the PR, and read-only project context.

The official MCP specification includes security and trust concerns for a reason: once an agent can call tools, the review path is no longer just prompt in and comment out. A review agent may fetch documents, inspect resources, or call tools that expose sensitive context. That is powerful, but it needs scoped access and clear logging.

The tradeoff is real. Narrow access makes reviews more predictable, but the agent may miss context. Broad access gives better recall, but it is harder to audit. I would rather have a small, reliable review receipt than a wide agent pass that nobody can replay.

Further reading

Put the receipt in one PR today

Pick one active pull request and require the receipt before merge. Do not change the whole review process until the team has seen what the agent can prove and what it leaves for humans.

Where does your team stand?

Each team member completes the proficiency matrix individually. You receive a PDF with the team baseline and a recommended next step.

Assess your team