Codex automatic code review in your team review policy

This research library uses AI-assisted source research and drafting. Linked sources support product claims; analysis and proposed exercises are our interpretation. Unless an article documents a test and its results, do not read it as a hands-on review or an independently verified benchmark.
Codex automatic code review gives every GitHub pull request an AI first pass before a person opens it. At DevDay on 29 September 2026, OpenAI announced a new Code Review experience for all plans: Codex reviews pull requests in the cloud while you are away, and the ChatGPT desktop app shows summaries, diffs, and findings before you post feedback. The GitHub review guide is the primary source for how it works. The decision for a team is what the human reviewer still owns once it is switched on.
This page is a policy guide. For the click-by-click setup, see the Code Review docs.
What does Codex automatic code review cover?
The first pass reviews the pull request diff and follows the review rules in AGENTS.md. Codex applies the root file and any nested file that covers each changed path, so service rules stay with the service. On GitHub it flags only P0 and P1 issues, which keeps its comments on high-priority risks.
Reviews run automatically, based on repository settings and each person's Review trigger, or on request with an @codex review comment. A separate Security Review, in research preview, goes deeper on security risks and runs with @codex security review or alongside code review.
The limits matter as much as the features:
- It does not approve or merge. A review in the desktop app chat posts nothing, and people choose which findings to share.
- It does not replace tests, branch protections, or required approvals. The GitHub guide says so directly about review rules.
- GitLab coverage is partial. GitLab merge requests in the desktop app are in preview, and the GitLab integration is in beta.
Who owns what in the review?
| Check | Owner | Evidence the owner signs off on |
|---|---|---|
| Formatting, lint, and other deterministic checks | CI | A passing pipeline, kept out of review rules |
| High-priority risks in the diff (P0, P1) | Codex first pass, verified by the reviewer | A finding with file, lines, and explanation |
| Repository-specific rules | One rule owner per area, through AGENTS.md | A rule that names the risk and the safe path |
| Fit with requirements and acceptance criteria | Human reviewer with the task owner | Behaviour checked against the agreed brief |
| Design and maintainability | Human reviewer | An explanation of the approach and its trade-offs |
| Whether each AI finding is real | Human reviewer | The finding checked against the latest diff and tests |
| Approve, request changes, merge | Human reviewer, under branch protection | A submitted review decision |
The reviewer's job shrinks in one place and grows in another. Codex may catch a missing null check before anyone looks, but someone now has to judge whether its findings are real. Count that judgement as review work when you plan capacity.
Which policy decisions come before you enable it?
Scope comes first. Decide which repositories get automatic reviews and whose pull requests. Repository settings can apply reviews to all pull requests or follow each person's preferences, and each person picks a Review trigger.
The status of findings comes second. We suggest that the author answers every P1 finding with a fix or a short reason for dismissing it, and the reviewer checks that answer. Treat findings as advisory until that habit is in place.
Rule ownership comes third. Give each AGENTS.md review section one owner who adds, narrows, or removes rules. The guide recommends starting with two or three concise rules and leaving mechanical checks to CI.
Fix requests need a rule too. A comment such as @codex fix the P1 issue starts a cloud chat that can push a fix to the branch when Codex has permission. Decide who may ask for that, and require a fresh human look at the pushed commit.
Posting habits need a sentence of training. In the desktop app, a comment posted from Summary or Changes goes to GitHub or GitLab immediately, without waiting for Submit review. Reviewers should draft in the chat and post deliberately.
Cost is the last decision. Reviews that Codex runs through GitHub count as Code Review usage, and local reviews count toward general limits. API-key sign-in does not include GitHub code review.
Rollout checklist
- Pick one active repository with a named rule owner and a reviewer who will give feedback.
- Write two or three Code Review Rules in AGENTS.md for mistakes reviewers explain again and again.
- Turn on Automatic review for that repository only, and leave existing required approvals in place.
- For two sprints, have authors mark each Codex finding as fixed or dismissed, with a reason.
- Narrow or remove rules that produce noise, and add a rule only after a real miss.
- Review the measures below before you extend automatic review to the next repository.
What should you measure?
Track findings accepted and dismissed, reviewer active time per pull request, rework after the first review, and defects found after merge. A high dismissal rate suggests the rules need narrowing. Falling reviewer time alongside rising post-merge defects suggests people lean on the first pass too much.
Keep the definitions stable with our measurement guide. The evidence a reviewer should receive is listed in From AI-generated code to a reviewable change. Our free methodology guide explains Delegate, Review, Own, and AI training for teams can run review exercises on your own pull requests.
Choose the pilot repository and write its first two review rules before you switch anything on.