MCP Events and unattended agents: guardrails before rollout

This research library uses AI-assisted source research and drafting. Linked sources support product claims; analysis and proposed exercises are our interpretation. Unless an article documents a test and its results, do not read it as a hands-on review or an independently verified benchmark.
MCP Events lets a ChatGPT plugin start work when something happens in a connected app, such as a new task on a project board or a new comment on a document. OpenAI added support for the proposed MCP Events specification at DevDay 2026 on all plans, and the developer guide sets out how it works. Adopt it first for automations that produce drafts, and keep anything that changes a system of record behind human review.
This is a change in who starts the work. Until now, a person typed a request and read the answer. With events, an outside system starts the run, and the first person to see the result may be the reviewer.
What does MCP Events change?
The mechanics are simple. A plugin's MCP server lists the events it can send. A user tells ChatGPT what to watch and what to do when an event arrives. ChatGPT subscribes and gives the server a callback URL and a signing secret. When something matching happens, the server sends the event, and ChatGPT follows the user's instructions in that chat.
The OpenAI example watches a project board for new tasks, reads the linked documents and drafts a plan while you are away. The developer guide adds a sharper one: monitor a feedback channel for bug reports and open draft pull requests with fixes and tests.
The specification itself is still a proposal. The MCP Triggers and Events Working Group was chartered in March 2026, and its events SEP is listed as ideating. ChatGPT implements the webhook part of the draft. Treat the details as early and expect them to move.
Should your team use event-triggered agents now?
It depends on what the automation produces. The table below is the decision we would make for a first pilot.
| Automation output | Example | Start now? | Who reviews |
|---|---|---|---|
| A summary or plan in the chat | Draft a plan when a task lands on the board | Yes | The person who owns the task |
| A draft in another tool | Open a draft pull request for a reported bug | Yes, with branch protection | A code reviewer |
| A comment or status change in the source app | Reply to a document comment | After a month of clean drafts | The document owner |
| An irreversible action | Merge, send to a customer, delete a record | Not in the pilot | Not delegated |
The line between rows two and three matters. A draft waits for a person. A comment or status change is visible to others the moment it happens.
What can go wrong while nobody watches?
Instructions hidden in event text are the first risk. A bug report or a comment is written by someone, and that someone may not be your colleague. The guide tells server builders to treat user-authored text as data and never to put model instructions inside the payload. Your side of that bargain is to give the automation tools that cannot do much harm if it follows a bad instruction.
Feedback loops are the second. If an automation writes back to the app it watches, its own change can trigger the next run. The guide lists this as a test case. Ask for it in your acceptance criteria.
Duplicates and ordering are the third. Events can arrive out of order and deliveries are retried, so the guide asks for idempotent write tools. Check that a repeated event cannot open a second pull request or post the same reply twice.
Stale access is the fourth. A subscription can outlive the reason for it. Servers must recheck the user's access during the subscription's lifetime and stop delivery when access is revoked, and subscriptions carry an expiry that ChatGPT refreshes. Someone should still decide when a subscription is no longer needed.
Missed events are the fifth. For event types without replay, events missed during an interruption cannot be recovered through the protocol. Do not build a process where one lost event means a lost customer request.
Guardrails to agree before the first subscription
- Name an owner for every subscription and record what it watches and what it may do.
- Use the narrowest filter the plugin offers, such as one channel, one document or one project.
- Give the automation read tools plus draft-only write tools for the pilot.
- Protect branches so a draft pull request cannot merge without review.
- Write success criteria and guardrails into the instructions the user gives ChatGPT.
- Test revoked access, duplicate events and the feedback loop before the pilot starts.
- Set a review date for each subscription and end the ones nobody uses.
If you use ChatGPT Business or Enterprise, Team Tasks are a related option. They run on a schedule or on supported events, use the team's service account and configured connections, and admins control who may create them through the Create and manage team automations permission. Decide which route suits which work, and apply the same guardrails to both.
What should you measure in the pilot?
Count how many events arrived, how many runs they started, and how many outputs a person accepted, edited or discarded. Record reviewer time per output, because an automation that creates ten drafts a day can move work onto one reviewer. Log every run that acted on a false trigger or on text it should have ignored.
Compare against the manual baseline for the same task. Our guide to measuring an AI workflow before scaling it sets out the measures, and our checklist for reviewable AI-generated changes covers what each draft should carry for its reviewer.
Where does this fit in how teams work with agents?
Event-triggered work is delegation without a person at the start. That makes the review step carry more weight, not less. Teams that already agree task boundaries and evidence for interactive agent work in Codex, Claude Code or Cursor have most of the habits they need. Teams that do not should build those habits first. Our free Delegate, Review, Own guide is a place to start, and our AI training for teams practises it on your own repositories.
Choose one draft-only automation, give it an owner and a review date, and run it for two weeks before adding a second.