OpenAI Agents API computer use: risk controls for UI agents

This research library uses AI-assisted source research and drafting. Linked sources support product claims; analysis and proposed exercises are our interpretation. Unless an article documents a test and its results, do not read it as a hands-on review or an independently verified benchmark.
OpenAI added computer use to the Agents API at DevDay on 29 September 2026. Agents can now operate a browser in an OpenAI-hosted environment to test a site, collect information, or use an application through its UI. The same release brought Codex's multi-agent capabilities, tool search, tool calling and context compaction to the Agents API. The computer use guide documents the controls in detail, including what they do not cover.
This article is for engineering leaders deciding whether to let agents operate software UIs, and under what conditions. The short answer: yes for read-only work on sites you choose, with a named approver. Hold off on anything that can buy, delete or change access until you have a control the platform does not give you by default.
Why UI agents need a different review model
An agent calling an API tool does one declared thing with declared arguments. You can review the tool list and know the blast radius. An agent operating a browser does whatever the page lets a user do. A settings page, an admin console or a checkout flow is one click away from a read-only page on the same site.
That changes the review question. With API tools you ask "which tools does this agent have?" With a browser you ask "which websites can it reach, and what can a signed-in user do there?" The answer depends on your application's permissions as much as on the agent.
Two further risks come with the browser. Page content is untrusted input, and the docs are explicit that website content cannot grant permission or override the user's instructions. Screenshots can also contain account data, so they need the same handling as the pages themselves.
Controls in Agents API computer use, and the gaps
| Control | What the platform provides | What your team still owns |
|---|---|---|
| Website access | Your application must approve, deny or cancel each new origin, including public sites | Who approves, and a list of origins that are always denied |
| Network | network.access can be enabled, disabled or restricted to 1 to 100 exact host names |
Keeping the allowlist short and reviewed |
| Consequential actions | No per-action confirmation. Origin approval does not enforce one | Targets that cannot perform them, or a browser runtime you control |
| Sign-in | A dedicated flow; submitted values stay outside the model input and saved history | Deciding which accounts an agent may use at all |
| Unsupported sign-in | Passkeys and QR-code sign-in do not work | Not weakening authentication to make a task pass |
| Subagents | Only the main agent can request sign-in | Designing tasks so any sign-in happens in the main agent |
| History | Browser steps are saved as computer_use_call items |
Retention, access to screenshots, and log hygiene |
| Cleanup | Deleting the session requests environment cleanup | Making deletion part of every run |
The row to read twice is consequential actions. OpenAI's guide says that if your application must guarantee confirmation before purchases or destructive changes, restrict the hosted browser to resources that cannot perform them or use a runtime you control. A confirmation step built as a function tool relies on the agent choosing to call it.
Data and compliance questions to settle first
Some answers end the pilot before it starts, so ask them early.
- Zero Data Retention: the Agents API overview says the Agents API does not support ZDR. Workloads that require it are out of scope.
- Data residency: the overview says the Agents API currently supports data residency only in the United States.
- Plans: the recap lists computer use in the API, and in Codex and ChatGPT Work on Pro 500 and Enterprise.
- Cost: model usage bills at API rates, and hosted sandboxes at standard container rates. Browser tasks can run for many steps, so budget per task, not per call.
- Accounts: any account the agent signs into should be a dedicated, least-privilege account, not a person's own login.
Approval points your team must name
Each of these needs a person, not a role on a slide.
- The owner who approves which websites are on the allowlist and which are always denied.
- The approver who answers origin requests at run time, with deny as the default.
- The account owner who decides which credentials an agent may use through the sign-in flow.
- The reviewer who checks the saved browser activity after a run, before its result is trusted.
This maps onto our Delegate, Review, Own methodology. The agent is delegated a bounded task. A person reviews the evidence. A named owner answers for the allowlist and the accounts.
Rollout checklist for the first UI agent
- Choose a read-only task on a public or staging site you control.
- Set
network.accesstorestrictedand list only the hosts that task needs. - Keep origin approvals human, with deny as the default answer.
- Cancel every sign-in request in the first phase.
- Keep screenshots out of application logs and limit who can view them.
- Delete the session at the end of every run, whatever the outcome.
- Review the saved activity for five runs before adding a second site or a signed-in account.
The same principle as our guide to reviewable AI-generated changes applies: ask for evidence that matches the task, not a claim that it worked.
What to measure before widening access
Follow the approach in measuring an AI workflow before scaling and keep the record small.
- Origin requests per task, and the share denied. A rising count means the task wanders.
- Turns that fail or are cancelled, and why.
- Reviewer minutes per run spent checking activity and screenshots.
- Tasks where the result was wrong despite a completed turn.
- Cost per completed task, including sandbox time.
Widen access only when denials are rare and reviewers trust the result without replaying every step. If you want help designing the approval flow for your team, our training covers it on your own systems. This week, write down the one site and the one read-only task you would allow first.