OpenAI agents on AWS: governing Bedrock Managed Agents

This research library uses AI-assisted source research and drafting. Linked sources support product claims; analysis and proposed exercises are our interpretation. Unless an article documents a test and its results, do not read it as a hands-on review or an independently verified benchmark.
OpenAI agents on AWS became a procurement option on 29 September 2026, when OpenAI and Amazon announced Bedrock Managed Agents, powered by OpenAI. The service runs the OpenAI agent harness and model inference inside Amazon Bedrock, and agents authenticate with AWS IAM. It is a limited preview, so treat it as something to evaluate now and commit to later. The AWS product page and OpenAI's comparison with the Agents API are the two primary sources.
This guide is for engineering leaders, platform owners and security reviewers who already run workloads on AWS. It covers what the boundary does and does not settle, the questions procurement will ask, and a rollout path with explicit approval points.
Who should consider OpenAI agents on AWS?
Three situations make the case strongest. Your data policy says model inference must stay with your existing cloud provider. Your services already authenticate with IAM roles, and adding OpenAI project keys would create a new secret class to manage. Or your agents need to work close to data in AWS services you already run.
If none of those apply, the OpenAI Agents API is simpler to start with. It is in public beta, documented, and billed at model API rates. Teams that want their own compute but accept OpenAI inference can use its self-hosted sandbox option.
A caution for mixed teams: developers may already use Codex locally with OpenAI models served through Bedrock. That setup uses a model_provider setting in Codex and is a different product. It does not include Codex cloud or several hosted features. Do not let a working local setup stand in for an approved managed-agent platform.
What the AWS boundary settles, and what it leaves open
AWS states that every agent operates with its own identity and logs every action for auditability. It also states that all inference runs on Amazon Bedrock and that your data never leaves AWS. The managed runtime handles inference, memory and skills within your environment. Bedrock AgentCore is the default compute, and OpenAI's docs say self-hosted compute is also possible.
That settles residency and identity questions that often stall an AI pilot. It does not settle behaviour. An agent with a valid IAM identity can still do the wrong thing with the permissions you gave it. Some controls are also future tense. AWS says AgentCore and Bedrock Managed Agents "will provide" authorization policy enforcement, agent and tool discovery, and observability and evaluation. Plan as if those are not available on day one.
OpenAI's docs send you to AWS for data-handling guidance on session state, execution files, logs and model inference. Ask for that guidance in writing before a security review signs off. A marketing sentence about data not leaving AWS is a starting point, not evidence.
Procurement questions to ask before the preview ends
Procurement will want answers that the public pages do not give yet. Collect them early, because preview terms often differ from general availability terms.
- What does the service cost, and how is it metered? AWS does not list managed agents pricing on the product page.
- Which Regions are supported, and do they match your data residency commitments?
- What service limits apply to sessions, concurrency and long-running tasks?
- Which OpenAI models are available, and in which Regions? OpenAI's Bedrock model guide already lists different Regions for different models.
- What preview terms apply, and what changes at general availability?
- Does spend count toward your existing AWS agreement? Neither source says, so ask your AWS account team.
Approval points and owners
The table below is a starting point for your own approval model. Adjust owners to your organisation.
| Decision | Evidence to collect | Suggested owner |
|---|---|---|
| Join the preview | Preview terms, intended workflow, named sponsor | Engineering leader |
| Approve data handling | AWS data-handling guidance for state, files and logs | Security and privacy |
| Grant each IAM permission | The task that needs it and a least-privilege policy | Platform or cloud owner |
| Allow write actions in connected systems | A recovery path and a named person who reviews results | The system's business owner |
| Choose AgentCore or self-hosted compute | Isolation, network egress rules and cost estimate | Platform owner |
| Widen from pilot to more teams | Pilot results against the measures below | Engineering leader with security |
Keep one record per agent: its purpose, its IAM identity, the permissions granted, who approved each, and where its action logs land. That record is what an auditor will ask for.
A rollout checklist
- Pick one workflow that mostly reads data, such as summarising incidents or answering questions over internal records.
- Write the agent's task, allowed tools and stop conditions before anyone requests preview access.
- Create a dedicated IAM identity with only the permissions that workflow needs.
- Confirm where action logs go and who reads them each week.
- Run the pilot with a human reviewing every result that leaves the agent.
- Review after a fixed period and decide to keep, adjust or stop.
OpenAI's sandbox guidance for the Agents API is a sensible baseline here as well. Isolate workloads, restrict outbound network access to approved endpoints, and keep long-lived credentials out of the agent's environment. Agent-generated code can read what the environment contains.
What should you measure?
Measure the whole path, not agent activity. Use the same definitions we describe in our guide to measuring an AI workflow before scaling it.
Track time from task to accepted result, active review time per result, and corrections needed after review. Add two platform measures: the number of permission requests the pilot raised, and how long each took to approve. If approvals take longer than the work, fix the approval model before adding agents.
Log every action the agent took that a reviewer later reversed. A short list of reversed actions tells you more about readiness than a success rate.
Where training fits
The platform decision is only half the work. People still need a shared method for briefing agents and reviewing their output. Our checklist for reviewable AI-generated changes covers the review side. Our AI training for teams runs the same exercises on your own workflows, and the free Delegate, Review, Own guide explains the method.
Next step: send the six procurement questions above to your AWS account team this week, before anyone builds against the preview.