When Codex Ultrafast is worth paying for in a team

This research library uses AI-assisted source research and drafting. Linked sources support product claims; analysis and proposed exercises are our interpretation. Unless an article documents a test and its results, do not read it as a hands-on review or an independently verified benchmark.
Codex Ultrafast is worth paying for when a person is waiting on every agent turn and that wait is the slowest part of the work. It is rarely worth it for background tasks. OpenAI launched Ultrafast at DevDay on 29 September 2026: GPT-6 Astra generates tokens up to 8x faster in Codex, about 300 tokens per second (Codex speed docs).
This is our proposed decision frame for engineering leads. The prices and plan rules are OpenAI's. Whether the speed is worth it depends on your workflow.
What Ultrafast changes, and what it leaves alone
Ultrafast changes token generation speed. OpenAI states that the 8x figure is not a measure of total task time. Tool calls, test runs, builds, and the reviewer's reading time stay the same.
It also changes the bill. In Codex, Ultrafast uses included subscription limits at 8x the standard rate. Purchased credits and Enterprise pay-as-you-go usage are billed at 6x. In the API, GPT-6 Astra Ultrafast costs $60 per million input tokens and $300 per million output, six times the standard price.
So the question for a team is narrow. Where does waiting for tokens cost more than six to eight times the model usage?
Where speed pays, and where it does not
| Workflow | Is a person waiting on each turn? | Ultrafast worth it? |
|---|---|---|
| Pairing with an agent on a live bug | Yes, every few minutes | Often, if the loop is long |
| Incident response with a fix under time pressure | Yes | Often, with the same review rules |
| Demo or workshop where people watch the agent | Yes | Sometimes, for the session only |
| Feature work split into cloud tasks | No, results arrive later | Rarely |
| Overnight or scheduled tasks | No | No |
| Large refactor with long test runs | Partly; tests dominate the wait | Rarely |
A useful test is to time one real session. If most of the wall-clock time is test runs or thinking about the diff, faster tokens will not change much.
Who should get access?
Access is limited by plan. On consumer plans, only Pro 500 at $500 a month includes Ultrafast. Pro 100 and Pro 200 do not, and buying credits on them does not add it. On business plans, eligible Enterprise workspaces on credit-based or USD usage-based agreements can use it, as can eligible Edu plans. Legacy Enterprise plans on rate limits are not supported.
For Enterprise, Ultrafast is off by default. Workspace owners can enable it for selected users or the whole workspace. Existing per-user spend controls apply. That gives a lead a clean way to run a trial: named people, a spend cap, and an end date.
Workspaces that require inference residency outside the United States cannot use Ultrafast. Check that before you plan a trial in an EU-resident workspace.
Review discipline when agents get faster
Speed moves the bottleneck. When an agent produces a diff every minute instead of every few minutes, the reviewer becomes the limit. A team that does not plan for this can end up merging more code with less reading.
Four rules keep review honest:
- Keep the evidence bar fixed. A faster change still needs the test that proves it and a note on what was not checked. Our checklist for reviewable AI-generated changes sets out that evidence.
- Keep changes small. Faster generation tempts people to accept larger diffs in one go. Split at a boundary the reviewer can understand.
- Put the rules in the repo. The test command and done criteria belong in
AGENTS.mdor its equivalent, so every agent run ends with the same checks. - Watch the queue. If pull requests from Ultrafast sessions wait longer for review, the speed is being spent in the wrong place.
The same applies whichever tool your team uses. Faster agents in Claude Code or Cursor raise the same review question.
A two-week trial plan
- Pick two or three people who do long interactive sessions, such as on-call engineers or people pairing on hard bugs.
- Enable Ultrafast for them only, with a per-user spend cap.
- For each session, record wall-clock time, usage consumed, changes accepted, and review minutes.
- Run a similar set of sessions on standard speed for comparison. Keep the task mix as close as you can.
- At the end, compare cost per accepted change and time to accepted change, not tokens per second.
- Decide to keep, narrow, or switch off. Write down the reason.
Our guide to measuring an AI workflow before scaling it explains how to keep those definitions stable across the two weeks.
What to watch next
GPT-6.1 Sol Ultrafast is listed as coming soon. Sol's standard price is a fifth of Astra's, so the cost side of this trade may shift. Run the same trial again when it arrives rather than assuming the result carries over.
Start by timing one interactive agent session this week and noting how much of it was spent waiting for output. If you want help designing the trial, book a short call.