OpenAI Ultrafast in Codex and the API: plans, setup, cost

This research library uses AI-assisted source research and drafting. Linked sources support product claims; analysis and proposed exercises are our interpretation. Unless an article documents a test and its results, do not read it as a hands-on review or an independently verified benchmark.
OpenAI Ultrafast is a premium speed tier that OpenAI launched at DevDay on 29 September 2026. GPT-6 Astra Ultrafast generates tokens up to 8x faster in Codex, about 300 tokens per second, and up to 6x faster in the API (Ultrafast mode guide). It is available today for GPT-6 Astra; GPT-6.1 Sol Ultrafast is listed as coming soon.
Which plans get Codex Ultrafast?
Ultrafast in Codex and ChatGPT Work is limited to the top Pro tier and to usage-billed Enterprise and Edu plans. The Codex speed docs and the Pro tiers help article give this picture:
| Plan | Codex Ultrafast on 29 September 2026 |
|---|---|
| Pro 100 ($100 a month) | Not included |
| Pro 200 ($200 a month) | Not included |
| Pro 500 ($500 a month) | Included; uses included usage first, then credits |
| Enterprise on credit-based or USD usage-based agreements | Eligible; off by default until a workspace owner enables it |
| Enterprise on legacy rate-limit plans | Not supported |
| Edu plans that use credits | Eligible |
| Other self-serve plans | Not available at launch, even with purchased credits |
| OpenAI API | GPT-6 Astra for all API users, at low default rate limits |
Buying credits on Pro 100 or Pro 200 does not give you Ultrafast. Workspaces that require inference residency outside the United States cannot use it, and the API version supports only US data residency and global processing.
How to turn on Ultrafast in Codex
On Pro 500, Ultrafast is selectable in the model picker. On Enterprise, a workspace owner first enables it for selected users or the whole workspace through workspace permissions. Existing per-user spend controls apply to Ultrafast usage.
The Codex docs are specific about Fast mode, the cheaper speed tier, but not about an Ultrafast switch in the CLI. Fast mode toggles with /fast, and you can make it the default in config.toml:
service_tier = "fast"
[features]
fast_mode = true
At the time of writing, the docs do not list a CLI command or config value for Ultrafast. Use the model picker, and check the speed page before you script it.
If you run Codex with an API key instead of ChatGPT sign-in, Codex uses API token pricing and the ChatGPT credit multipliers do not apply.
Ultrafast mode API setup
In the API, set model to gpt-6-astra and service_tier to ultrafast. This is the plain HTTP request from OpenAI's guide:
curl https://api.openai.com/v1/responses \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6-astra",
"input": "Explain why the sky is blue in one sentence.",
"service_tier": "ultrafast"
}'
For agent loops, OpenAI strongly recommends a persistent WebSocket connection. Without it, network overhead between tool calls eats into the speed gain. The Python version from the guide keeps one connection open and chains turns with previous_response_id:
from openai import OpenAI
client = OpenAI()
previous_response_id: str | None = None
prompts = [
"Explain why the sky is blue in one sentence.",
"Now explain why sunsets look red.",
]
with client.responses.connect() as connection:
for prompt in prompts:
connection.response.create(
model="gpt-6-astra",
service_tier="ultrafast",
previous_response_id=previous_response_id,
input=prompt,
)
for event in connection:
if event.type == "response.output_text.delta":
print(event.delta, end="", flush=True)
elif event.type == "response.completed":
previous_response_id = event.response.id
print()
break
elif event.type in {"response.failed", "response.incomplete", "error"}:
raise RuntimeError(event.to_json())
else:
raise RuntimeError("Connection closed before the response finished.")
Default Ultrafast rate limits for GPT-6 Astra are 500,000 tokens per minute on usage tiers 1 to 3, 1,000,000 on tier 4, and 5,000,000 on tier 5. Teams with an OpenAI account team can ask for higher limits.
What does Ultrafast cost?
The API pricing page lists GPT-6 Astra Ultrafast at six times the standard price. Short context means up to 272K input tokens.
| GPT-6 Astra, per 1M tokens | Input | Cached input | Cache writes | Output |
|---|---|---|---|---|
| Standard, short context | $10 | $1 | $12.50 | $50 |
| Ultrafast, short context | $60 | $6 | $75 | $300 |
| Ultrafast, long context | $120 | $12 | $150 | $450 |
Take an agent loop of 30 requests. Each one sends the same 40K-token prefix plus 5K new tokens, and gets 1.5K tokens back. The first request writes the prefix to the cache and the next 29 read it. On standard Astra that loop costs about $5.41. On Ultrafast it costs about $32.46.
In Codex, Ultrafast uses included subscription limits at 8x the standard rate. Purchased credits and Enterprise pay-as-you-go usage are billed at 6x. OpenAI notes these multipliers describe billing, not speed. The 8x speed figure also measures token generation, not total task time, so tool calls and test runs take as long as before.
A Codex workflow where the speed pays off
Ultrafast earns its price when you sit and wait on each turn. A tight interactive loop on a failing test is the clearest case. You read the diff, ask for a change, and wait again. Faster generation shortens every wait in that loop.
It is poor value for work you do not watch. A Codex cloud task that runs while you are in a meeting gains little from faster tokens, and you pay the multiplier anyway.
Speed also raises the number of diffs you have to read per hour. Keep your review bar fixed. Put the test command and the evidence you expect in the repo's AGENTS.md, so a faster turn still ends with a test result as well as a patch. Our Delegate, Review, Own methodology covers that review step.
Pick one interactive bug-fix session this week, run it once on standard Astra and once on Ultrafast, and compare wall-clock time with the usage it consumed.