OpenAI Decisions API: where Luna fits in an agent pipeline

By Rogier Muller09.29.26
OpenAI Decisions API: where Luna fits in an agent pipeline

This research library uses AI-assisted source research and drafting. Linked sources support product claims; analysis and proposed exercises are our interpretation. Unless an article documents a test and its results, do not read it as a hands-on review or an independently verified benchmark.

The OpenAI Decisions API is a new API that points Luna at a set of questions you define, each with a finite list of answers you also define. You send context as text or images and get back answers your code can branch on: classify content, route a request, or choose an agent's next action. OpenAI announced it in the DevDay 2026 recap as a limited preview, with broad release planned in the coming days.

That is most of what is public today. There is no Decisions API guide or API reference page on the developer site yet, so this article does not show a request body. It covers where a finite-answer call belongs in an agent pipeline and how to use Codex to get your code ready for it.

What the OpenAI Decisions API does today

Here is what the launch material states, and what it leaves open.

Question Status on 29 September 2026
What it answers User-defined questions with finite, pre-defined answers
Input Context as text or images
Stated uses Classify content, route requests, choose an agent's next action
Model Luna, per the recap
Availability Limited preview, broad release planned in the coming days
Request and response format Not documented publicly yet
Pricing and limits Not documented publicly yet
Confidence scores or probabilities Not documented publicly yet

The recap names Luna without a version number. The API models page lists GPT-6 Luna as "our most efficient model for focused, high-volume tasks," introduced on 22 September 2026. The recap does not say whether Decisions API calls are billed at GPT-6 Luna token rates, so do not budget on that assumption.

Where a finite-answer call fits in an agent pipeline

Most agent code mixes two kinds of model calls. Some produce work: a patch, a summary, a reply. Others produce a choice that code then acts on: which queue, which tool, which subagent, stop or continue. The second kind is usually a free-text prompt followed by string parsing, and it breaks in dull ways. The model returns "Billing." with a full stop, or a label nobody handles.

A finite-answer call removes that class of bug by design. The answer is always one of the values you listed, so your switch statement has no unknown branch to guess about. Whether the answer is right is a separate question, and you still have to test it.

Three places in a typical Codex-built agent are good candidates:

  • Intake routing: A ticket, issue or chat message arrives. The decision is which handler, team or agent receives it.
  • Triage: A failing test, alert or security finding arrives. The decision is severity or category, which sets what happens next.
  • Next action: Mid-task, the agent has a set of allowed moves, such as run tests, ask the user, open a pull request, or stop. The decision picks one.

Keep generation calls where they are. A decision call cannot write the patch, and nothing in the announcement says it tries to.

Map your decision points with Codex

You can do the useful prep work today, before you have preview access. Ask Codex to find the places where your code already treats model output as a decision. Run this in Codex CLI or a Codex cloud task, in read-only mode if your setup allows it:

Find every place in this repository where a model call returns a label,
category, or choice that code then branches on (if/else, switch, match,
dictionary lookup on the model output).

For each one, report:
- file and line
- the question the prompt is really asking
- every value the code handles, and what happens on any other value
- whether a person reviews the result before an action runs

Do not change any code. Return a markdown table sorted by how often the
call runs, most frequent first.

Review the table Codex returns before you trust it. Grep for your model client yourself and check that no call site is missing. Codex finds the obvious switch statements quickly. It may miss a label that is parsed in one module and acted on in another.

Write the answer set before the prompt

For each decision point on the list, write the question and its allowed answers in plain language, in the repository. This is your own design note, not the API format. Something like:

Decision Allowed answers Action per answer Fallback
Which handler owns this issue? bug, feature, question, security Route to that queue security goes to a person
Is this test failure flaky? flaky, real, unclear Retry once, open ticket, ask a person unclear goes to a person
What should the agent do next? run_tests, edit, ask_user, stop Call the matching tool ask_user

Three rules make these sets hold up. Every answer maps to exactly one action. There is always an answer that means "a person should look." And the answers are mutually exclusive, so two reasonable reviewers would pick the same one for the same input.

Add the table to your AGENTS.md or a linked design doc so Codex keeps the same vocabulary when it edits the routing code later. Our methodology treats these explicit decision points as the places where a human reviews agent behaviour, so they belong in version control.

What to hold until the docs ship

Do not write a wrapper around a guessed request shape. When the reference appears, the field names, the way you declare answers, and any confidence output will decide how your code looks. Build the inventory and the answer sets now. Swap in the real call when the docs land.

Also keep a baseline. Log what your current prompt-and-parse calls return for a week, so you can compare accuracy, latency and cost against the Decisions API on the same inputs.

Pick the single most frequent decision call Codex found, write its answer set today, and start logging its current outputs.

Further reading