OpenAI Decisions API as a control point for agent workflows

By Rogier Muller09.29.26
OpenAI Decisions API as a control point for agent workflows

This research library uses AI-assisted source research and drafting. Linked sources support product claims; analysis and proposed exercises are our interpretation. Unless an article documents a test and its results, do not read it as a hands-on review or an independently verified benchmark.

The OpenAI Decisions API, announced at DevDay on 29 September 2026, points Luna at questions you define and returns answers from a finite set you also define. OpenAI's DevDay recap names three uses: classify content, route requests, and choose an agent's next action. It is in limited preview, with broad release planned in the coming days.

For teams running coding agents, the interesting part is not speed. It is that a decision with a fixed answer set can become an explicit control point in the workflow: a place where behaviour is predictable, logged and reviewable. This article covers when to use one, who should own it, and how to pilot it before the documentation arrives.

What a fixed answer set gives you, and what it does not

Most agent workflows already contain decisions. Which queue gets this issue. Is this change in scope. Should the agent stop and ask. Today those are usually free-text prompts whose output code then parses. The failure mode is an answer the code never expected, handled by a default branch nobody reviewed.

A fixed answer set closes that gap. Every possible output is known in advance, so every output can have a reviewed action attached. That is what makes it a control and not just a cheaper model call.

It does not make the decision correct. A confidently wrong route to the wrong team is still wrong. We made the same point about Jev, another decision model, in our analysis of typed decisions in the agent harness: schema-valid output is a property of the format, not of the judgement.

What OpenAI has not published yet matters for governance. There is no public request reference, pricing, or statement about whether answers come with a confidence value. Jev's documented design leans on calibrated probabilities to decide when to escalate. If the Decisions API returns only an answer, your escalation logic has to come from the answer set itself.

Which decisions belong in a fixed-answer call?

Not every choice should go to a model. Use this table to sort the decisions in one workflow.

Decision type Example Best control Why
Rule is written and stable Files under migrations/ need a DBA review Plain code A rule you can write down should not depend on a model
Judgement on messy input, low impact Label an incoming issue as bug, feature or question Fixed-answer decision call Wrong answers are cheap to correct and easy to sample
Judgement, medium impact Choose which subagent handles a task Decision call with a person-review answer The answer set itself offers escalation
Judgement, high impact Merge, deploy, change access, spend money Human approval A decision call can suggest, but a person decides
Open-ended Explain why a test fails Generation model Finite answers cannot carry the explanation

The middle rows are where a decision call earns its place. The top and bottom rows are where teams most often misuse one: replacing a clear rule with a model, or letting a model approve something consequential.

Who owns the answer set?

Treat each answer set like a small policy document. It decides what the system is allowed to do, so it needs the same care as a permission.

  • Name one owner per decision, usually the person who owns the downstream action.
  • Keep the question, the answers and the action per answer in version control, reviewed like code.
  • Always include an answer that routes to a person, and make it the default for anything unclear.
  • Change the answer set only through review, and record the date, because a new answer changes behaviour for every past log you compare against.
  • Log the input reference, the answer and the action taken for every call, so a reviewer can reconstruct what happened.

This is the Delegate, Review, Own pattern from our methodology applied to a single call. The model gets delegated a narrow judgement. A person reviews samples. A named owner is accountable for the answer set.

A pilot plan for the Decisions API preview

You do not need preview access to start. The control design is independent of the API.

  1. Pick one workflow with a repeated, low-impact decision, such as issue labelling or test-failure triage.
  2. Write the question, the answer set and the action per answer. Add the person-review answer.
  3. Run the decision with your current method for two weeks and log every call.
  4. Have the owner review a random sample each week and record the correct answer.
  5. When the Decisions API is available to you, run it on the same inputs in shadow mode, with no actions attached.
  6. Compare agreement, latency and cost, then decide whether to switch the action over.

Keep high-impact actions behind human approval throughout. A pilot that only proves a decision call is fast has not tested whether it is safe to act on.

What to measure per decision

Use the same discipline as our guide on measuring an AI workflow before scaling. Measure per decision, not per workflow, because one weak decision can hide inside good averages.

  • Agreement with the owner's sampled answers.
  • Rate of the person-review answer. Near zero may mean the set is missing an escape; very high means the question is too vague.
  • Cost of a wrong answer, in rework time or in tickets that reached the wrong team.
  • Latency and cost per call against your current method, once pricing is published.

If agreement is low on one answer, split or rename that answer before blaming the model. Vague answer sets produce vague routing with any model.

Our team training works through this exercise on your own workflows. This week, write down the answer set for the one decision your agents make most often, and name its owner.

Further reading