A GPT-6.1 Sol model policy: default, escalate, measure

By Rogier Muller09.29.26
A GPT-6.1 Sol model policy: default, escalate, measure

This research library uses AI-assisted source research and drafting. Linked sources support product claims; analysis and proposed exercises are our interpretation. Unless an article documents a test and its results, do not read it as a hands-on review or an independently verified benchmark.

GPT-6.1 Sol gives most teams a reason to change their default model. OpenAI released it at DevDay on 29 September 2026 at a fifth of GPT-6 Astra's standard token prices, and says it comes close to Astra on coding and computer-use work (OpenAI launch post). Our advice is to make it the default, keep Astra as a named escalation, and decide with evidence from your own repository.

This is a proposed policy shape for engineering teams using Codex, Claude Code, or Cursor with OpenAI models. The prices and benchmarks are OpenAI's. The decisions are yours.

Why price per token is the wrong unit

GPT-6.1 Sol costs $2 per million input tokens and $10 per million output tokens in the API. GPT-6 Astra costs $10 and $50. Cached input is $0.10 against Astra's $1.00. On paper, Sol is five times cheaper.

The bill a team pays is per task, not per token. A task includes reasoning effort, tool calls, retries, and the minutes a reviewer spends on the result. A cheaper model that needs two attempts and a longer review can cost more than one Astra run.

OpenAI's own per-task figures show why the unit matters. On Terminal-Bench Science, GPT-6.1 Sol cost $5.47 per task and Astra $23.80. Astra still scored highest, at 68.1%. Whether that gap matters depends on what a failed task costs you.

When is GPT-6 Astra still worth it?

OpenAI's Codex guidance recommends Astra for demanding end-to-end work that needs sustained reasoning. It recommends GPT-6.1 Sol for repeated, long-running work where cost matters. That maps onto a decision table most teams can adopt:

Work type Default Escalate to Astra when
Bug fix with a clear reproduction GPT-6.1 Sol Sol fails twice with the same evidence
Feature work inside one service GPT-6.1 Sol The change crosses several services or data models
Test writing and refactors GPT-6.1 Sol Rarely; retry at higher effort first
Code review pass GPT-6.1 Sol The change touches security or payments
Architecture investigation GPT-6 Astra Not applicable
Summaries and extraction GPT-6 Luna Not applicable

The "fails twice" rule stops people from reaching for Astra out of habit. It also creates a record of where Sol falls short, which feeds the next policy review.

Run an evaluation on your own repository

Public benchmarks tell you what to test. Your codebase tells you what to buy. A two-week comparison is enough for a first decision.

  1. Pick 10 to 20 closed tickets that represent normal work, with known accepted fixes.
  2. Freeze the instructions. Both models get the same repo AGENTS.md, the same task text, and the same reasoning effort.
  3. Run each ticket once per model in a non-interactive run, for example codex exec -m gpt-6.1-sol, so the runs are repeatable.
  4. Record tokens used, retries, whether a reviewer accepted the change, and review minutes.
  5. Calculate cost per accepted task for each model. Include failed runs in the cost.
  6. Write down the decision and the task types where Astra earned its price.

Two mistakes to avoid. Do not compare Sol at max effort with Astra at low effort. And do not let the person who wrote the prompt be the only reviewer.

Where the policy lives, and where it does not

A model policy has two parts: the configuration that sets defaults and the review step that checks them.

In Codex, a model = "gpt-6.1-sol" line in config.toml sets the default. Administrators can supply managed defaults under [models.new_thread]. The Codex configuration reference is explicit that these are defaults, not enforcement. A person who passes --model or a profile overrides them.

Two other gaps are worth writing into the policy. Enterprise model controls and defaults do not apply to ChatGPT dots. And GPT-5.5 retires from ChatGPT, ChatGPT Work, and Codex on 14 October 2026, so scripts and scheduled tasks that name it need an owner before then.

A short policy note is enough:

# Model policy (review every quarter)

Default: gpt-6.1-sol at medium effort.
Escalate to gpt-6-astra: cross-service changes, security or payment
code, or two failed Sol attempts with the same evidence.
High-volume summaries: gpt-6-luna.
Every agent change: same review standard, whichever model wrote it.
Owner: <name>. Next review: <date>.

What should you measure after the switch?

Track cost per accepted change, escalation rate to Astra, rework after first review, and reviewer minutes. A falling token bill with rising rework is not a saving. Our guide to measuring an AI workflow before scaling it explains how to keep those definitions stable.

Look at the escalation rate monthly. If it climbs, your task mix may have changed, or the rule is too loose. If it stays near zero on hard work, check that people are not forcing Sol through tasks it keeps failing.

Start by pulling ten closed tickets this week and running the comparison above. If you want help setting the evaluation up for your team, book a short call.

Further reading