GPT-6.1 Sol costs $2 per million input tokens, $10 per million output tokens, and $0.10 per million cached input tokens: the same list price as Claude Sonnet 5.5 and half that of Opus 5.5. In most turns of a code agent it’s cheaper than both, except when the prompt exceeds 272K input tokens, where OpenAI’s surcharge reverses the outcome.
OpenAI launched GPT-6.1 Sol at DevDay on September 29, a week after GPT-6 Sol. The headline OpenAI chose is “near-Astra intelligence at a fifth of the price”. For those paying per token, that’s not the useful comparison. The useful one is against the model you’re already running inside your agent, which on many teams is a Claude model. And against Claude, the answer depends on how much your context grows.
What is GPT-6.1 Sol?
GPT-6.1 Sol is the middle tier of OpenAI’s GPT-6 family, between GPT-6 Astra (the most capable and most expensive) and GPT-6 Luna (the cheap one, built for volume). It’s an update to GPT-6 Sol. OpenAI presents it as a model that nearly matches Astra in agentic coding, computer use, and professional work, at a fifth of Astra’s standard input and output prices. If you want the other side of that comparison, we break it down in how much GPT-6 Astra costs.
Data from the model page in OpenAI’s documentation:
- Model ID:
gpt-6.1-sol - Context: 1,050,000 token window and up to 128,000 output tokens
- Knowledge cutoff date: April 30, 2026
- Reasoning effort:
low,medium(default),high,xhigh, andmax. Does not supportnoneorminimal. - Tool calling: via the Responses API. Chat Completions works, but without tool calling.
How much does GPT-6.1 Sol cost?
| Per 1M tokens | GPT-6.1 Sol |
|---|---|
| Input | $2.00 |
| Cached input | $0.10 |
| Cache writes | $2.50 |
| Output | $10.00 |
Cached input costs 5% of the standard input rate: that’s the 95% discount OpenAI highlights. Cache writes are charged at 1.25 times the input rate. Batch and Flex cost 50% less than the standard tier, and Fast mode costs double.
And there’s a line that changes the math in long sessions: prompts over 272K input tokens are charged double on input and cache, and 1.5 times on output, on the entire request, not just on tokens exceeding the threshold.
These are the prices as of September 30, 2026. GPT-6 Sol was replaced seven days after its launch, so check OpenAI’s pricing page before building a budget on this table.
What changed from GPT-6 Sol?
In price, only one number moved. Standard input and output stay at $2 and $10; cached input dropped by half, from $0.20 to $0.10. If your workload is mostly new prompts, the update costs you the same. If it’s an agent that rereads the same repository context each turn, your bill goes down.
In capability, OpenAI reports the following. All figures are from OpenAI itself and there’s still no independent replication:
- DeepSWE v1.1: 6.4 points above GPT-6 Sol’s best score, with lower reasoning effort and lower cost. OpenAI says it matches Astra there at around a fifth of the cost.
- OSWorld 2.0 offline (computer use): 7 points above GPT-6 Sol at maximum effort, and 2.1 points from Astra at around a seventh of the cost per task.
- Terminal-Bench Science 0.1: more than double GPT-6 Sol’s score at maximum effort.
- Factuality: at low effort, the proportion of responses with at least one factual error drops from 11.4% to 7.7%.
Claude vs ChatGPT: Is GPT-6.1 Sol cheaper than Sonnet 5.5 and Opus 5.5?
Yes, in most turns, because the only thing that changes from Sonnet 5.5 is the cached input read price, and there GPT-6.1 Sol charges half. These are list prices, taken from Anthropic’s pricing documentation:
| Per 1M tokens | GPT-6.1 Sol | Sonnet 5.5 | Opus 5.5 |
|---|---|---|---|
| Input | $2 | $2 | $4 |
| Output | $10 | $10 | $20 |
| Cache writes | $2.50 | $2.50 (5 min) | $5 |
| Cached input read | $0.10 | $0.20 | $0.20 |
The first two rows are identical to Sonnet 5.5’s. And an agentic coding session is mostly cached input reads: each turn resends the conversation, the repository context, and the tool results.
I used the illustrative turn from our Sonnet 5.5 vs Opus 5.5 analysis and moved the cached context up and down. It’s my arithmetic with list prices, not a measurement:
| Cached context (+5K input, 2K output) | GPT-6.1 Sol | Sonnet 5.5 | Opus 5.5 |
|---|---|---|---|
| 100K | $0.040 | $0.050 | $0.080 |
| 265K | $0.057 | $0.083 | $0.113 |
| 300K | $0.110 | $0.090 | $0.120 |
What if your context exceeds 272K tokens?
Up to the threshold, GPT-6.1 Sol clearly wins, and its advantage grows as the cached context grows. Above 272K, the entire request shifts in price and Sonnet 5.5 becomes the cheapest. The calculation assumes the cached context counts toward the 272K, which it should, because it’s part of the prompt. Anthropic, meanwhile, charges Claude’s full 1M token window at standard rate. It’s a cliff, not a slope, and long sessions with Codex or Claude Code approach it without anyone noticing.A caveat before taking the table as a verdict: token prices only compare cleanly if both models count the same text as the same number of tokens, and they don’t. Each provider uses its own tokenizer; Anthropic itself points out that its current tokenizer generates about 30% more tokens than the previous one for the same text. Run your own repository through both before drawing conclusions.
If what you’re looking for is a broader comparison between assistants and not just prices, we did one in which is the best AI for programming.
Is GPT-6.1 Sol better than Opus 5.5 on benchmarks?
On the tasks OpenAI chose to publish, GPT-6.1 Sol comes out ahead of Opus 5.5 and at lower cost. These are OpenAI’s measurements of a competitor’s model:
- AutomationBench: 2.2 points above Opus 5.5 at medium effort, at about a third of the cost.
- GDP.pdf: higher score than “Opus 5.5 with fallbacks”, at less than half the cost per task.
- Terminal-Bench Science: $5.47 per task at maximum effort, versus $23.21 for Opus 5.5 and $23.80 for Astra. That’s where the “more than 75% cheaper” figure circulating these days comes from. It applies to a benchmark, not the model in general.
In that same benchmark, OpenAI acknowledges that Astra still has the highest score, 68.1%, and should remain the option for the most difficult research. It’s the provider itself telling you where their cheap model ends.
Can you use GPT-6.1 Sol in Codex?
Yes. It’s available in ChatGPT Work and in Codex for Plus, Pro, Business, Enterprise, and Edu users, and in the API as gpt-6.1-sol. At the time of publishing this note, it’s still not available in the regular ChatGPT chat.
The fastest tier still hasn’t arrived. OpenAI’s Ultrafast tier offers token generation up to 8 times faster in Codex (about 300 tokens per second) and up to 6 times in the API, but today it only runs on GPT-6 Astra, in the Pro 500 and Enterprise plans. As of September 30th, OpenAI lists GPT-6.1 Sol Ultrafast as “coming soon”.
A practical distinction: if you use Codex with a ChatGPT subscription, the token counts above don’t affect you directly. They matter when you run agents with your own API keys, or in CI, bots, and pipelines. We covered that setup in Claude Code vs Codex: how to run both with your own keys.
What should a technical team do with this?
My take: GPT-6.1 Sol makes the decision by token more interesting, not simpler.
- For well-scoped agent work, with less than 272K tokens of context (review bots, agents in CI, specified features, bugs with known reproduction), today it’s the cheapest capable option by list price. It deserves a test with your own workload.
- For long sessions, context size becomes a cost variable you have to manage on purpose. Compression, new sessions, and tighter repository context were hygiene; with GPT-6.1 Sol they decide whether a turn costs $0.06 or $0.11.
- For the most difficult open work, none of this changes what we said last week: the providers themselves point to Opus 5.5 and Astra as the models for sustained judgment.
The underlying pattern for CTOs is that two providers now sell their mid-tier models at almost identical list prices, and the difference has moved to caching rules and thresholds. Those rules go into a written routing policy that your team reviews every time a provider changes their pricing page. At the current rate, that’s every couple of weeks.