On September 29, OpenAI introduced GPT-6.1 Sol at DevDay 2026: an upgrade to GPT-6 Sol that, according to OpenAI, nearly matches GPT-6 Astra's intelligence on agentic coding, computer use, and professional workloads, at one-fifth of Astra's standard input and output token prices.
Price is the heart of this release. Standard API pricing is $2 per million input tokens and $10 per million output tokens, with cached input at just $0.10 — 95% less than standard input pricing and half of GPT-6 Sol's cached input price. OpenAI says this gives developers more room to build and run agents that reuse context across requests. As of September 29, GPT-6.1 Sol is available to Plus, Pro, Business, Enterprise, and Edu users in ChatGPT Work and Codex; developers can access it via the API as gpt-6.1-sol. It is not yet available in Chat. A GPT-6.1 Sol Ultrafast option, with up to 8x faster token generation in Codex, is coming in the next few days.
Benchmarks: Where It Caught Up With Astra
- DeepSWE v1.1 (long-horizon software engineering tasks in real codebases): GPT-6.1 Sol matches Astra at roughly one-fifth the cost, beating GPT-6 Sol's best score by 6.4 percentage points at lower reasoning effort.
- GDP.pdf (answering professional questions from complex PDFs with tables, charts, and fine print): scores above Opus 5.5 with fallbacks at less than half the cost per task, approaching Astra's state-of-the-art performance at roughly one-fifth the cost.
- AutomationBench (multi-step business workflows across 47 tools spanning sales, marketing, and operations): 2.2 percentage points above Opus 5.5 at medium reasoning effort at about a third of the cost; 4.8 points above GPT-6 Sol.
- OSWorld 2.0 offline set (long-horizon computer-use workflows): 7 percentage points above GPT-6 Sol; within 2.1 points of Astra at maximum reasoning effort at roughly one-seventh the cost per task.
- Terminal-Bench Science 0.1 (scientific workflows like data analysis, simulation, and theorem proving): more than double GPT-6 Sol's score, at an average $5.47 per task versus $23.21 for Opus 5.5 and $23.80 for Astra.
Factuality and Safety: Fewer Errors, Better Behaved
In factuality evaluations, the share of answers containing a factual error at low reasoning effort fell from 11.4% to 7.7% — a reduction of about 32% — staying within 1.9 percentage points of Astra across settings. Alignment evaluations show GPT-6.1 Sol failing less often than GPT-6 Sol on challenges like disclosing a broken search tool, respecting explicit restrictions, and avoiding unauthorized outcomes in agentic tasks; OpenAI says it observed no attempts to bypass an automated safety reviewer, matching Astra and GPT-6 Sol.
What It Means: After Astra Hit Pause, Sol Becomes the Pragmatic Pick
Just days earlier, OpenAI canceled the planned GPT-6.1 Astra release after it failed to meet internal safety and alignment standards. Astra stands for maximum intelligence; Sol stands for good-enough intelligence at a price you can actually use. For developers, the real signal is a changed cost structure for long-horizon agent tasks: cached input at $0.10 per million tokens turns "feeding an entire codebase to an agent for weeks" from a luxury into a routine operation. The tradeoff is equally clear: it is still not Astra — Astra's 68.1% on Terminal-Bench Science remains the top score, and OpenAI still recommends Astra for the hardest scientific research.