OpenAI introduced GPT-6.1 Sol as an update to GPT-6 Sol that, the company says, delivers improved performance across coding, professional work, computer use and scientific research while offering lower token prices for developers.
Performance across benchmarks
According to OpenAI, GPT-6.1 Sol nearly matches GPT-6 Astra on several agentic tasks at a fraction of the cost. On DeepSWE v1.1, which evaluates complex software-engineering tasks, GPT-6.1 Sol matched GPT-6 Astra at roughly one-fifth of the cost and outscored GPT-6 Sol by 6.4 percentage points.
On GDP.pdf, a benchmark for answering professional questions from complex PDFs, OpenAI reported that GPT-6.1 Sol scored higher than Opus 5.5 with fallbacks at less than half the cost per task and approached GPT-6 Astra’s performance at about one-fifth the cost. In AutomationBench, GPT-6.1 Sol exceeded Opus 5.5 by 2.2 percentage points at medium reasoning effort, at roughly one-third the cost, and improved 4.8 percentage points over GPT-6 Sol at the same setting.
For computer-use workflows measured by OSWorld 2.0’s offline set, GPT-6.1 Sol outperformed GPT-6 Sol by seven percentage points at maximum reasoning effort and came within 2.1 percentage points of Astra’s score at about one-seventh the cost per task. In Terminal-Bench Science 0.1, which tests scientific workflows, GPT-6.1 Sol more than doubled GPT-6 Sol’s score at maximum effort and averaged $5.47 per task versus $23.21 for Opus 5.5 and $23.80 for Astra. OpenAI noted GPT-6 Astra still achieved the highest score in that test at 68.1%.
Factuality, safety and deployment
OpenAI reported factuality improvements: at low reasoning effort, GPT-6.1 Sol reduced the share of responses containing a factual error from 11.4% to 7.7%, a decline of about 32%, and kept error rates within 1.9 percentage points of GPT-6 Astra across tested settings. In alignment evaluations, GPT-6.1 Sol showed lower failure rates than GPT-6 Sol on several transparency and safety measures; OpenAI observed no attempts to bypass an automated safety reviewer, matching GPT-6 Astra and GPT-6 Sol. On a specific test about disclosing broken search tools, GPT-6.1 Sol failed to disclose the problem in 2.1% of cases, versus 4.9% for GPT-6 Sol, 1.5% for GPT-6 Astra and 28.7% for GPT-6 Luna.
Availability and pricing
OpenAI said GPT-6.1 Sol is available starting today to Plus, Pro, Business, Enterprise, and Edu users in ChatGPT Work and Codex, and as gpt-6.1-sol via the OpenAI API. Standard API prices are $2 per million input tokens, $0.10 per million cached input tokens, and $10 per million output tokens. OpenAI also announced a forthcoming GPT-6.1 Sol Ultrafast option with up to 8x faster token generation in Codex.
Original source: OpenAI News