Introducing GPT-6.1 Sol
- ID
- 29901
- Status
- summarized
- Published
- 29 Sep 2026, 6:00 PM
- Fetched
- 30 Sep 2026, 2:08 AM
- Provider
- OpenAI News
- Category
- ai-labs
- Original URL
- https://openai.com/index/introducing-gpt-6-1-sol
- Source URL
- https://openai.com/news/rss.xml
Summary
- Score
- 5.0
- Created
- 30 Sep 2026, 2:08 AM
- Tags
- Audience
- developersai_agent_usersai_ml_learnerssaas_founders
What happened
OpenAI announced GPT-6.1 Sol, an upgrade to GPT-6 Sol that it claims nearly matches GPT-6 Astra on agentic coding, computer use and professional work at one-fifth of Astra's standard input/output token prices. Cached input is listed at $0.10 per million tokens, which OpenAI says is 95% below its standard input pricing and 50% below GPT-6 Sol's cached rate. The post cites self-reported results including matching GPT-6 Astra on DeepSWE v1.1 at roughly one-fifth the cost, beating GPT-6 Sol's best DeepSWE score by 6.4 percentage points at lower reasoning effort, and scoring 2.2 points above Opus 5.5 on AutomationBench at medium effort for about a third of the cost; the excerpt cuts off mid-sentence in the OSWorld 2.0 computer-use section, so those numbers are not visible here.
Why it matters
The only decision-grade number in this post is cached input at $0.10 per million tokens, 50% below GPT-6 Sol's cached rate — if your agent loop resends the same system prompt, tool schemas or document context on every call, that is the line item that changes your bill, not the headline token price. Every capability claim (DeepSWE v1.1, GDP.pdf, AutomationBench) is OpenAI's own benchmark run with no independent replication, and the OSWorld 2.0 section is truncated, so treat this as a reason to re-run your own eval on one cached-context workload, not as a reason to migrate production traffic.
Discussion angle
Pull your own token logs and estimate what fraction of input tokens are cached-context reuse — if it's high, this pricing change is worth a test; if it's low, the 95%-off cached rate is irrelevant to you. Then ask what independent evidence (someone else's DeepSWE v1.1 or AutomationBench run, or the missing OSWorld 2.0 numbers) would be needed before switching a production agent off your current model.