OPENAI · MODEL RELEASE TRACKER
GPT-6.1 Sol
GPT-6.1 Sol is an OpenAI model positioned as a major upgrade to GPT-6 Sol, offered via Amazon Bedrock. It is presented as delivering strong reasoning for agentic coding, computer use, and professional workflows while matching higher-tier models on specific benchmarks at lower cost per task.CURRENT SNAPSHOT3/5 DIMENSIONS WITH DATA
The dimensions that change the decision.
Not established from the available sources.
Not established from the available sources.
RELEASE TIMELINE
Published, source-backed release events only.
Google DeepMind unveils Gemini 4 Argon, a frontier model with 1M-token context
Google DeepMind announced Gemini 4 Argon, a new frontier multimodal model optimized for long-horizon reasoning and complex workflows. Argon is rolling out initially to trusted cyber defenders via the Fairwind Program, expands context length to 1 million tokens, reports leading benchmark performance across coding, finance, legal, and video understanding, and will be made more widely available after phased safety testing and engagement with U.S. government pre-release processes; Google also published introductory pricing.
OpenAI launches Dots — always-on assistants powered by GPT-6 Astra
OpenAI unveiled Dots at DevDay: always-on, agentic assistants powered by the GPT-6 Astra model that run on their own cloud computer, can access a web browser and 4,000+ apps, and interact via text or voice across ChatGPT, Teams, and Slack. Dots are rolling out to ChatGPT Pro and Business Premium users, start with one per account, include custom rules and an auto-review safety feature, and conversations with a Dot do not count toward ChatGPT usage limits.
VERIFIABLE FACTS
Every value stays attached to a source and date.
According to AWS, GPT-6.1 Sol is generally available on Amazon Bedrock.
Runs on an inference engine on Amazon Bedrock described as built for performance, security, and reliability at scale.
Developers can access the model in the API as 'gpt-6.1-sol'; an 'Ultrafast' variant for faster token generation in Codex was announced as forthcoming.
OpenAI published API pricing for GPT-6.1 Sol at $2 per million input tokens and $10 per million output tokens; the company also says Sol delivers near-Astra performance at roughly one-fifth the standard input/output token prices.
OpenAI reported that Sol ties Astra on the DeepSWE v1.1 coding benchmark (at ~one-fifth the cost), scores 6.4 percentage points higher than GPT-6 Sol on that benchmark, beats GPT-6 Sol by ~7 points on OSWorld 2.0 and sits ~2.1 points behind Astra, outperforms Anthropic's Opus 5.5 on a GDP.pdf document benchmark at lower cost, finishes 2.2 points ahead of Opus 5.5 on AutomationBench (medium reasoning), and more than doubles GPT-6 Sol on Terminal-Bench Science.
OpenAI said GPT-6.1 Sol delivers significant improvements over GPT-6 Sol on complex tasks including programming and debugging, understanding documents, and executing multi-step workflows; the company claims Sol approaches GPT-6 Astra's performance on agentic coding, computer use, and professional work.
OpenAI reported that at low reasoning effort the share of responses containing a factual error fell from 11.4% (GPT-6 Sol) to 7.7% (GPT-6.1 Sol), and that across all reasoning settings Sol's error rate stayed within 1.9% of GPT-6 Astra.
OpenAI said GPT-6.1 Sol is more upfront about limitations and more reliable at honoring user intent and safety constraints than GPT-6 Sol in challenging evaluations (e.g., flagging broken search tools, following explicit restrictions, avoiding unauthorized outcomes); OpenAI reported observing no attempts by Sol to circumvent its automated safety reviewer.
Described as a major upgrade to GPT-6 Sol that delivers strong performance in agentic coding, computer use, and professional workloads; positioned to bring near‑Astra intelligence to everyday workflows.
According to OpenAI (via the AWS post), GPT-6.1 Sol matches GPT-6 Astra on DeepSWE v1.1 at roughly one‑fifth the cost per task, and exceeds the best score from GPT-6 Sol by 6.4 percentage points while using lower reasoning effort.
WHAT CHANGED