OpenAI has introduced GPT-6.1 Sol and Ultrafast mode within its Responses API, offering a significant boost in speed and cost efficiency for production AI workloads. The new model and mode are now available, with a comprehensive 13-step setup guide to help developers migrate from the older Chat Completions endpoint.
TL;DR
- OpenAI launches GPT-6.1 Sol and Ultrafast mode within its Responses API, priced at $2 per million input tokens and $10 per million output tokens.
- The new model and mode offer faster processing times and lower costs for production AI workloads.
- A 13-step setup guide is provided to help developers migrate from the older Chat Completions endpoint.
What happened
OpenAI has released GPT-6.1 Sol, a mid-tier model positioned between the flagship GPT-6 Astra and the budget GPT-6 Luna. Priced at $2 per million input tokens and $10 per million output tokens, Sol offers near-Astra performance at a lower cost. Cached input tokens on Sol run at a 95% discount, costing just $0.10 per million.
On October 8, 2026, OpenAI introduced Ultrafast mode within the Responses API. This mode reduces the time between generated output tokens, making it ideal for latency-sensitive interfaces like voice agents or live coding assistants. Ultrafast mode is available to all API users and works across both US and EU data residency regions.
The Responses API is designed to replace both the older Chat Completions API and the deprecated Assistants API. It offers a unified endpoint for text generation, tool calls, file inputs, multi-turn state, and background processing. The API also includes features like server-side storage of responses and typed response objects for easier reference and management.
Why it matters
The introduction of GPT-6.1 Sol and Ultrafast mode provides developers with more options for optimizing the speed and cost of their AI workloads. The new model and mode are particularly beneficial for enterprises looking to deploy generative AI applications in production, as they offer a balance between performance and cost.
The 13-step setup guide provided by OpenAI makes it easier for developers to migrate from the older Chat Completions endpoint. This guide covers everything from creating API keys and installing SDKs to sending the first API call, turning on Ultrafast mode, and shipping a complete working project.
The Responses API's unified endpoint and advanced features make it a more powerful and flexible tool for building AI applications. Developers can now handle text generation, tool calls, file inputs, multi-turn state, and background processing all in one place, simplifying the development process.
Key facts
- GPT-6.1 Sol is priced at $2 per million input tokens and $10 per million output tokens.
- Cached input tokens on Sol cost $0.10 per million, a 95% discount off the standard input rate.
- Ultrafast mode reduces the time between generated output tokens, making it ideal for latency-sensitive interfaces.
- The Responses API offers a unified endpoint for text generation, tool calls, file inputs, multi-turn state, and background processing.
- The API includes features like server-side storage of responses and typed response objects for easier reference and management.
- The 13-step setup guide covers creating API keys, installing SDKs, sending the first API call, turning on Ultrafast mode, and shipping a complete working project.
- GPT-6.1 Sol supports function calling, web search, file search, and computer-use tools natively, with data residency limited to the US and the EU.
- The model has a 1.05-million-token context window and a 128,000-token maximum output, with a knowledge cutoff of April 30, 2026.
Context
The launch of GPT-6.1 Sol and Ultrafast mode comes at a time when enterprises are increasingly adopting generative AI APIs. According to Gartner, more than 80% of enterprises are expected to have used generative AI APIs or deployed generative-AI-enabled applications in production by 2026, up from under 5% in 2023.
The Responses API is part of OpenAI's efforts to provide a more unified and powerful tool for building AI applications. By replacing the older Chat Completions API and the deprecated Assistants API, the Responses API offers a more streamlined and feature-rich solution for developers.
The introduction of GPT-6.1 Sol and Ultrafast mode is also a response to the growing demand for more efficient and cost-effective AI solutions. As enterprises look to deploy AI applications at scale, the need for models and modes that offer a balance between performance and cost has become increasingly important.
