OpenAI Introduces Ultrafast Speed Tier for Codex and API
OpenAI has launched Ultrafast, a new premium speed tier offering significantly increased token generation rates for Codex and API users.
Ultrafast mode optimizes token throughput for high-performance development tasks.
- Ultrafast is a new premium speed tier offering up to 8x faster token generation in Codex and 6x in the API.
- The mode reaches speeds of 300 tokens per second in Codex.
- Access is available through the new Pro 500 plan and for all developers via the API.
- Integration is supported for GPT-6 Astra, with GPT-6.1 Sol support arriving soon.
Overview of Ultrafast Mode
OpenAI has officially unveiled "Ultrafast," a new performance-focused tier designed to accelerate token generation across its model lineup. The initiative, announced in late September 2026, aims to improve development workflows and support real-time application requirements by reducing the latency between input and output generation.
According to official documentation and developer announcements, Ultrafast is currently integrated with GPT-6 Astra, with support for the upcoming GPT-6.1 Sol model expected to follow. The tier represents a significant step in the company's effort to provide more granular control over model inference speeds for heavy users and enterprise developers.
Performance Benchmarks
The primary benefit of the Ultrafast tier is its increased throughput. For users operating within the Codex environment, OpenAI reports token generation speeds reaching up to 300 tokens per second. This represents an improvement of up to 8x compared to standard operating modes. For those utilizing the OpenAI API, the service offers up to 6x faster performance compared to standard configurations. These figures are calculated against existing benchmarks, specifically contrasting Ultrafast with "Astra Standard" and "Astra Fast" variations.
Access and Availability
OpenAI has structured access to Ultrafast through a combination of new consumer plans and standard API access. For individual users and developers looking for high-capacity environments, the company introduced the "Pro 500" plan. This subscription tier is positioned as a high-usage alternative, offering up to 25 times the usage limits compared to the standard Plus plan, while including direct access to the Ultrafast features in both Codex and ChatGPT Work environments.
For developers building on the platform, Ultrafast is accessible via the API. The company encourages developers to pair this mode with the WebSocket-based "Responses API" to facilitate the creation of real-time intelligence experiences. This combination is intended to allow for more seamless, high-speed iteration when moving from an initial idea to code execution.
Strategic Context
The launch of Ultrafast coincides with updates to the broader OpenAI subscription ecosystem. Alongside the introduction of the Pro 500 plan, the company has reopened the Pro 200 subscription, which provides access to frontier models like Astra. By segmenting their service offerings, OpenAI appears to be addressing the divergent needs of casual users and power users who require higher speed and throughput for industrial-scale code generation or agentic workflows.
As of late 2026, the documentation indicates that while Ultrafast is optimized for current Astra deployments, the infrastructure is designed to accommodate new models as they reach production readiness. Developers are advised to review the API reference guides regarding latency optimization and WebSocket integration to properly leverage the new performance tiers.
Enjoyed this?
Get more posts like this delivered to your inbox.