GPT-6.1 Sol gets an Ultrafast tier: 8x speed, 6x price
OpenAI switched on Ultrafast for GPT-6.1 Sol: up to 8x speed at $12/$60 per million tokens, gated to the $500 Pro 500 plan in ChatGPT and Codex.

On October 8-9, OpenAI brought its Ultrafast high-speed mode down from the flagship Astra to the volume tier, GPT-6.1 Sol: developers can call gpt-6.1-sol with Ultrafast enabled and get up to 8 times standard speed while the company claims “near-Astra intelligence,” per coverage by Machine Heart and others. The price jumps with it — $12 per million input tokens and $60 per million output tokens, six times the standard tier. It is day four of product lead Tibo’s “28 days of shipping” streak, one day after GPT-6 reached every ChatGPT user.
Key points
- 8x the speed, 6x the price: Sol Ultrafast costs $12/$60 per million input/output tokens and runs up to eight times faster than standard Sol.
- A gate on the consumer side: the API is open to all developers at metered rates; ChatGPT and Codex access is limited to the $500/month Pro 500 plan plus some enterprise and education contracts, and developers are already complaining about how fast the usage allowance burns.
- Whose silicon: per SemiAnalysis analysis relayed by OfficeChai, this Ultrafast generation runs on Nvidia GPUs; the August preview of GPT-5.6 Sol Ultrafast ran on Cerebras.
Background
Ultrafast debuted as a preview tier with GPT-5.6 Sol on August 13, in partnership with Cerebras, with a cited 14x speedup; at DevDay on September 29 the mode landed on the flagship Astra first. Sol is OpenAI’s volume tier — per VentureBeat’s late-September report, GPT-6.1 Sol delivers near-Astra performance at about one-fifth of Astra’s price, at roughly 300 output tokens per second. Extending Ultrafast to the volume tier turns low latency from a flagship privilege into a line item. Worth remembering: Astra itself was pulled briefly over safety concerns, and its Ultrafast tier has had an EU data-residency gap.
The facts
| Tier | API price per 1M tokens (in / out) | Notes |
|---|---|---|
| GPT-6.1 Sol standard | $2 / $10 | volume tier, DevDay pricing |
| GPT-6.1 Sol Ultrafast | $12 / $60 | up to 8x speed, new |
| GPT-6 Astra (reference) | $10 / $50 | flagship; OpenAI puts Sol near it |
Details worth checking yourself: per MIXED Reality News, EU data residency is supported on Sol Ultrafast while Astra’s Ultrafast lacks it — European enterprises that need low latency currently have one option. On the API side, Ultrafast is metered with no usage cap; on the consumer side, the gate is the $500/month Pro 500 subscription plus some metered enterprise and education contracts. Machine Heart relays the early developer gripe: one user says Pro 500’s 25x usage allowance can burn “1% in a few minutes” on the fast tier — speed is a direct multiplier on token spend.
What others say
OfficeChai relays SemiAnalysis’s supply-chain read that this generation of Ultrafast runs on Nvidia GPUs rather than Cerebras, whose stock has already swung on earlier reports of an OpenAI shift to Nvidia; if accurate it confirms a partner migration, though OpenAI has not confirmed it. MIXED Reality News approaches from compliance: the residency difference between tiers will shape procurement for European companies. Machine Heart captures developer sentiment — at six times standard pricing on the API and a top-tier subscription on the consumer side, the reaction in both places has been vocal.
Our take
Paying for latency is now a first-class axis of the API, and that matters more for engineering decisions than another model bump: real-time voice, coding agents and browser automation can now run cost-benefit math between two price tags on the same model — the extra dollars per second have to pay for themselves in wait time and task throughput. Who should care: teams running interactive agents. Who can wait: batch and analytics workloads, where standard Sol’s economics are unchanged. Honest uncertainty: “up to 8x” is the company’s framing and there is no third-party percentile-latency benchmark under real load yet; the Nvidia claim comes from an analysis firm, not OpenAI.
How to try it
API users: call gpt-6.1-sol with Ultrafast enabled — metered, no usage cap. ChatGPT and Codex users need Pro 500 ($500/month) or a qualifying enterprise or education contract. Pressure-test p99 latency with your real workload before switching anything latency-critical over.