· 6 min read
OpenAI Ultrafast mode: how the GPT-6 Astra tier works
OpenAI Ultrafast mode makes GPT-6 Astra stream output up to several times faster at six times the price. How it works, what it costs and when to use it.

OpenAI's Ultrafast mode is a premium service tier for GPT-6 Astra that speeds up how quickly the model streams its output. You turn it on by setting service_tier to "ultrafast" on a Responses API request, and you pay six times the Standard token price for it. It shipped on 29 September 2026 and is meant for latency-sensitive work: interactive coding agents, tool-heavy loops and anything where a person is waiting on the answer.
The switch is one field. Deciding when it is worth six times the price takes a bit more thought, so here is how the tier works, what it costs, and how to use it without wasting money.
What is OpenAI Ultrafast mode?
A service tier is a processing option you choose per request that trades price for speed or cost. The OpenAI API now has five of them: Batch, Flex, Standard, Fast and Ultrafast. Batch and Flex are cheaper and slower. Standard is the default. Fast, which was called priority processing until July 2026, costs more and is documented as up to 2.5x faster with more consistent latency.
Ultrafast sits above Fast. It is generally available for GPT-6 Astra, OpenAI's top model, with other models in limited preview. OpenAI's launch post quoted up to 6x faster in the API, and the Ultrafast guide says up to 8x faster than Standard. Treat both numbers as ceilings, not promises, and measure your own workload.
The changelog is specific about what changes: Ultrafast reduces the time between output tokens. That matters because a model response has two latency parts. Time to first token is how long you wait before anything streams back, and it covers network, queueing and reading your prompt. Inter-token latency is the gap between each generated token after that. Ultrafast attacks the second part. Long outputs gain the most; a one-word answer barely changes.
How do you turn on Ultrafast mode?
Set the model and the tier on the request. Nothing else in the payload changes:
from openai import OpenAI
client = OpenAI()
response = client.responses.create(
model="gpt-6-astra",
service_tier="ultrafast",
input="Refactor this function to remove the nested loops.",
)
print(response.output_text)
That works over plain HTTP with the SDKs or curl. For agents, OpenAI recommends going further and using WebSocket mode, a way of talking to the Responses API over one persistent connection instead of a new HTTP request per turn.
The reason is simple arithmetic. An agent that makes twenty quick tool calls pays connection and request overhead twenty times. If each turn generates only a short tool call, most of the wall-clock time is that overhead, and a faster decoder cannot help with it. OpenAI's own guide warns that without a persistent connection, network overhead can eat the latency gains.
import OpenAI from "openai";
import { ResponsesWS } from "openai/resources/responses/ws";
const client = new OpenAI();
const ws = new ResponsesWS(client);
const events = ws.stream();
ws.send({
type: "response.create",
model: "gpt-6-astra",
service_tier: "ultrafast",
previous_response_id: lastResponseId,
input: toolOutputs,
});
for await (const event of events) {
if (event.type !== "message") continue;
if (event.message.type === "response.completed") {
lastResponseId = event.message.response.id;
break;
}
}
In WebSocket mode, each turn is a response.create event. You continue a chain by sending previous_response_id plus only the new input, such as tool results. The server keeps recent response state in memory for that connection, so it does not reload the whole conversation every turn. A few limits to plan around: a connection lasts at most 60 minutes, holds up to 16 in-flight responses, and loses its cached state when it closes. If you run with store=false, a reconnect means you need to resend context, because the old response IDs are gone.
What does Ultrafast cost?
This is where the decision gets real. Prices per million tokens for GPT-6 Astra, short context:
| Tier | Input | Cached input | Output | Notes |
|---|---|---|---|---|
| Standard | $10 | $1 | $50 | Default |
| Fast | premium over Standard | discounted | premium over Standard | Up to 2.5x faster; no latency SLA on GPT-6 Astra |
| Ultrafast | $60 | $6 | $300 | Up to 6-8x faster output |
Long-context requests cost more again, at $120 input and $450 output per million on Ultrafast.
Put numbers on a single turn. A response that reads 20,000 input tokens and writes 2,000 output tokens costs about $0.30 on Standard and about $1.80 on Ultrafast. Across an agent session with fifty turns, that is the difference between roughly $15 and $90.
Cached input still gets its discount on Ultrafast, which makes caching more valuable, not less. If your agent resends a long system prompt and tool definitions every turn, getting those cached cuts the biggest line item. The mechanics are the same as in prompt caching on GPT-6: keep the stable prefix stable and put the changing parts last.
What are the limits of Ultrafast mode?
Three constraints are worth knowing before you build on it.
- Rate limits are separate and modest. Default Ultrafast limits for GPT-6 Astra are 500,000 tokens per minute on the Build usage tier, 1,000,000 on Launch and 5,000,000 on Grow. Higher limits go through OpenAI's account team. If you put Ultrafast behind a shared gateway, budget for these the way you would any other quota, as covered in rate limiting with token buckets.
- Data residency is narrow. Ultrafast supports global processing and US data residency only. EU and other regional endpoints are not available, so teams with regional requirements cannot use it yet.
- Check what you actually got. Responses carry a
service_tierfield. OpenAI documents that Fast requests can be downgraded to Standard when traffic ramps too quickly, and the response then reportsdefault. The Ultrafast guide does not describe a downgrade path, but logging the field costs nothing and tells you whether you paid for the speed you asked for.
When should you use Ultrafast mode?
The tier is priced for cases where waiting is the expensive part.
- Use it for interactive agents where a developer watches tokens stream, for long generated outputs such as code edits where decode speed dominates, and for multi-step tool loops on WebSockets where each step's output gates the next.
- Skip it for background jobs, evals, nightly pipelines and anything nobody waits on. Batch or Flex will do that work far cheaper.
- Try Fast first if your latency problem is mostly inconsistency rather than raw speed. It costs less and may be enough.
A practical pattern is to choose the tier per request, not per project. The same agent can run its planning step and final answer on Ultrafast, where a user is watching, and its background summarisation on Standard.
def pick_tier(user_is_waiting: bool, expected_output_tokens: int) -> str:
if not user_is_waiting:
return "flex"
if expected_output_tokens > 500:
return "ultrafast"
return "default"
The thresholds are yours to tune. The point is that the tier is a request parameter, so you can spend the premium only where it buys something.
Key takeaways
- Ultrafast is a
service_tiervalue for GPT-6 Astra on the Responses API, released 29 September 2026. - It speeds up the gap between output tokens, so long outputs gain the most.
- It costs six times Standard: $60 input and $300 output per million tokens on short context.
- Pair it with WebSocket mode for tool-heavy agents, or network overhead hides the gain.
- Pick the tier per request and log the
service_tierthe response reports.
FAQ
Which models support OpenAI Ultrafast mode?
GPT-6 Astra has general availability. Other models are in limited preview, and access goes through OpenAI's account team.
Does Ultrafast make time to first token faster?
OpenAI's changelog describes it as reducing the time between output tokens. Time to first token still depends on your prompt size, network and connection reuse, which is why WebSocket mode is recommended alongside it.
Can I use Ultrafast with EU data residency?
No. At launch Ultrafast supports global processing and US data residency only.