· 6 min read
GPT-6 prompt caching: breakpoints, pricing and cache misses
How GPT-6 prompt caching works: explicit breakpoints, the 30-minute TTL, 1.25x writes vs 0.1x reads, and how to change tools or effort without a cache miss.

GPT-6 prompt caching lets the API reuse the work it already did on the unchanged start of your prompt, and bills those reused tokens at a tenth of the normal input rate. With the GPT-6 Sol and Luna release on 22 September 2026, OpenAI added explicit cache breakpoints, a fixed 30-minute cache lifetime, and ways to change reasoning effort or tools mid-conversation without throwing the cache away.
If you run agents or long chats against the API, those three changes decide most of your input bill. Here is how the mechanism works and what to change in your requests.
What is prompt caching in GPT-6?
When a model reads your prompt it computes intermediate key-value state for every token. Prompt caching saves that state for a prefix, meaning the unchanged tokens at the start of a request. If the next request begins with exactly the same tokens, the model loads the saved state instead of processing them again.
Two consequences follow. Only the beginning of the prompt can be reused, so anything that changes early in the prompt invalidates everything after it. And the match is exact: one different character in the system message is a different prefix.
For GPT-5.6 and later models, a prefix must be at least 1,024 visible input tokens before it can be cached. Short prompts never hit the cache, which is fine, because short prompts are cheap anyway.
How much does caching save?
On GPT-5.6 and later, caching is no longer free to write. Writing a prefix into the cache costs 1.25× the normal input rate. Reading it back costs 0.1× on most models, and 0.05× on GPT-6.1 Sol, which shipped on 29 September.
The documentation gives the break-even case: write a prefix once and reuse it once, and you pay 1.35× its normal cost instead of 2×. Every extra reuse widens the gap.
A worked example on GPT-6 Sol, which costs $2 per million input tokens. Take a 20,000-token prefix (system prompt, tool definitions, a reference document) sent ten times, ignoring the short part that changes each turn:
| Approach | Cost of the prefix over 10 requests |
|---|---|
| No caching | $0.400 |
| Cached, 0.1× reads | $0.086 |
| Cached on GPT-6.1 Sol, 0.05× reads | $0.068 |
The cached prefix costs about a fifth of the uncached one. The catch is the word "unchanged". A cache that keeps missing gives you the 1.25× write premium with none of the reads.
How do explicit cache breakpoints work?
A breakpoint marks where a cacheable prefix ends. By default, in implicit mode, OpenAI places one at the end of the latest eligible message for you. That is enough for a plain chat.
Explicit mode lets you choose the boundary yourself. You set the request-level mode, then mark the content block that ends the stable part:
{
"model": "gpt-6-sol",
"prompt_cache_options": { "mode": "explicit", "ttl": "30m" },
"input": [
{
"role": "developer",
"content": [
{
"type": "input_text",
"text": "Long, stable instructions and reference material...",
"prompt_cache_breakpoint": { "mode": "explicit" }
}
]
},
{ "role": "user", "content": "Today's question" }
]
}
A few rules to remember:
- A request can create at most four cache writes.
- The top-level
instructionsfield cannot hold a breakpoint, so put stable instructions in a developer message if you want to mark them. - In explicit mode, a request with no breakpoints does not use the cache at all. Forgetting the marker is a silent full-price request.
Explicit mode is worth it when your prompt has layers that change at different speeds. A typical agent has fixed instructions, a per-user document that stays for a session, and a conversation that grows every turn. A breakpoint after each stable layer means a new user still reuses the shared instructions, even though their document differs.
How long does the cache last?
The new prompt_cache_options.ttl field accepts one value, "30m", which is also the default. A cached prefix stays eligible for 30 minutes after it was last written or read, so it stays available as long as some request reuses it at least every 30 minutes.
Older models used a different field, prompt_cache_retention, with in_memory and 24h options. If you share request-building code across model generations, branch on the model rather than sending both.
There is also a prewarm option. Setting prompt_cache_options.prewarm to true prepares the cache for a prompt without generating any output. That suits a known burst, such as a batch job that is about to send thousands of requests with the same instructions.
{
"model": "gpt-6.1-sol",
"input": [{ "role": "developer", "content": "Shared instructions..." }],
"prompt_cache_options": { "prewarm": true }
}
What breaks the cache, and how do you avoid it?
Some request settings are part of the cache identity. Changing any of these discards existing entries for later requests:
modeltools, including names, descriptions, schemas and their orderparallel_tool_callstext.format, the Structured Outputs schemareasoning.effortat the request leveltext.verbositycontext_management, the compaction setting
The two that agents change most often, effort and tools, now have cache-safe alternatives.
For reasoning effort, keep the request-level value fixed and append a configuration_update item to the input when you want the next response to think harder or less:
{
"type": "configuration_update",
"reasoning": { "effort": "high" }
}
The earlier prefix survives because nothing before that item changed.
For tools, treat the list as append-only. Keep every definition in tools stable from the first turn, use allowed_tools to restrict which ones the model may call right now, and set tool_choice to "none" when it should call none. If you need a tool that appears partway through a thread, a developer-role additional_tools input item adds it at that point in the input, after the cached part. Note that these items cannot carry a breakpoint themselves.
When you are not sure why hits are low, the new diagnostics report a reason for each miss, such as tools_changed, along with how many tokens could have been reused. The usage object also splits input into cached_tokens and cache_write_tokens under input_tokens_details, so you can log the hit rate per route.
Key takeaways
- GPT-6 prompt caching reuses the exact unchanged start of a prompt; anything that changes early invalidates what follows.
- Cache writes cost 1.25×, reads 0.1× (0.05× on GPT-6.1 Sol), and the prefix must be at least 1,024 tokens.
- Use explicit breakpoints after each stable layer of an agent prompt, and never send an explicit-mode request without one.
- Change effort with
configuration_updateand limit tools withallowed_toolsinstead of editing the request settings. - Watch
cached_tokensandcache_write_tokens: a high write count with few reads means you are paying the premium for nothing.
FAQ
Do I need prompt_cache_key on GPT-6?
No. On GPT-5.6 and later it is optional and not needed for good hit rates. It is still useful if you want cache usage accounted separately per customer or workspace.
Should I switch every request to explicit mode?
Not necessarily. Implicit mode already places a breakpoint at the latest eligible message, which suits simple chats. Explicit mode pays off when your prompt has several stable layers that change at different rates.
Why is cached_tokens zero on a short prompt?
On GPT-5.6 and later, a prefix must be at least 1,024 visible input tokens before it can be cached. Shorter prompts are always processed in full.