· 7 min read
Cache stampede: how to prevent the thundering herd
A cache stampede hits your database when a hot key expires. How request coalescing, locks, early refresh and stale-while-revalidate stop it.

A cache stampede happens when a popular cache entry expires and every request that misses it goes to the database at the same moment. Instead of one query rebuilding the value, you get hundreds or thousands of identical queries, and the database falls over exactly when traffic is highest. The fixes are request coalescing, a short rebuild lock, refreshing before expiry, and serving stale data while one worker refreshes.
The problem is easy to miss because the cache works perfectly almost all of the time. It only fails during the few hundred milliseconds after a hot key expires, and those are usually the busiest few hundred milliseconds of the day.
What is a cache stampede?
A cache stampede (also called a thundering herd or dog-piling) is a burst of duplicate work caused by many clients missing the same cache key at once. A hot key is a single cache entry that a large share of traffic reads, like a homepage, a results page or a config blob.
The standard read-through pattern looks harmless:
async function get(key) {
const hit = cache.get(key);
if (hit && hit.expires > Date.now()) return hit.value;
const value = await loadFromDb(key); // slow
cache.set(key, { value, expires: Date.now() + 60_000 });
return value;
}
Picture 1,000 requests arriving in the 50 ms it takes loadFromDb to run. Every one of them checks the cache, finds nothing, and calls the database. Nothing in this code tells request 2 that request 1 is already doing the work. A quick simulation with a 50 ms fake query and 1,000 concurrent callers confirms it: 1,000 database calls for one value.
The cost compounds. The database slows down under the duplicate load, so each rebuild takes longer, so the window where the key is missing gets wider, so more requests pile into it. That feedback loop is why stampedes turn a small blip into an outage.
How does request coalescing stop a stampede?
Request coalescing (sometimes called single-flight) means only one caller rebuilds a missing key, and everyone else waits for that same result. Inside one process it takes a few lines: keep a map of in-flight promises.
const inflight = new Map();
async function get(key) {
const hit = cache.get(key);
if (hit && hit.expires > Date.now()) return hit.value;
if (inflight.has(key)) return inflight.get(key);
const p = loadFromDb(key)
.then(value => {
cache.set(key, { value, expires: Date.now() + 60_000 });
return value;
})
.finally(() => inflight.delete(key));
inflight.set(key, p);
return p;
}
Run the same 1,000 concurrent calls through this version and the database sees exactly one query. The finally matters: if the load throws, the entry is removed so the next request can retry instead of awaiting a rejected promise forever.
The limit is scope. This deduplicates within one process. With 40 app instances behind a load balancer, you can still get 40 rebuilds. That is usually fine, since 40 is not 1,000, but for expensive values you need coordination across instances.
Using a lock across many servers
The cross-instance version uses the shared cache itself as a lock. In Redis, SET with NX (only set if the key does not exist) and PX (expiry in milliseconds) is atomic, so exactly one caller wins.
SET lock:results <token> NX PX 5000
The winner rebuilds the value, writes it, and deletes the lock. Losers do not hit the database. They either wait briefly and re-read the cache, or return a stale copy if one is available.
A few details decide whether this works in practice:
- Always set an expiry on the lock. If the winning process crashes mid-rebuild, the lock must free itself, or nobody refreshes the key again.
- Make the expiry longer than a normal rebuild but short enough that a crash only blocks refreshes briefly.
- Store a random token and check it before deleting, so a slow worker whose lock already expired does not delete a lock someone else now holds. Do the check and delete atomically, for example in a small Lua script.
- Bound the waiting. Losers that poll forever just move the stampede from the database to the cache.
A lock like this is a coordination hint, not a correctness guarantee. If two workers occasionally rebuild the same key, nothing breaks. If your use case needs strict mutual exclusion, a cache lock is the wrong tool.
Refreshing before the key expires
Coalescing and locks reduce the damage after a miss. A better approach avoids the miss entirely: refresh hot keys before they expire.
The simplest version is TTL jitter, which adds a random offset to each expiry so that keys written together do not all expire together. If you warm 10,000 keys at deploy time with a flat 60-minute TTL, they will all expire in the same second an hour later. A random 0 to 10 percent extension spreads that out.
Jitter helps with many keys expiring at once, but not with one very hot key. For that, use probabilistic early expiration. Each reader, on a cache hit, rolls a die. The closer the entry is to expiry, the more likely the roll says "refresh now". The well-known formula, often called XFetch, from a 2015 research paper on optimal stampede prevention, is:
import math, random
def should_refresh(now, expiry, delta, beta=1.0):
# delta: how long the last rebuild took, in seconds
# beta: > 1 refreshes earlier, < 1 refreshes later
return now - delta * beta * math.log(random.random()) >= expiry
Because log(random()) is negative, the subtraction pushes now forward by a random amount scaled by how slow the rebuild is. The probability of refreshing with r seconds left works out to exp(-r / (delta * beta)). Simulating it with a 0.5 second rebuild shows the shape:
| Seconds before expiry | Chance a request refreshes |
|---|---|
| 10 | ~0% |
| 2 | ~2% |
| 1 | ~14% |
| 0.5 | ~37% |
| 0.1 | ~82% |
Under heavy traffic, some request almost certainly triggers a refresh in the last second or two, while the old value is still valid. Under light traffic, nothing refreshes early, which is what you want: you do not pay for rebuilds nobody reads. You need to store delta alongside the value, which just means timing the rebuild when you write it.
Serving stale while revalidating
The last tool changes what a miss means. Instead of deleting the value at expiry, keep two times: a soft TTL, after which the value is stale and should be refreshed, and a hard TTL, after which it must not be served at all.
Between the two, a request returns the stale value immediately and triggers one background refresh, guarded by coalescing or a lock. Users never wait on the database for a hot key, and the database only ever sees one rebuild.
HTTP caches already support this with the stale-while-revalidate directive on Cache-Control:
Cache-Control: max-age=60, stale-while-revalidate=300
That tells a supporting browser or CDN the response is fresh for 60 seconds, and for the next 300 seconds it may serve the stale copy while fetching a new one in the background. Support varies between CDNs, so check how your provider handles it, but the idea is the same one you would build in your application cache. It pairs naturally with the cache-first read path described in designing for the spike, not the average.
When should you use which technique?
| Technique | Prevents | Cost | Use when |
|---|---|---|---|
| TTL jitter | Many keys expiring together | One line | Always, it is nearly free |
| In-process coalescing | Duplicate rebuilds per instance | A promise map | Any read-through cache |
| Distributed lock | Duplicate rebuilds across instances | Lock logic, timeouts | Rebuilds are expensive |
| Probabilistic early refresh | The miss itself, for hot keys | Store rebuild time | A few keys take most traffic |
| Stale-while-revalidate | Users waiting on rebuilds | Two TTLs, background work | Slightly old data is acceptable |
They stack. A common setup is jitter on everything, coalescing in every process, and stale-while-revalidate for the handful of keys that matter. Add the lock only where a single rebuild is expensive enough that even one per instance hurts.
Key takeaways
- A cache stampede is many requests rebuilding the same expired key at once, and it hits hardest at peak traffic.
- Request coalescing turns N concurrent misses into one rebuild inside a process with a few lines of code.
- Use a short, expiring lock with a unique token to coordinate rebuilds across instances.
- Jitter TTLs, and refresh hot keys probabilistically before they expire.
- Serving stale data during a refresh means no user waits on the database for a hot key.
FAQ
Is a cache stampede the same as a thundering herd?
They are used interchangeably. Thundering herd is the broader term for many waiters waking up for one event; a cache stampede is the specific case where the event is a cache entry expiring.
Does a longer TTL fix cache stampedes?
No. A longer TTL makes stampedes less frequent, but each one is just as large when it happens. It also means users see older data for longer.
Do I need Redis to prevent cache stampedes?
No. In-process coalescing, jitter and probabilistic early refresh work with any cache, including an in-memory map. You only need a shared store when you want to coordinate rebuilds across multiple servers.