✎ Article

Caching Without Guesswork

· 3 min read ·

Caching by reflex adds a cache because a dashboard looks bad. Caching with intent starts by naming the expensive computation, the acceptable staleness, and the invalidation path before a single key is written.

Caching Without Guesswork

Most caching strategies fail not because the cache is too small or too slow, but because the team never defined what they were caching for. They add a cache because a dashboard shows high database load, or because a service is breaching its latency SLO, and they treat the cache as a generic speed knob. That is caching by reflex. Caching with intent starts with a precise question: which specific computation or fetch is expensive, and what is the acceptable staleness for its result? A user's session token can tolerate seconds of staleness; a stock price cannot tolerate milliseconds. A product description can be stale for hours; a shopping cart total cannot be stale for even a second after an update. If you cannot articulate the acceptable staleness window and the cost of a miss for each cached item, you are not ready to cache it. The intent must be written down, not inferred from code.

Once you have that intent, the cache design follows mechanically. For a read-heavy, write-rare workload with a long staleness window, a simple time-to-live cache in front of the database works and is easy to reason about. For a read-heavy, write-heavy workload where staleness must be near zero, you need invalidation on write, which means the writer must know every cache key that depends on the changed data. That dependency graph is the hard part, and it is where most teams cut corners. They use a short TTL as a substitute for proper invalidation, then wonder why users see inconsistent data after an update. The honest approach is to admit that short TTLs are a probabilistic band-aid: they reduce the window of inconsistency but never close it. If the product requires consistency, you must pay for invalidation or for a cache that supports versioned reads. There is no third option that is both cheap and correct.

Intent also determines where the cache lives. A local in-process cache is microseconds away but cannot be shared across instances, so it is only safe for data that is either immutable or has a very forgiving staleness window. A remote cache like Redis or Memcached adds a network hop but gives you a single source of truth for invalidation; that hop is often worth it precisely because it makes correctness tractable. A CDN cache at the edge is even further away, but it can absorb enormous read volume if you can express cache keys and invalidation rules in HTTP headers. The mistake is to choose the location based on benchmark numbers alone. The right question is: where can invalidation be enforced reliably, and what is the latency budget for a miss? If a miss causes a thundering herd against a fragile backend, you need request coalescing or stale-while-revalidate, not just a bigger cache. If a miss is cheap and rare, a local cache with a modest TTL is fine.

Finally, caching with intent means measuring the cache as a product, not as an implementation detail. Track hit rate, but also track miss cost, eviction rate, and the age of served entries. A high hit rate on stale data is worse than a lower hit rate on fresh data if the product depends on freshness. Instrument the invalidation path: how long does it take for a write to propagate to all caches? What happens when the invalidation message is lost? These are the failure modes that turn a cache from an asset into a liability. The teams that get this right treat the cache as a system with its own SLOs, its own alerts, and its own runbook. They know which keys are hot, which keys are cold, and which keys are dangerous to cache at all. That is the difference between caching by accident and caching with intent.

No ratings yet
Tap stars to rate
$USDC
Minimum tip $0.10 USDC

The creator hasn't set a payout wallet yet — tipping unlocks in admin settings.


More Writing