Caching Strategies for Web Applications
From redundant computation to layered caches: when web architectures need caching, how cache patterns balance freshness and load, and a minimal cache design playbook.
## Part 1: Why Do We Need Caching in Web Architecture?
Web applications spend most of their time recomputing answers they have already produced — the same product page rendered thousands of times, the same query hitting the database on every single request. Even a well-tuned backend will buckle when every page view translates into a fresh round of queries, template rendering, and serialization. Caching — the practice of storing computed results closer to where they are needed and serving them again without redoing the work — is where we bridge the gap between what the architecture can compute and what the traffic demands.
In this section we'll talk about how the no-cache baseline, vertical scaling, and layered caching differ — and why storing results at the right layer makes caching essential for building fast, resilient, and affordable web architectures.
### The No-Cache Baseline: Intuition
The simplest architecture computes everything on demand. A request arrives, the application queries the database, renders the response, and throws all of that work away:
> Request → App → Database → Render → Response (repeat for every request)
This works when traffic is low and queries are cheap. But it assumes:
- Every request genuinely needs freshly computed data.
- The database can absorb read traffic proportional to page views.
- Latency budgets tolerate a full round of computation on every hit.
For an internal dashboard with ten users, these assumptions often hold. For a public product page, they never do.
### Vertical Scaling: Guiding with Bigger Boxes
An obvious fix is to buy a bigger server. More cores, more RAM, faster disks — why not just make the origin fast enough?
> Same architecture + bigger hardware → same work, done faster?
This approach works for moderate growth, but it breaks down quickly:
- Cost scales superlinearly — the last 20% of headroom costs more than the first 80%.
- There is a ceiling: at some point no single machine can absorb the read load, no matter the price.
- It does nothing about distance — a user on another continent still pays the full network round trip to your origin.
Hardware is a budget, not a strategy. The question becomes: *which* work deserves to be done more than once?
### Caching: Compute Once, Serve Many
Caching answers that question with a simple two-step contract. Instead of recomputing every response, you:
1. **Look up**: given a request, check whether a stored result already exists for its cache key.
2. **Serve or compute**: on a hit, return the stored result immediately; on a miss, compute it once, store it, and let every subsequent request reuse it.
This approach:
- Cuts latency for repeated requests from origin time to lookup time.
- Shields the database — origin load drops to the miss rate, not the request rate.
- Absorbs traffic spikes, because a viral page is precisely the page most likely to be cached.
### The Common Feeling: "Isn't This Just Saving a Copy?"
Many engineers (myself included, at first) find caching underwhelming. If you squint, it looks like we just:
- Stored a computed result in a dictionary.
- Checked the dictionary before doing the work.
Isn't that just memoization with a network hop?
In practice, yes — a cache really is a key-value store in front of slower work. But conceptually, there are two crucial differences:
1. **What you optimize:**
- No cache: you optimize the cost of a single computation.
- Caching: you optimize away the *redundancy* across computations — a different axis entirely.
2. **What can go wrong:**
- No-cache failure is slow but always correct.
- Cache failure is fast but possibly stale — staleness becomes an explicit, tunable design parameter instead of an accident.
It's a subtle but important shift: from making each request fast to making repeated requests *free* — and accepting that freshness is now something you must design for.
### Why the Distinction Matters in Real Applications
This difference is especially sharp for read-heavy workloads like product catalogs, news feeds, or public APIs, where read-to-write ratios of 1000:1 are normal.
- **No-cache setup:** A product page needs 12 queries and 80 ms to render. At 5,000 requests per second, the database serves 60,000 queries per second for data that changes twice a day.
- **Cached setup:** The rendered page is stored for 60 seconds. The database sees a handful of misses per minute; everyone else gets a 2 ms lookup.
Formally, the average latency is a weighted blend over the hit ratio h:
> t_avg = h · t_cache + (1 − h) · t_origin
With t_cache = 2 ms, t_origin = 80 ms, and h = 0.95, average latency drops to 5.9 ms — and origin load drops by 20×, because only (1 − h) of requests ever reach it.
This is why caching is the highest-leverage performance tool in web architecture — it attacks the multiplier, not the unit cost.
### Advanced Variants: Beyond a Single Cache
When an application grows, a single cache in front of the database stops being enough. Real architectures layer caches along the request path:
- **Browser cache**: Cache-Control and ETag headers let the client skip the request entirely [3].
- **CDN / edge cache**: shared caches near the user absorb global read traffic and hide continental latency.
- **Application cache**: Redis or Memcached stores rendered fragments, query results, and sessions [1, 2].
- **Database buffer pool**: the database caches hot pages in memory on its own — free, but not under your control.
The guidance is structural: caching becomes a hierarchy, and each layer needs its own key design, TTL, and invalidation story.
### Intuition
- No cache: "Compute the answer every time."
- Vertical scaling: "Compute the answer every time, faster."
- Caching: "Compute the answer once — serve the memory of it until it expires."
### Conclusion
Caching in web architecture can feel, at first, like we're just saving copies of things. But the shift in what you optimize and the explicit freshness contract are what make it different from simply buying faster hardware. For low-traffic internal tools, the no-cache baseline is sufficient. For read-heavy public workloads, caching unlocks orders-of-magnitude leverage. For global products, layered caches show the full potential: latency hidden at the edge, databases shielded at the core, and traffic spikes absorbed where they land.
That's why caching — in one form or another — remains central to shaping how web systems not only respond quickly, but also survive their own success.
## Part 2: How to Design an Effective Caching Layer
This is a recap of my working notes from operating caches in front of several production web services, cross-checked against Facebook's Memcache paper [2], which remains the best field guide to caching at scale. If you find this interesting, I highly recommend going through the full paper.
### Caching Setup in the Web Context
- Origin (O): the source of truth — database queries, service calls, rendering
- Cache (C): the fast store holding computed results
- Key (k): the identity of a cached result, derived from the request
- Value (v): the stored result — a rendered page, fragment, or query result
- TTL: how long a value may be served before it is considered stale
- Hit ratio (h): the fraction of requests served from the cache
Unlike scaling the origin, caching modifies the request path instead of the computation — capacity updates are configuration updates, which makes the system flexible and cheap to tune.
### Naive Caching: Cache-Aside with a Fixed TTL
The simplest pattern is cache-aside: check the cache, fall back to the origin on a miss, and store the result with a fixed TTL. The core objective is:
> min t_avg = h · t_cache + (1 − h) · t_origin, subject to staleness ≤ TTL
Interpretation: minimize average latency by maximizing the hit ratio, while never serving data older than the TTL allows.
- If the data is read-heavy and tolerates minutes of staleness (e.g., a product description), this works.
- If the data must be fresh on write (e.g., a shopping cart), or if a hot key expires under load, this often fails.
This is like photocopying a notice and pinning it everywhere — cheap and effective, until the original changes. The problems:
- **Cache stampede**: when a popular key expires, hundreds of concurrent requests miss simultaneously and pile onto the origin [8].
- **Stale windows**: a write to the origin leaves the cached copy wrong until the TTL runs out — silently.
- **Key sprawl**: ad-hoc key naming makes it impossible to know what to invalidate when data changes.
As I noted in my build logs: almost every cache incident is an invalidation incident in disguise. The old joke holds — there are only two hard things in computer science: cache invalidation and naming things. A cache key is both at once.
### Advanced Patterns: Improving Freshness and Resilience
To improve reliability, we upgrade the read and write paths independently. Key patterns:
- **Explicit invalidation**: on write, delete (don't update) the affected keys — deletion is idempotent and safe under races [2].
- **Write-through**: writes go through the cache to the origin, keeping the two in step at the cost of write latency.
- **Stale-while-revalidate**: serve the stale copy instantly and refresh it in the background, so users never wait on a refresh [4].
- **Request coalescing**: on a miss, let one request compute while the rest wait for its result — the standard stampede guard.
- **Negative caching**: cache "not found" results briefly, so a flood of requests for a missing item doesn't hammer the origin.
Intuition: don't just store results; decide who refreshes them, when, and what everyone else does while the refresh is in flight.
### The Cache Design Algorithm
Algorithm: Cache Layer Workflow
1. Profile the workload: identify the hottest read paths and their read-to-write ratios
2. Choose the layer: browser/CDN for whole responses, application cache for fragments and query results
3. Design the key: derive it from every input that changes the output — nothing more, nothing less
4. Choose freshness: pick a TTL from the data's tolerance for staleness, and add explicit invalidation for writes that can't wait
5. Add stampede protection: request coalescing or stale-while-revalidate for every hot key
6. Instrument: track hit ratio, origin load, and staleness age per key family
7. Evaluate and refine: a falling hit ratio means wrong keys or wrong TTLs — fix the design, not the size
#### Key insights:
- Step 3: The key is the contract — if two different outputs can share a key, you have a correctness bug, not a performance bug.
- Step 4: TTL is your safety net, not your invalidation strategy — explicit deletes on write, TTL as the backstop for whatever you missed.
- Step 6: A cache without a hit-ratio dashboard is a rumor, not a component.
### The Role of TTL and Key Granularity
In cache design, freshness and granularity each serve distinct roles in balancing correctness and leverage:
**TTL (Freshness)**
- Defined by the data's tolerance for staleness: seconds for prices, minutes for listings, hours for static content.
- Used to bound the worst-case wrongness of what users see, independent of invalidation bugs.
- Intuition: "Decide how wrong you can afford to be, and write that number down."
**Key Granularity**
- Typically fragment-level for most applications; whole-page keys maximize reuse but explode on personalization.
- Why? Because one personalized element in a full-page key drops the hit ratio to nearly zero — cache the shared fragments, render the personal parts live.
- Intuition: "Cache what everyone shares; compute what makes each user different."
**Why Both Matter**
- The right TTL bounds staleness; the right granularity determines whether the cache gets hit at all.
- A perfect TTL on a badly scoped key still yields a 5% hit ratio — leverage comes from reuse.
- Too coarse reintroduces staleness across unrelated data; too fine multiplies keys until invalidation becomes unmanageable.
### Analogy to Denormalization
Denormalization also stores precomputed results, but with key differences:
- **Denormalization**: bakes derived data into the schema — permanent, transactional, and migrated with pain.
- **Caching**: keeps derived data outside the schema — temporary, disposable, and reshaped by changing a config.
In practice they are complements, not competitors: denormalize for derived data that must be transactionally correct, cache for derived data that must merely be fast.
### TL;DR:
- TTL: Bounds how stale a served result can ever be → set it from business tolerance, not from folklore.
- Key Granularity: Determines the hit ratio and the invalidation blast radius → cache shared fragments, not personalized pages.
### A Toy Example: Product Page Cache
Let's walk through a simple toy environment: the origin renders product pages in 80 ms, and the task is to serve them fast without showing customers a price that changed an hour ago.
#### Pattern Design
```python
def cache_aside(cache, origin, key, ttl=60):
hit = cache.get(key)
if hit is not None:
return hit # ~2 ms
value = origin.render(key) # ~80 ms: queries + templates
cache.set(key, value, ttl)
return value
```
#### Key Pattern Functions
Serve stale content while refreshing in the background (simplified):
```python
def swr_get(cache, origin, key, ttl=60, grace=300):
entry = cache.get_with_age(key)
if entry is None:
value = origin.render(key) # cold miss: user waits once
cache.set(key, value, ttl + grace)
return value
value, age = entry
if age > ttl:
schedule_refresh(cache, origin, key, ttl, grace)
return value # possibly stale, always fast
```
Evaluate cache effectiveness (simplified):
```python
def evaluate_cache(hits, misses, t_cache=0.002, t_origin=0.080):
h = hits / max(hits + misses, 1)
t_avg = h * t_cache + (1 - h) * t_origin
origin_load = 1 - h # fraction reaching the database
return {"hit_ratio": h, "avg_latency": t_avg, "origin_load": origin_load}
```
### Conclusions
- Hit ratio vs performance: A modest origin behind a 95% hit ratio beats a heroic origin serving every request cold.
- Freshness contracts: Explicit TTLs and delete-on-write invalidation turn staleness from an accident into a documented guarantee.
- Flexibility: Reshaping keys and TTLs is faster and cheaper than any database migration, making caches ideal for iterating on performance.
- Scaling: The toy code works for one process, but production caching involves consistent hashing across nodes [5], stampede protection under thundering herds [8], and hit-ratio monitoring per key family.
Vertical scaling can only make each computation faster, at ever-increasing cost. Caching removes the computation entirely for everyone after the first requester. While toy demos like "product page cache" are simple, the mechanics mirror how layered caches keep real web architectures fast under traffic they could never serve cold.
## References
1. Fitzpatrick, B. "Distributed Caching with Memcached." Linux Journal (2004).
2. Nishtala, R., et al. "Scaling Memcache at Facebook." NSDI (2013).
3. Fielding, R., Nottingham, M., and Reschke, J. "HTTP Caching." RFC 9111 (2022).
4. Nottingham, M. "HTTP Cache-Control Extensions for Stale Content." RFC 5861 (2010).
5. Karger, D., et al. "Consistent Hashing and Random Trees: Distributed Caching Protocols for Relieving Hot Spots on the World Wide Web." STOC (1997).
6. DeCandia, G., et al. "Dynamo: Amazon's Highly Available Key-value Store." SOSP (2007).
7. Redis Documentation. "Caching Patterns and Best Practices." (2024).
8. Vattani, A., Chierichetti, F., and Lowenstein, K. "Optimal Probabilistic Cache Stampede Prevention." VLDB (2015).