Tags: web-dev concept

Build Caching

Date: 2026-08-17


Skipping work whose inputs haven’t changed. It’s the difference between a two-minute CI run and a twenty-minute one — and it only works if builds are deterministic, which most aren’t until someone makes them so.


Build caching stores the output of a build step keyed by a hash of its inputs, so an identical set of inputs returns the stored result instead of re-running.

INPUTS          source files, config,
                dependency versions, env
   │ hash
CACHE KEY       a3f9c2e1…
   │
HIT             return stored output
MISS            run the step, store it

Determinism is the prerequisite. If the same inputs can produce different outputs, the cache returns something wrong — Idempotency.

The layers

LayerScopeTypical saving
Dependency installnode_modulesLarge — most of a cold CI run
CompilationPer file or moduleLarge on big codebases
Task-levelA whole build or test stepLargest, hardest to get right
Docker layerImage build stepsLarge — Containers

Dependency caching is the cheapest win in CI and the one most often missing:

- uses: actions/setup-node@v4
  with:
    node-version: 24
    cache: 'npm'

Keyed on the lockfile hash, so it invalidates exactly when dependencies change and not otherwise — Lockfiles.

Local versus remote

LOCAL     one machine, one developer
          → your second build is fast
          → CI gets nothing

REMOTE    shared across machines and CI
          → a colleague's build, or CI's,
            populates yours
          → a PR that changed nothing in
            a package skips its tests
            entirely

Remote caching is where the large wins are in a monorepo — most pull requests touch one package, and everything else can be restored rather than rebuilt — Monorepos.

The trust question is real, since a poisoned remote cache serves malicious output to everyone. Cache writes should come from trusted CI only, with read-only access for developers.

What breaks determinism

The things that silently make a build unreproducible:

  • Timestamps embedded in output
  • Absolute paths baked into source maps or bundles
  • Non-deterministic module IDs — hash-based ordering that shifts between runs
  • Math.random() or Date.now() at build time
  • Unpinned dependency versions, so a fresh install resolves differently — Lockfiles
  • Environment variables varying between machines and not being part of the key
  • Filesystem ordering, where a glob returns files in a different order
# the test
npm run build && mv dist a
npm run build && mv dist b
diff -r a b

If that diff isn’t empty, your cache is unreliable — and so is your content hashing, which means every deploy invalidates every cached asset in browsers — Module Bundling.

Getting the key right

Too broad — the whole repository — means the cache misses on every commit and buys nothing.

Too narrow — only source files — means missing an input that matters, and returning a stale result. The classic omissions:

FORGOTTEN INPUTS
  the tool's own version
  the config file
  environment variables that affect
    output
  the Node version
  the OS / architecture

A cache returning a wrong result is worse than no cache, because it produces a build nobody can reproduce and a bug nobody can locate.

Docker specifically

Layer order determines how much is reused:

# BAD — any source change reinstalls
COPY . .
RUN npm ci
 
# GOOD — deps cached until the
# lockfile changes
COPY package*.json ./
RUN npm ci
COPY . .

Copy the dependency manifest, install, then copy the source. One of the highest-value four-line changes available in any Dockerfile — Containers.

Measuring it

CACHE HIT RATE       by step
COLD vs WARM TIME    what the cache buys
CACHE SIZE           and eviction rate

A hit rate near zero means the key is wrong, not that caching doesn’t help. Investigate the key rather than removing the cache — the usual cause is an input in the key that changes every run, like a timestamp or a commit SHA that didn’t need to be there.