Blue-Green and Rolling Deployments
Date: 2026-08-17
Two ways to replace a running version without downtime. Blue-green swaps everything at once between two full environments; rolling replaces instances a few at a time. The choice is between paying for double capacity and accepting that both versions run simultaneously.
Blue-green deployment runs two complete production environments and switches all traffic from old to new at once; rolling deployment replaces the instances of one environment in batches.
The two mechanics
BLUE-GREEN ROLLING
blue v1 ██████ ← all traffic ██ ██ ██ ██ ██ ██ 6 instances on v1
green v2 ██████ ← idle, warmed
░░ ██ ██ ██ ██ ██ replace 1, health check
switch router ░░ ░░ ██ ██ ██ ██ replace next
…
blue v1 ██████ ← idle, kept ░░ ░░ ░░ ░░ ░░ ░░ done
green v2 ██████ ← all traffic
cutover: seconds cutover: minutes
rollback: flip back, seconds rollback: roll forward again, slow
cost: 2× capacity during deploy cost: ~1× plus one spare
mixed versions live: no* mixed versions live: YES, always
* Blue-green avoids mixed versions serving traffic, but in-flight requests, background workers and anything already queued still straddle the switch. The database is shared throughout in both models, so the compatibility constraint below applies to both regardless.
The constraint neither escapes
The database is not blue and green. There is one, both versions use it, so every schema change must satisfy the old code and the new code simultaneously — expand-contract, always — Database Migrations, Backwards Compatibility.
This is why “we use blue-green so we can roll back instantly” is only half true. The code rolls back in seconds; the migration doesn’t roll back at all, and if v2 wrote data in a shape v1 can’t read, flipping the router restores a broken site rather than a working one.
Choosing
| Blue-green | Rolling | |
|---|---|---|
| Cost | Double capacity for the deploy window | One extra instance |
| Rollback | Instant — flip the router | Another rolling deploy |
| Version mixing | Brief and bounded | Guaranteed, for the whole deploy |
| Long-lived connections | Awkward — WebSockets and streams must drain | Same problem, spread out |
| Verification before traffic | Yes — smoke-test green in place | No — instances take traffic as they come up |
| Suits | Monoliths, big-bang releases, regulated cutovers | Many small instances, frequent deploys, the default in Kubernetes |
Rolling is the sane default for a service deployed several times a day. Blue-green earns its cost when the release is large, infrequent, or needs verification against production infrastructure before anyone sees it.
The pieces that make either work
- Health checks that mean something. An instance reporting healthy because the process started, before it can serve a request, produces a deploy that “succeeded” into an outage. Check dependencies, not liveness
- Connection draining. Remove an instance from the load balancer, let in-flight requests finish, then stop it. Without this, every deploy drops requests
- Warm-up. JIT compilation, connection pools and caches mean a cold instance is slow. Rolling straight through six instances can brown out a site that never technically went down
- Session state must not live in an instance. If it does, replacing instances logs people out — Sessions and Tokens
- Background workers deploy too, and they’re routinely forgotten. A worker on v1 consuming messages produced by v2 is the same compatibility problem with less visibility
Where this sits relative to progressive delivery
They answer different questions, and conflating them is common:
- Blue-green and rolling are about replacing infrastructure — how instances get swapped, with no reference to who sees what
- Progressive Delivery is about exposing behaviour — 1% of users, then 5%, watching guardrail metrics between steps
- Feature Flags are how you get the second without the first. Deploy to 100% of instances with the feature off, then ramp exposure independently
A canary is the overlap: a small number of instances on the new version taking real traffic while metrics are compared. It’s a rolling deploy that pauses after the first step, and it’s the cheapest version of progressive delivery available if you don’t have flags.
Where it interacts
- Rollback and Forward Fix — blue-green makes rollback genuinely cheap, which changes the decision, but only for stateless changes
- Environments — blue and green are not staging and production; they’re two halves of production
- Observability — during any deploy, metrics must be split by version or you’re averaging the old and new behaviour and seeing neither