Environments
Date: 2026-08-17
The ladder a change climbs before reaching customers. Every rung costs money and slows delivery, and earns its place only by catching a class of bug the rung below can’t — which is a much higher bar than most staging environments clear.
An environment is a complete, separately deployed instance of a system — its code, config and data — used for one stage between a developer’s machine and customers.
The ladder
LOCAL one developer, fast feedback, fake everything
↓
PREVIEW one per pull request, real deploy, real URL ← Preview Environments
↓
STAGING production-shaped, shared, integration-tested
↓
PRODUCTION customers
Each rung must catch something the one below cannot, or it’s queue time with extra steps:
| Rung | Catches | Can’t catch |
|---|---|---|
| Local | Logic, most bugs | Anything about the deployed shape |
| Preview | ”Does it look and work right”, stakeholder review | Scale, real data volume, third-party sandboxes lying |
| Staging | Integration with real services, migrations against real schema | Production traffic, production data, production scale |
| Production | Everything else — which is why you need flags and observability here | — |
Drift is the default state
Environments diverge from the moment they’re created, and the divergence is what makes staging pass while production fails.
staging production
data volume 12,000 rows 14 million rows ← the query is fine here, not there
third parties sandbox live
traffic one tester 4,000 concurrent
config copied 8 months ago current
scale 1 instance 12 behind a load balancer
TLS/CDN often absent always present
The data volume row causes more escaped bugs than the rest combined. A missing index is invisible at 12,000 rows and an outage at 14 million — Indexing, N+1 Queries.
Reducing drift, in order of value:
- Same deployment mechanism everywhere. If production is deployed differently from staging, staging tests the application but not the deploy — and the deploy is a common failure
- Infrastructure as code, so environments are rebuilt from a definition rather than accumulated by hand. The test is whether you could delete staging and recreate it by Friday
- Configuration by injection, not by copy — Environment Configuration
- Production-scale data in at least one environment, anonymised — see below
- Rebuild staging periodically. A long-lived environment accumulates manual fixes that exist nowhere else
Data in non-production
The tension: staging is useless without realistic data, and a copy of production is a copy of everyone’s personal data sitting in a less-defended environment with wider access.
- Never a raw production dump. Under UK GDPR — the UK’s retained General Data Protection Regulation — a test environment holding personal data is processing it, needs a lawful basis, and is in scope for a breach notification. It’s also usually the environment where credentials are shared and logging is verbose — PII in Analytics, UK GDPR and PECR for Analytics
- Anonymise on extraction, not after loading. The window between the two is the leak
- Anonymisation must preserve shape. Replacing every postcode with
AA1 1AAbreaks the address validation you were trying to test; replacing every order total with10.00makes the totals meaningless. Consistent pseudonymous substitution, preserving distributions and cardinality — Cardinality - Synthetic data at production volume is the better answer for performance testing and the worse one for “does this actually work with real-world mess”
- Never point a non-production environment at a live third party. Sandbox credentials, enforced by configuration, checked at startup — the failure mode is a test run emailing real customers
How many rungs
The honest default for a small team is three: local, preview, production. Preview Environments do most of what staging was for, per change rather than in a shared queue, and a shared staging environment has costs that get understated:
- It’s a queue. One broken change blocks everyone else’s verification
- Nobody owns it, so it’s broken often and its failures get ignored — which trains people to ignore failures
- A green staging run stops meaning anything once “staging is just flaky” is a sentence people say
Staging earns its place when there’s something genuinely untestable elsewhere: a third-party integration with only one sandbox, a regulated sign-off, a payment provider that won’t issue per-branch credentials, or a data migration to rehearse.
The real backstop is production itself, made safe: Feature Flags to keep changes dark, Progressive Delivery to expose them slowly, and Observability to see the result. That combination catches more than any staging environment, because it’s testing against the only traffic that’s real.
Where it interacts
- Blue-Green and Rolling Deployments — blue and green are both production; don’t confuse a deployment slot with an environment rung
- Database Migrations — rehearsing against production-scale data is the main thing a staging environment is genuinely good for
- Secrets Management — per-environment credentials, never shared, and a production secret reachable from staging means you have one environment
- Multi-Tenancy — with per-tenant databases, “an environment” is a fleet, and rebuilding it is a different scale of job