Tags: web-dev concept

Environments

Date: 2026-08-17


The ladder a change climbs before reaching customers. Every rung costs money and slows delivery, and earns its place only by catching a class of bug the rung below can’t — which is a much higher bar than most staging environments clear.


An environment is a complete, separately deployed instance of a system — its code, config and data — used for one stage between a developer’s machine and customers.

The ladder

LOCAL        one developer, fast feedback, fake everything
                ↓
PREVIEW      one per pull request, real deploy, real URL      ← Preview Environments
                ↓
STAGING      production-shaped, shared, integration-tested
                ↓
PRODUCTION   customers

Each rung must catch something the one below cannot, or it’s queue time with extra steps:

RungCatchesCan’t catch
LocalLogic, most bugsAnything about the deployed shape
Preview”Does it look and work right”, stakeholder reviewScale, real data volume, third-party sandboxes lying
StagingIntegration with real services, migrations against real schemaProduction traffic, production data, production scale
ProductionEverything else — which is why you need flags and observability here—

Drift is the default state

Environments diverge from the moment they’re created, and the divergence is what makes staging pass while production fails.

                    staging              production
data volume         12,000 rows          14 million rows      ← the query is fine here, not there
third parties       sandbox              live
traffic             one tester           4,000 concurrent
config              copied 8 months ago  current
scale               1 instance           12 behind a load balancer
TLS/CDN             often absent         always present

The data volume row causes more escaped bugs than the rest combined. A missing index is invisible at 12,000 rows and an outage at 14 million — Indexing, N+1 Queries.

Reducing drift, in order of value:

  • Same deployment mechanism everywhere. If production is deployed differently from staging, staging tests the application but not the deploy — and the deploy is a common failure
  • Infrastructure as code, so environments are rebuilt from a definition rather than accumulated by hand. The test is whether you could delete staging and recreate it by Friday
  • Configuration by injection, not by copy — Environment Configuration
  • Production-scale data in at least one environment, anonymised — see below
  • Rebuild staging periodically. A long-lived environment accumulates manual fixes that exist nowhere else

Data in non-production

The tension: staging is useless without realistic data, and a copy of production is a copy of everyone’s personal data sitting in a less-defended environment with wider access.

  • Never a raw production dump. Under UK GDPR — the UK’s retained General Data Protection Regulation — a test environment holding personal data is processing it, needs a lawful basis, and is in scope for a breach notification. It’s also usually the environment where credentials are shared and logging is verbose — PII in Analytics, UK GDPR and PECR for Analytics
  • Anonymise on extraction, not after loading. The window between the two is the leak
  • Anonymisation must preserve shape. Replacing every postcode with AA1 1AA breaks the address validation you were trying to test; replacing every order total with 10.00 makes the totals meaningless. Consistent pseudonymous substitution, preserving distributions and cardinality — Cardinality
  • Synthetic data at production volume is the better answer for performance testing and the worse one for “does this actually work with real-world mess”
  • Never point a non-production environment at a live third party. Sandbox credentials, enforced by configuration, checked at startup — the failure mode is a test run emailing real customers

How many rungs

The honest default for a small team is three: local, preview, production. Preview Environments do most of what staging was for, per change rather than in a shared queue, and a shared staging environment has costs that get understated:

  • It’s a queue. One broken change blocks everyone else’s verification
  • Nobody owns it, so it’s broken often and its failures get ignored — which trains people to ignore failures
  • A green staging run stops meaning anything once “staging is just flaky” is a sentence people say

Staging earns its place when there’s something genuinely untestable elsewhere: a third-party integration with only one sandbox, a regulated sign-off, a payment provider that won’t issue per-branch credentials, or a data migration to rehearse.

The real backstop is production itself, made safe: Feature Flags to keep changes dark, Progressive Delivery to expose them slowly, and Observability to see the result. That combination catches more than any staging environment, because it’s testing against the only traffic that’s real.

Where it interacts

  • Blue-Green and Rolling Deployments — blue and green are both production; don’t confuse a deployment slot with an environment rung
  • Database Migrations — rehearsing against production-scale data is the main thing a staging environment is genuinely good for
  • Secrets Management — per-environment credentials, never shared, and a production secret reachable from staging means you have one environment
  • Multi-Tenancy — with per-tenant databases, “an environment” is a fleet, and rebuilding it is a different scale of job