Tags: web-dev concept

Rollback and Forward Fix

Date: 2026-08-16


Two ways out of a bad release, and the choice has to be made before you need it. Rollback is fast and often impossible; forward fix is always possible and slower — and mid-incident is the worst moment to work out which applies.


What it is

  • Rollback — return to the previous known-good state
  • Forward fix — ship a correction

Rollback is the default instinct and it’s frequently unavailable, because the previous version can no longer run against the current world.

What blocks a rollback

deploy v2  →  migration adds a column, code writes to it
           →  10,000 orders written with the new field
           →  incident
           →  roll back to v1?
                v1 doesn't know the column exists
                v1 can't read those 10,000 orders correctly
                the migration can't be reversed without losing data

The general rule: anything that changed state which the old version can’t interpret blocks a rollback. Common cases:

  • Schema changes the old code doesn’t understand — Database Migrations
  • Data written in a new format
  • Messages published in a new shape to a queue another service consumed
  • External side effects — emails sent, payments taken, webhooks delivered
  • Cached artefacts in a new format

This is exactly what expand-contract prevents. If every change is backwards-compatible when it lands, rollback stays available — which is the real argument for the discipline, more than elegance.

Choosing

Rollback whenForward fix when
The change is backwards-compatibleA migration or data change can’t be reversed
The blast radius is wide and growingThe problem is narrow and understood
You don’t yet know the causeThe fix is small and obvious
A flag can disable it instantlyRolling back would break something newer

The fastest rollback is a feature flag. Turning a flag off is seconds and touches no infrastructure. That’s the strongest practical argument for shipping behind flags: it converts “roll back the release” into “switch off the feature”, which is faster, narrower, and doesn’t revert everything else that shipped alongside.

The decision belongs upfront

Mid-incident, under pressure, with an audience, is the worst time to discover the rollback path doesn’t exist.

Decide at design time, and write it down:

  • Is this change backwards-compatible? If not, why, and what’s the recovery path
  • Is it behind a flag? If not, why not
  • What’s the rollback command, precisely, and has anyone run it
  • What’s the threshold at which we stop rather than diagnose

That last one matters most. In an incident the temptation is always to understand before acting. Restore service first, diagnose afterwards — a rolled-back system with an unexplained bug is a far better position than a broken one with a promising theory.

Practicalities

  • Test the rollback. An untested rollback path is a hypothesis. Exercise it in staging as part of release readiness
  • Keep the previous artefact deployable. Retention policies that delete last week’s build remove your fastest option
  • Roll back the whole release, not a cherry-picked commit. Partial reverts create states nobody has tested
  • Watch after rolling back. Confirm recovery rather than assuming it, and check for data written during the bad window that now needs repair
  • Record the incident and the decision. Whether you rolled back or fixed forward, and why, is the most useful thing in the postmortem — Incident Response

Where it connects

  • Progressive Delivery — a ramp limits how much damage happens before you have the choice at all
  • Feature Flags — the mechanism that makes rollback instant and narrow
  • Database Migrations — the usual reason rollback is unavailable
  • Continuous Deployment — deploying frequently makes each rollback small, which is what makes deploying frequently safe rather than reckless