Rollback and Forward Fix
Date: 2026-08-16
Two ways out of a bad release, and the choice has to be made before you need it. Rollback is fast and often impossible; forward fix is always possible and slower — and mid-incident is the worst moment to work out which applies.
What it is
- Rollback — return to the previous known-good state
- Forward fix — ship a correction
Rollback is the default instinct and it’s frequently unavailable, because the previous version can no longer run against the current world.
What blocks a rollback
deploy v2 → migration adds a column, code writes to it
→ 10,000 orders written with the new field
→ incident
→ roll back to v1?
v1 doesn't know the column exists
v1 can't read those 10,000 orders correctly
the migration can't be reversed without losing data
The general rule: anything that changed state which the old version can’t interpret blocks a rollback. Common cases:
- Schema changes the old code doesn’t understand — Database Migrations
- Data written in a new format
- Messages published in a new shape to a queue another service consumed
- External side effects — emails sent, payments taken, webhooks delivered
- Cached artefacts in a new format
This is exactly what expand-contract prevents. If every change is backwards-compatible when it lands, rollback stays available — which is the real argument for the discipline, more than elegance.
Choosing
| Rollback when | Forward fix when |
|---|---|
| The change is backwards-compatible | A migration or data change can’t be reversed |
| The blast radius is wide and growing | The problem is narrow and understood |
| You don’t yet know the cause | The fix is small and obvious |
| A flag can disable it instantly | Rolling back would break something newer |
The fastest rollback is a feature flag. Turning a flag off is seconds and touches no infrastructure. That’s the strongest practical argument for shipping behind flags: it converts “roll back the release” into “switch off the feature”, which is faster, narrower, and doesn’t revert everything else that shipped alongside.
The decision belongs upfront
Mid-incident, under pressure, with an audience, is the worst time to discover the rollback path doesn’t exist.
Decide at design time, and write it down:
- Is this change backwards-compatible? If not, why, and what’s the recovery path
- Is it behind a flag? If not, why not
- What’s the rollback command, precisely, and has anyone run it
- What’s the threshold at which we stop rather than diagnose
That last one matters most. In an incident the temptation is always to understand before acting. Restore service first, diagnose afterwards — a rolled-back system with an unexplained bug is a far better position than a broken one with a promising theory.
Practicalities
- Test the rollback. An untested rollback path is a hypothesis. Exercise it in staging as part of release readiness
- Keep the previous artefact deployable. Retention policies that delete last week’s build remove your fastest option
- Roll back the whole release, not a cherry-picked commit. Partial reverts create states nobody has tested
- Watch after rolling back. Confirm recovery rather than assuming it, and check for data written during the bad window that now needs repair
- Record the incident and the decision. Whether you rolled back or fixed forward, and why, is the most useful thing in the postmortem — Incident Response
Where it connects
- Progressive Delivery — a ramp limits how much damage happens before you have the choice at all
- Feature Flags — the mechanism that makes rollback instant and narrow
- Database Migrations — the usual reason rollback is unavailable
- Continuous Deployment — deploying frequently makes each rollback small, which is what makes deploying frequently safe rather than reckless