Refactoring Strategy
Date: 2026-08-17
Changing the structure of code without changing what it does, in steps small enough that the system stays shippable throughout. The strategy part is the sequencing — a refactor that can’t be released until it’s finished has become a rewrite, and inherits a rewrite’s failure rate.
The definition is strict
Refactoring changes structure, not behaviour. If the output changes, it isn’t a refactor — it’s a change, and it needs testing as one.
This matters because “refactoring” is routinely used to describe a mixed commit that reorganises code and fixes a bug and adjusts behaviour. When that breaks, nobody can tell which part did it.
✗ one commit: extract the pricing service, fix the rounding bug,
add VAT for Ireland
→ reverting loses the fix; keeping it keeps the bug
✓ three: 1. extract the pricing service (no behaviour change)
2. fix the rounding bug (behaviour change, tested)
3. add Irish VAT (feature, flagged)
→ each revertible on its own
The safety net comes first
Refactoring without tests is editing and hoping. But the code most in need of refactoring is usually the code hardest to test — that’s related, not coincidental.
The way through is characterisation tests: tests that assert what the code currently does, bugs included, with no judgement about whether it’s right.
1 find the seam the narrowest input → output boundary you can call
2 capture reality run real inputs, record actual outputs verbatim
3 assert those including the wrong ones. this is deliberate
4 refactor freely any change in output is now a failing test
5 fix the bugs after separately, with the test updated in that commit
Step 3 is the counterintuitive one. You’re locking in behaviour you know is wrong, because right now the goal is to change structure safely — and you cannot tell an accidental change from an intended one unless the baseline is exact.
Keeping it shippable
The strategy is choosing an order where every intermediate state is releasable.
Parallel change — also called expand-contract, the same shape as Database Migrations:
1 EXPAND add the new implementation alongside the old
nothing calls it yet. ship this.
2 MIGRATE move callers across, a few at a time
both exist. ship after each batch.
3 CONTRACT delete the old implementation
ship this.
Every one of those is independently deployable and independently revertible. Compare with the alternative — change the implementation and every caller in one branch — which is unreviewable, un-mergeable for a fortnight, and conflicts with everyone.
Supporting techniques worth naming:
- Branch by abstraction. Put an interface in front of the thing being replaced, then swap the implementation behind it. Lets a large replacement live on
mainin pieces - Feature Flags for risky swaps. Both implementations shipped, flag decides which runs, ramp it — Progressive Delivery
- Parallel run for critical logic. Run both, serve the old result, log where the new one disagrees. The only genuinely safe way to replace pricing, tax or anything financial — and the disagreement log is usually more informative than the tests
- Long-lived refactor branches are the failure mode. Two weeks of restructuring on a branch means two weeks of conflicts with everyone else’s work — Branching Strategies
Choosing what to refactor
Not everything, and not by taste. Two signals worth trusting:
- Change frequency × difficulty. Code touched weekly that takes a day to change is where the money is. Git history gives you the first factor for free — the files with the most commits are usually the ones worth attention
- It’s blocking the work in front of you. The most fundable refactoring is the kind that makes the next feature cheaper, and it should be part of that feature’s estimate rather than a separate request — Technical Debt
Explicitly not worth refactoring: stable code you dislike, and code you’re about to delete. Both are common and both are pure cost.
Where it goes wrong
- Refactoring while also adding features. The commit becomes unreviewable and the bug becomes unattributable
- Big-bang restructure. Beyond a certain size the correct comparison isn’t “refactor or not” but The Strangler Pattern
- Improving the abstraction rather than the problem. Three layers of indirection added to avoid a duplicated conditional is a worse system with better vocabulary
- No end state. A refactor that stops after the expand phase leaves two implementations, both live, permanently — this is a very common way to increase debt while believing you reduced it
- Renaming across a public boundary. Fine internally, a breaking change for anyone outside — Backwards Compatibility, Deprecation
Where it interacts
- Technical Debt — this is the repayment mechanism, and the case for it is made in that note’s terms
- Code Review — pure-refactor pull requests should be reviewed differently: the question is “is behaviour unchanged”, not “is this right”
- Test Doubles and Testing Strategy — tests coupled to implementation break during refactoring and make it more expensive, which is the practical argument for testing at the seam rather than the unit