Tags: experimentation concept

Test Prioritisation

Date: 2026-08-17


Deciding which test to run next. The scoring frameworks — ICE, PIE, PXL — are mostly theatre: they convert guesses into numbers, and the numbers inherit none of the authority the arithmetic implies. What they’re genuinely good for is forcing a conversation, and they should be judged on that rather than on their output.


Test prioritisation is ranking candidate tests to decide which gets the next slot of traffic and build time.

The frameworks

ICE — Impact, Confidence, Ease. Each scored 1–10, averaged or multiplied.

PIE — Potential, Importance, Ease. Same shape; “Importance” weights by traffic, which is the one real improvement.

PXL — a CRO-specific variant that replaces subjective 1–10 ratings with binary yes/no questions (“is it above the fold?”, “does it address an issue found in user research?”), which removes some of the guesswork by removing the scale.

ICE, worked

test                                  I     C     E    score
──────────────────────────────────────────────────────────────
Simplify checkout to one page         9     5     2     5.3
Change CTA colour                     2     6    10     6.0   ← "wins"
Add delivery estimates to PDP         7     8     6     7.0
Remove the carousel                   5     7     8     6.7

The carousel and the CTA colour outrank the checkout rebuild, because ease is weighted identically to impact and is the only one of the three anyone can estimate accurately. This is the structural bias of every additive framework: it promotes cheap, low-impact work.

Why the numbers don’t mean what they look like

  • Impact is the thing you’re running the test to find out. If you could estimate it, you wouldn’t need the test. Estimated impact correlates poorly with measured impact — this is one of the more consistently reported findings in the experimentation literature, and it’s the central argument of Kohavi - Trustworthy Online Controlled Experiments - 2020
  • Confidence is a restatement of impact, so the two multiply the same guess by itself
  • Ease is the only honest number, so it dominates
  • Scores are not comparable across people. One person’s 7 is another’s 4, and averaging across a team launders that disagreement into a decimal
  • A 1–10 scale invites false precision. Nobody can distinguish a 6 from a 7, and the difference decides the ordering

Two decimal places on a guess is not analysis. The output of ICE is an ordering of opinions, and it should be described that way in the room.

What they’re actually for

The score is worthless; the argument that produces it is not.

The value is that scoring forces three questions to be asked out loud: what do we think will happen, how sure are we, and what will it cost. Teams that skip prioritisation entirely don’t ask any of them — they run whatever the loudest person suggested. Judged as a discussion protocol rather than as a ranking algorithm, ICE earns its place.

So: use it to surface disagreement, then discard the number. Where two people score Impact 3 and 9, the gap is the interesting object — it usually means they disagree about the mechanism, and that conversation is worth more than the whole exercise.

Better inputs

If you’re going to score, score things that are knowable:

  • Traffic to the page — knowable exactly, and it determines whether the test is powerable at all. A brilliant hypothesis on a page with 400 weekly visitors cannot be tested — Sample Size Calculation
  • Baseline conversion and its variance — determines detectable effect
  • Evidence behind the hypothesis — from user research, session replay, funnel analysis, a prior test. Ranked evidence beats guessed impact, and it’s the substance of PXL’s approach — Hypothesis Design
  • Value per unit of the metric — a 1% lift on a page in the £2m path is worth twenty times the same lift on a £100k path
  • Reversibility — cheap to undo means cheap to be wrong about
  • What you learn if it fails. A test whose null result is informative is worth more than one that just stops

The strongest reframe: prioritise by expected value in pounds, not by score.

test: delivery estimates on PDP

annual revenue through the page          £2,400,000
plausible relative lift, if it works           +2%
probability it works (base rate)               25%     ← your own win rate,
                                                          not optimism
expected value    2,400,000 × 0.02 × 0.25 = £12,000/yr
cost              6 dev-days + 3 weeks of traffic

That’s a number with real inputs, and the honest one is the 25% — use your programme’s actual win rate rather than a feeling, which is one more reason the Experiment Archive pays for itself — Win Rate and Expected Value.

The bigger lever

Prioritisation matters much less than people expect, because the variance in test outcomes swamps the variance in ordering. If a quarter of tests win and effect sizes are unpredictable, running twelve tests in the “wrong” order beats running six in the right one.

Optimise throughput before ordering: reduce the time from idea to live, run tests concurrently where they don’t interact, and stop spending two weeks ranking a backlog you could have half-finished — Experimentation Velocity, Interaction Effects.

The exception is when tests are genuinely scarce — low traffic, one test at a time, each taking six weeks. Then ordering is most of the value and the discussion deserves the time.

Where it interacts