Tags: experimentation concept

Overall Evaluation Criterion

Date: 2026-08-16


One metric, nominated before launch, that decides the test. Everything else is diagnostic. More than one primary metric isn’t thoroughness — it’s a licence to pick whichever one won.


What it is

The overall evaluation criterion (OEC) is the single metric a test is judged on, chosen in advance. It’s the answer to “if only one number could decide this, which?”

The name comes from Microsoft’s experimentation literature and is used interchangeably with “primary metric”.

Why exactly one

Report six metrics and you have made six comparisons. At 95% significance the chance of at least one crossing the line by luck is:

1 − 0.95⁶ = 26.5%

Better than one in four tests will produce a “winner” on some metric while nothing is happening. If the decision rule is “did anything win”, you’ve built a machine that ships noise. See The Multiple Comparisons Problem.

Nominating one metric in advance costs nothing and removes the problem entirely — the other five can still be reported, they just can’t decide anything.

In plain terms: if you don’t say which number counts before you look, you’ll choose the one that agrees with you afterwards, and you won’t notice you did it.

Choosing it

The OEC has to satisfy three things at once, and they pull against each other.

RequirementMeaningPulls towards
SensitiveDetectable at your traffic — see Metric SensitivityUpper-funnel rates
MeaningfulA real move in it is a real business gainRevenue
AttributablePlausibly influenced by this specific changeThe step being tested

Revenue per visitor is maximally meaningful and needs 2.4× the traffic of conversion rate. Add-to-cart rate is cheap to detect and may not translate. The usual resolution for retail CRO:

  • OEC: conversion rate, user-scoped
  • Secondary: revenue per visitor, reported with its interval, explicitly not decision-eligible
  • Diagnostic: add-to-cart, checkout entry, step completion — for explaining why the OEC moved

Attributability is the underrated one. Testing a category page filter and judging it on site-wide revenue per visitor means most of the metric’s variance comes from things your change can’t touch, which destroys sensitivity for no benefit.

Composite OECs

Where no single metric works, some teams combine several into one weighted score — conversion plus retention plus satisfaction, say.

The appeal is that it captures trade-offs a single metric misses. The costs are real:

  • Weights are arbitrary and quietly encode a strategy nobody agreed
  • The result is uninterpretable. “The score went up 3%” doesn’t say what happened
  • It’s gameable. A variant that wrecks one component and lifts another can win

Worth it when a single metric is genuinely perverse — a change that lifts conversion by degrading the product would win on conversion alone. Usually better handled by an OEC plus Guardrail Metrics, which keeps the primary interpretable and still prevents harm.

Long-term versus short-term

The deepest problem with any OEC: a two-week test measures two weeks. Changes that lift immediate conversion while damaging trust, repeat purchase or brand will win.

The available defences, none complete:

  • Guardrails on leading indicators of harm — returns rate, support contacts, unsubscribes
  • Holdout Groups — a permanently untreated slice, compared after months, which is the only way to see cumulative effect
  • Post-Test Validation — checking the shipped change still delivers after novelty fades
  • Judgement. Some changes shouldn’t be tested because a win would be the wrong answer — see Ethics of Experimentation

Failure modes

  • Changed mid-test, or after the result. This is the one that most destroys a programme’s credibility
  • Chosen for sensitivity alone — a test optimised to produce significant results rather than useful ones
  • Different OEC per test, making the archive incomparable and meta-analysis impossible
  • Session-scoped when the Randomisation Unit is the user — the denominator can move with the treatment
  • Never written down, so “primary” is decided in the readout meeting