Tags: experimentation content

Kohavi - Trustworthy Online Controlled Experiments - 2020

Author: Ron Kohavi, Diane Tang, Ya Xu
Published: 2020, Cambridge University Press
Format: Book, ~250pp
Read:


The field’s standard reference, and the source of most of its vocabulary. Cited constantly, read properly by few — its value is that everyone in experimentation has agreed to treat it as the baseline, so knowing what’s in it is knowing the shared language.


Why it’s the one that gets named

Experimentation had no canonical text before it. The knowledge lived in conference papers, company engineering blogs and individual practitioners’ heads. This book collected it and gave it a shared vocabulary, which is why it’s the default citation and why “have you read Kohavi” functions as a shibboleth.

Referred to as “the Kohavi book” or just “Kohavi”, despite three authors.

Who wrote it, and why that matters

Worth knowing, because the authors’ positions are the book’s authority:

Ron Kohavi     built and led Microsoft's
               experimentation platform (ExP)
               later Airbnb, Amazon

Diane Tang     Google Fellow, Google's
               experiment infrastructure

Ya Xu          led data science at LinkedIn

Three of the largest experimentation platforms in the world, written up together. That’s the claim to authority, and it’s a fair one — but it’s also the source of the book’s main weakness, below.

The organising idea

Trustworthiness, not statistics. The title is the argument: the hard part isn’t running a test, it’s knowing whether to believe the result.

The book’s implicit thesis is that most experiment results are wrong for boring, mechanical reasons, and that a mature programme spends most of its effort on validity rather than on analysis.

What those reasons look like in practice:

instrumentation  the event didn't fire for
                 some users, browsers or
                 devices — the metric is
                 measuring coverage, not effect

assignment       users landed in the wrong arm,
                 or moved between arms across
                 visits or devices

sample ratio     the split isn't the split you
                 asked for, which means
                 something upstream is
                 filtering one arm

dilution         users who could never have
                 seen the change are sitting
                 in the analysis

peeking          the result was read early and
                 often, so the stated error
                 rate is not the real one

Every one of these produces a plausible-looking number. That’s the point — none of them announce themselves, and collectively they are more likely than the result being a genuine effect of the size you were hoping for.

That framing is its real contribution, more than any individual technique.

What it put into circulation

These are the terms people use because of this book:

TermWhat it means
OEC — Overall Evaluation CriterionOne metric a test optimises, causally linked to a long-term objective and measurable short-term — Overall Evaluation Criterion
Twyman’s law”Any figure that looks interesting or different is usually wrong”
SRM — Sample Ratio MismatchArms not splitting as designed; the single highest-value automated check — Sample Ratio Mismatch
Guardrail metricsThings that must not get worse — Guardrail Metrics
Triggering and dilutionAnalyse only users who could have been affected, or the effect is diluted by everyone who never saw it
Crawl · walk · run · flyThe experimentation maturity model
Institutional memoryThe accumulated record of what was tried — Experiment Archive

Twyman’s law is not theirs. It comes from W. A. Twyman, a UK audience-measurement researcher, and he apparently never wrote it down. Kohavi popularised it in this context — a useful thing to get right, since crediting it to Kohavi is a common slip.

The three worth being able to explain

Naming these is table stakes. Explaining them is the thing that separates having read the book from having heard of it.

OEC — why one metric

The argument: a team optimising several metrics optimises none, because every ambiguous result becomes a negotiation about which metric mattered. An OEC forces that trade-off to be settled before the result is known, which is the only time it can be settled honestly.

Two conditions, and both are hard:

  • Measurable within a test’s duration — a metric that takes a year to move can’t referee a two-week test
  • Causally linked to the long-term objective — moving it must actually cause the thing you care about, not merely correlate

The failure it exists to prevent is optimising a proxy that diverges from the goal. Clicks are measurable in days and can be increased by making search results worse — so a team optimising clicks can degrade the product while reporting success every quarter — Overall Evaluation Criterion, Metric Design.

Triggering and dilution

The most practically useful technique in the book, and the least known outside it.

If a change can only affect some users, including everyone else drags the measured effect towards zero:

change affects only users who reach
the delivery step — say 20% of visitors

MEASURED ON EVERYONE
  true effect 5% on those who saw it
  5% × 20%  =  1% observed
  → underpowered, reads as flat

MEASURED ON TRIGGERED USERS ONLY
  → 5% observed
  → the effect is actually visible

Analyse only the users who could have been affected. The catch is instrumentation: you must log the trigger point in both arms — recording who would have seen the change in the control group, where by definition nothing happened. That’s a decision made before the test runs and cannot be reconstructed afterwards, which is why it’s so often missed — Experiment Assignment Tracking.

Crawl, walk, run, fly

The maturity model, referenced constantly and rarely explained:

CRAWL   a handful of tests a quarter
        analysis is manual
        the platform itself is unproven

WALK    tests are routine
        metrics standardised
        SRM and guardrail checks automatic

RUN     experimentation is the default way
        changes ship; dozens run at once
        the org expects tests, not opinions

FLY     nearly every change is an experiment,
        infrastructure included
        results feed institutional memory

Its actual purpose is to stop teams comparing themselves to Microsoft. The common failure is a programme at crawl adopting fly-stage practices — sophisticated variance reduction on twelve tests a quarter — instead of fixing the instrumentation that makes those twelve untrustworthy.

The stories everyone cites

The Bing ad headline experiment. The canonical anecdote of the whole field.

early 2012   an idea to change how ad titles
             displayed — one of hundreds
             stack-ranked below other work,
             so it sat for months

later        an engineer built it in days
             and ran it

             a revenue alert fired —
             Bing was making TOO MUCH money
             ↑ the detail that makes the story

result       +12% revenue
             ~$100 million/year
             Bing's best revenue idea ever

Why it gets told. Three separate arguments come out of one story, which is what makes it efficient — but each needs drawing out, because the story doesn’t state them.

1. Human prioritisation is unreliable. The idea was ranked below hundreds of others by experienced people exercising judgement, and it turned out to be worth more than all of them. The lesson isn’t that the rankers were incompetent — it’s that nobody can rank ideas by expected value before testing them, because the information needed to rank is exactly what the test produces. That’s the argument for cheap, high-volume testing over careful selection — Test Prioritisation.

2. Effect size doesn’t track effort or ambition. A cosmetic change to how ad titles rendered, built in days, beat every strategic initiative in the company. Meanwhile large redesigns routinely produce nothing. There is no reliable relationship between how significant a change feels and how much it moves — which is why pre-test effect estimates are closer to guesswork than anyone admits — Minimum Detectable Effect.

3. Trustworthiness machinery is what lets you believe a good result. The revenue alert existed to catch bugs. It fired, and the team’s first assumption was that the experiment was broken — Twyman’s law behaving exactly as advertised. They only accepted +12% after checking it wasn’t an error.

That third beat is the one most retellings drop, and it carries the book’s whole thesis. Without the checking apparatus, a 12% revenue result is indistinguishable from a bug, so it either gets dismissed or gets shipped on faith. An organisation that cannot verify its own wins cannot act on them.

The speed experiments — that page performance measurably moves revenue at large scale — are the other frequently-cited set. [CHECK: the specific millisecond-to-revenue figures before quoting any; several circulating versions conflate different companies and years.] The durable point is the direction, not the number — Performance and Conversion.

The numbers people quote

Verified, and worth having exactly right because they’re used to set expectations:

ACROSS MICROSOFT
  ~1/3 of ideas positive and significant
  ~1/3 flat
  ~1/3 negative and significant

AT BING SPECIFICALLY
  ~15% of launched experiments succeeded

BING RELEVANCE TEAM
  0.1–0.2% improvement per iteration
  ~2% accumulated annually

The “one third” figure is the one to know. It’s the standard counter to “our idea will obviously work”, and it reframes a programme’s purpose from validating ideas to killing them cheaply — Inconclusive Results, Win Rate and Expected Value.

The Bing relevance number is the underrated one: a serious programme’s wins are tiny and compound. That’s a more useful expectation-setter than any case study.

How it’s built

Five parts, and knowing which is which is the difference between finding it useful and finding it impenetrable:

1  Introductory topics for everyone
2  Selected topics for everyone
3  Complementary and alternative techniques
4  Advanced — BUILDING A PLATFORM
5  Advanced — ANALYSING EXPERIMENTS

Parts 1 and 2 are the practitioner book. Part 4 is for people building an experimentation platform, which almost no reader is. Part 5 is statistically dense.

Part 3 is the underrated one and the most relevant to a CRO role at ordinary scale. It covers what to do when a controlled experiment isn’t available — observational and quasi-experimental methods, interrupted time series, instrumented variables, difference-in-differences — plus user research as a complement rather than a rival. Most commercial work runs into “we can’t randomise this” constantly, and this is the part that addresses it — Difference-in-Differences, Incrementality Testing, Geo Holdout Tests.

The honest verdict

Reception is genuinely mixed, which is worth knowing rather than deferring to the reputation — the reputation is for its contribution to the field, and that’s a separate claim from it being a good read.

One reviewer’s summary is the fairest short version: “indispensable, but some assembly required.”

The recurring criticisms:

  • It reads as expanded bullet points rather than as a written argument — closer to a busy academic’s class notes, with much of the detail deferred to the papers it cites
  • Examples arrive without setup. They’re numerous, and their significance often goes unexplained, so the reader is left working out why each one is in the book
  • Heavy self-citation. Reasonable, given the authors built the platforms in question — but it means the evidence base is narrower than the citation count suggests
  • Three platform-builders writing for platform-builders, even in the parts labelled “for everyone”. The implicit reader has thousands of experiments a year and a dedicated data science team
  • Uneven in difficulty, moving between “what is an A/B test” and variance-reduction methods without signalling the change
  • Nothing is worked end to end. Concepts arrive in the authors’ order rather than the order the work happens

It also draws substantial praise, with plenty of readers finding it accessible and practical. The split appears to track what the reader came for.

What it’s genuinely good for: the vocabulary, the trustworthiness checklist, the numbers, and Chapter 3 on Twyman’s law. Excellent as a reference to look things up in; frustrating as a book to read through — which is roughly the opposite of how it presents itself.

Where this vault does it better

Not a boast — a filing note. The concepts are covered here in the order the work happens, so this note deliberately doesn’t repeat them:

The one-paragraph take

“It’s the standard reference, and it’s where most of the vocabulary comes from — OEC, SRM, guardrails, Twyman’s law. The framing I took from it is that the hard part is trustworthiness rather than statistics: most results are wrong for mechanical reasons before they’re wrong for statistical ones. The one-third-positive figure is the useful one for setting expectations. It’s written by platform-builders for platform-builders though, so it’s better as a reference than a read.”