Tags: experimentation concept
The HiPPO Problem
Date: 2026-09-27
Decisions made by the highest-paid person’s opinion rather than by evidence. The problem isn’t that senior people have bad judgement — it’s that nobody’s judgement predicts which ideas will work, and seniority makes that judgement hardest to question. Experiments move the argument from rank to results.
HiPPO — highest-paid person’s opinion — is the decision-making pattern where the most senior person’s view settles a question that evidence could have answered.
Where the term comes from
Popularised in web analytics and experimentation circles in 2006–07. The Experimentation Platform site run by Ronny Kohavi gives the history: an Intuit employee, Dylan Lewis, used “HiPO” (highest-paid opinion) in 2006; Avinash Kaushik picked it up and it became “HiPPO” in conversation with Kohavi — at Microsoft “HiPO” already meant high-potential employee. Kaushik blogged about it that year, and it appeared in a peer-reviewed paper at KDD (a data mining conference) in 2007 — Kohavi - Trustworthy Online Controlled Experiments - 2020.
Why opinions — anyone’s — are weak evidence
The case against HiPPO decisions isn’t about the HiPPO. It’s the base rate of ideas working:
MICROSOFT, ACROSS EXPERIMENTS (Kohavi et al.)
~1/3 positive and significant
~1/3 flat
~1/3 negative and significant
If experienced, well-informed people’s ideas succeed about a third of the time, then shipping on conviction is shipping something that has a one-in-three chance of actively making things worse — Win Rate and Expected Value.
Why seniority makes it worse, not better:
- Confidence isn’t calibration. Seniority comes from being right about big, slow decisions, where feedback is rare and noisy. It says little about predicting a checkout change
- The room defers. Once the senior view is stated, the junior evidence isn’t offered
- Nobody learns. An idea shipped without a test can never be shown to have failed, so the judgement is never corrected — Institutional Learning
What it looks like
- A redesign launched to company-wide applause and never measured
- “We don’t need to test that, it’s obviously better”
- A test result overruled because the loser was the boss’s idea
- Tests stopped early when the preferred variant is ahead — Stopping Rules
- Prioritisation where the ranking follows who proposed each idea — Test Prioritisation
Moving the argument off rank
- Agree the metric and the decision rule before the test. “If conversion rises by at least X without hurting Y, we ship” — written down, before anyone knows the result. Then the result decides, not the rank of whoever dislikes it — Pre-Registration, Overall Evaluation Criterion
- Frame tests as protecting the idea. “Let’s make sure it’s not costing us” gets a yes where “let’s check if you’re right” doesn’t
- Share the base rate early. Once a team knows most ideas fail, testing stops sounding like doubt
- Test the HiPPO’s ideas properly, and let some win. A programme that only ever says no loses its sponsor
- Keep an archive of surprises — results that went against expert prediction are the most persuasive evidence a programme has — Experiment Archive
Where it’s overused
Not every decision needs a test, and “the HiPPO” is sometimes a label for legitimate strategic judgement — brand, legal risk, long-term positioning, or changes too small or too rare to test. The problem is opinion standing in for evidence that could have been collected, not seniority making decisions at all.
A programme where a test has overruled something senior is a sign of maturity — Experimentation Maturity.