Tags: experimentation concept

Institutional Learning

Date: 2026-08-17


Turning individual results into beliefs the organisation holds, and revising those beliefs when new results contradict them. It’s the output that outlasts every shipped winner — and the one that requires someone to do a job nobody is assigned, which is why most programmes accumulate results without accumulating knowledge.


Institutional learning is the process of turning test results into beliefs the organisation holds and acts on — and revising them when later results disagree.

Result versus belief

RESULT                               BELIEF

"Test 214: adding delivery dates     "Our customers' main uncertainty at the
 to the PDP lifted checkout           PDP is when it arrives, not whether it
 completion +4.1%"                    fits. Reducing delivery uncertainty is
                                      our most reliable lever."
one test, one page, one moment       a claim about customers
tells you what happened              tells you what to try next
can be looked up                     can be argued with, and revised

A programme with 200 archived results and no beliefs will test delivery messaging again in eighteen months, because nobody generalised the first answer. The archive is the raw material; the belief is the product — Experiment Archive.

How a belief gets formed

Not from one test. The honest sequence:

1  one result                 interesting. mechanism proposed, not established
2  a second, different test   same mechanism, different surface —
   pointing the same way      now it's a hypothesis worth holding
3  a deliberate replication   pre-registered, on a new surface. this is the
                              step that gets skipped and the one that matters
4  a stated belief            written down, with the evidence attached
                              and a confidence level
5  ongoing challenge          any contradicting result reopens it

Step 3 is the whole difference between a learning organisation and a superstitious one. Without deliberate replication, beliefs form from whichever result was most memorable — usually the largest, which is also the most inflated by Winner’s Curse.

Writing them down

A belief register is a short document, not a wiki. Each entry:

BELIEF     Delivery-date certainty is a stronger lever than price
           framing on product pages.

CONFIDENCE Moderate. Three tests, two surfaces, consistent direction.
           Never tested on first-time visitors specifically.

EVIDENCE   T214 PDP delivery dates        +4.1%  (CI +1.2 to +7.1)
           T251 basket delivery countdown +2.6%  (CI +0.4 to +4.9)
           T263 price-match badge         +0.2%  (CI −1.8 to +2.2)  ← the
                                                    contrast that supports it

CONTRADICTS  T198 (2024) delivery banner on category pages: flat.
             Possibly wrong surface — intent is lower at category level.

STATUS     Held since 2026-03. Next challenge: first-time visitors.

The CONTRADICTS field is the one that keeps it honest. A belief register that only records supporting evidence is a marketing document, and it will be defended rather than tested.

Why beliefs need to be revisable

The main risk of writing beliefs down is that they calcify — a two-year-old finding becomes “we know that doesn’t work here”, and nobody tests it again.

Guards against that:

  • Date-stamp every belief and re-examine on a schedule. Annually is enough. The site, the traffic and the customers have all changed
  • Record confidence explicitly, and low confidence should read as an invitation
  • Treat a contradicting result as information, not as an error. The instinct is to find a flaw in the new test; sometimes that’s right, and the check must be applied symmetrically to the old one, which is usually less well documented
  • Note what has changed since. A belief formed before a replatform, a redesign or a shift in traffic mix is on thin ground — Replatforming
  • Distinguish “tested and false” from “never tested”. Both get stated as “that doesn’t work here” in meetings, and only one of them is evidence

The failures worth naming

  • Nobody owns it. Analysts run tests, product ships winners, and the synthesis is nobody’s objective. It needs an explicit owner and time allocated
  • Only winners get remembered. Nulls and losses carry as much information — often more, because they contradict a prior — and they leave no artefact unless someone writes one — Inconclusive Results
  • The person leaves. Undocumented belief is the most fragile asset a programme has, and this is the single most common way a mature programme regresses
  • Results without context. A result stripped of its date, its traffic conditions and its concurrent campaigns can’t be reinterpreted later, which is exactly when you need it — Annotation and Change Logs
  • Beliefs from one segment applied to all. “Our customers respond to urgency” formed entirely on mobile paid-social traffic — Heterogeneous Treatment Effects
  • Vendor case studies treated as evidence. Someone else’s win on someone else’s site is a hypothesis, and a heavily selected one

The output nobody counts

The strongest argument for this work is that the transferable knowledge usually exceeds the shipped uplift in value, and it’s completely absent from any programme report built on win rates and lifts.

A team that knows delivery certainty beats price framing on its own traffic makes better decisions in every project — including the ones that never get tested, which is most of them. That’s not measurable, which is precisely why it needs deliberate protection from a reporting culture that only counts what is.

Where it interacts