Tags: experimentation concept
Institutional Learning
Date: 2026-08-17
Turning individual results into beliefs the organisation holds, and revising those beliefs when new results contradict them. It’s the output that outlasts every shipped winner — and the one that requires someone to do a job nobody is assigned, which is why most programmes accumulate results without accumulating knowledge.
Institutional learning is the process of turning test results into beliefs the organisation holds and acts on — and revising them when later results disagree.
Result versus belief
RESULT BELIEF
"Test 214: adding delivery dates "Our customers' main uncertainty at the
to the PDP lifted checkout PDP is when it arrives, not whether it
completion +4.1%" fits. Reducing delivery uncertainty is
our most reliable lever."
one test, one page, one moment a claim about customers
tells you what happened tells you what to try next
can be looked up can be argued with, and revised
A programme with 200 archived results and no beliefs will test delivery messaging again in eighteen months, because nobody generalised the first answer. The archive is the raw material; the belief is the product — Experiment Archive.
How a belief gets formed
Not from one test. The honest sequence:
1 one result interesting. mechanism proposed, not established
2 a second, different test same mechanism, different surface —
pointing the same way now it's a hypothesis worth holding
3 a deliberate replication pre-registered, on a new surface. this is the
step that gets skipped and the one that matters
4 a stated belief written down, with the evidence attached
and a confidence level
5 ongoing challenge any contradicting result reopens it
Step 3 is the whole difference between a learning organisation and a superstitious one. Without deliberate replication, beliefs form from whichever result was most memorable — usually the largest, which is also the most inflated by Winner’s Curse.
Writing them down
A belief register is a short document, not a wiki. Each entry:
BELIEF Delivery-date certainty is a stronger lever than price
framing on product pages.
CONFIDENCE Moderate. Three tests, two surfaces, consistent direction.
Never tested on first-time visitors specifically.
EVIDENCE T214 PDP delivery dates +4.1% (CI +1.2 to +7.1)
T251 basket delivery countdown +2.6% (CI +0.4 to +4.9)
T263 price-match badge +0.2% (CI −1.8 to +2.2) ← the
contrast that supports it
CONTRADICTS T198 (2024) delivery banner on category pages: flat.
Possibly wrong surface — intent is lower at category level.
STATUS Held since 2026-03. Next challenge: first-time visitors.
The CONTRADICTS field is the one that keeps it honest. A belief register that only records supporting evidence is a marketing document, and it will be defended rather than tested.
Why beliefs need to be revisable
The main risk of writing beliefs down is that they calcify — a two-year-old finding becomes “we know that doesn’t work here”, and nobody tests it again.
Guards against that:
- Date-stamp every belief and re-examine on a schedule. Annually is enough. The site, the traffic and the customers have all changed
- Record confidence explicitly, and low confidence should read as an invitation
- Treat a contradicting result as information, not as an error. The instinct is to find a flaw in the new test; sometimes that’s right, and the check must be applied symmetrically to the old one, which is usually less well documented
- Note what has changed since. A belief formed before a replatform, a redesign or a shift in traffic mix is on thin ground — Replatforming
- Distinguish “tested and false” from “never tested”. Both get stated as “that doesn’t work here” in meetings, and only one of them is evidence
The failures worth naming
- Nobody owns it. Analysts run tests, product ships winners, and the synthesis is nobody’s objective. It needs an explicit owner and time allocated
- Only winners get remembered. Nulls and losses carry as much information — often more, because they contradict a prior — and they leave no artefact unless someone writes one — Inconclusive Results
- The person leaves. Undocumented belief is the most fragile asset a programme has, and this is the single most common way a mature programme regresses
- Results without context. A result stripped of its date, its traffic conditions and its concurrent campaigns can’t be reinterpreted later, which is exactly when you need it — Annotation and Change Logs
- Beliefs from one segment applied to all. “Our customers respond to urgency” formed entirely on mobile paid-social traffic — Heterogeneous Treatment Effects
- Vendor case studies treated as evidence. Someone else’s win on someone else’s site is a hypothesis, and a heavily selected one
The output nobody counts
The strongest argument for this work is that the transferable knowledge usually exceeds the shipped uplift in value, and it’s completely absent from any programme report built on win rates and lifts.
A team that knows delivery certainty beats price framing on its own traffic makes better decisions in every project — including the ones that never get tested, which is most of them. That’s not measurable, which is precisely why it needs deliberate protection from a reporting culture that only counts what is.
Where it interacts
- Experiment Archive — the record this reads from, and the reason its search and metadata quality matters
- Hypothesis Design — beliefs are where good hypotheses come from, closing the loop
- Experimentation Maturity — reaching this stage is roughly what distinguishes a mature programme from a busy one
- Win Rate and Expected Value — the arithmetic half of programme value, of which this is the unquantified other half