Tags: experimentation ux concept
Ethics of Experimentation
Date: 2026-08-17
Online experiments are run on people who didn’t agree to be in them, mostly can’t tell, and have no way to opt out. That’s usually fine and occasionally isn’t — and the line isn’t drawn by whether anyone complained, because the design guarantees they can’t.
The ethics of experimentation is the question of when it’s acceptable to vary someone’s experience, without their knowledge, in order to measure their response.
Why the usual defence isn’t sufficient
The standard justification is that A/B testing is just product development with measurement, and every site changes constantly without asking anyone. That’s true and it covers most tests.
Where it stops working:
covered by the defence not covered
which of two layouts is clearer which of two emotional appeals
makes people spend more than
they intended
whether a form is easier with whether inducing anxiety about
fewer fields stock levels increases orders
whether faster pages convert deliberately degrading service
to measure the tolerance threshold
any change you'd happily anything you'd rather customers
describe to the customer didn't know you were doing
The practical test: could you describe this test to the people in it without embarrassment? It’s not a rigorous principle, and it catches nearly everything that matters.
The three questions
1. Could a participant be harmed?
Financial harm (spending more than intended, worse terms), psychological harm (manufactured anxiety, shame), and harm through exclusion (a variant that fails for screen-reader users, or degrades on low-end devices) — Accessibility Testing.
Testing anything that induces a decision people would regret is the bright line. Scarcity messaging that’s accurate is a design choice; scarcity messaging that’s fabricated is a lie you’re measuring the profitability of — Scarcity and Urgency, Deceptive Design.
2. Is the loss to the losing arm acceptable?
Every test knowingly gives some users a worse experience — that’s what a control group is. The question is magnitude and duration, not existence.
- Small differences, short duration, easily reversed → fine
- A variant that materially damages someone’s experience of a service they’re paying for → not fine, and a guardrail should catch it fast — Guardrail Metrics
- Deliberately degrading service to measure sensitivity — slowing pages, hiding stock, worsening delivery estimates — is the case that needs explicit sign-off rather than an analyst’s judgement
3. Who’s in it?
Some populations warrant more care regardless of the test: people in financial difficulty, people seeking health-related products, children, and anyone in a moment of vulnerability. A test that’s unremarkable on a homepage can be indefensible on a page about debt or bereavement.
Consent, honestly
Meaningful consent to A/B testing isn’t available. A banner saying “we run experiments” that nobody reads isn’t consent; excluding non-consenters would bias the sample and defeat the test.
So the honest position is that legitimacy comes from restraint rather than permission: because you can’t ask, you take on the obligation to only run tests people would have agreed to. That’s a weaker guarantee, and pretending otherwise — pointing at a terms-of-service clause — is where organisations get into trouble.
Separately, the data protection side does have rules. Under UK GDPR — the UK’s retained General Data Protection Regulation — assignment records and outcome data tied to identifiable people are personal data, needing a lawful basis, a retention period and inclusion in your privacy notice. Legitimate interests generally covers product experimentation; it does not cover everything. [CHECK: whether your specific use — particularly anything involving profiling, special category data, or automated decisions with significant effects — needs a legitimate interests assessment or a DPIA. Confirm with whoever owns data protection.] — UK GDPR and PECR for Analytics, PII in Analytics
The cases worth deciding in advance
Rather than a principle, a list of the tests that should require someone senior to say yes:
- Anything involving price for equivalent customers. Differential pricing is a commercial, legal and reputational question long before it’s a statistical one — Price Testing
- Deliberate degradation of speed, availability or service quality
- Emotional manipulation — manufactured urgency, shame, fear of missing out
- Anything touching a regulated claim — health, financial promotions, environmental claims — Testing and Compliance
- Painted Door Tests, which work by deceiving customers, however briefly
- Tests on vulnerable segments, or on journeys where distress is likely
- Anything you wouldn’t want screenshotted and posted with your logo attached
The point of the list is that it exists before the test is proposed. Judging case by case under commercial pressure produces predictable answers.
The Facebook emotional contagion study
Worth knowing because it’s the reference case everyone in this field alludes to. In 2014 Facebook published a study in which the emotional valence of nearly 700,000 users’ news feeds was manipulated to measure whether emotional states spread through a network. The backlash was substantial, and it reshaped how the industry talks about this.
The instructive part isn’t the outrage — it’s the specifics of what people objected to. The users were not informed, there was no meaningful consent, the manipulation targeted emotional state rather than interface, and the research was published, which is what made it visible at all. The uncomfortable implication is that similar manipulations happen constantly without publication and therefore without objection — [CHECK: the exact sample size and the journal’s subsequent editorial statement if you need to cite it precisely.]
What it establishes for practice: the acceptability of a test depends on what’s being manipulated, not on whether it’s technically the same mechanism as a layout test.
Where it interacts
- Deceptive Design — the pattern catalogue this is the testing counterpart of; a test that optimises a dark pattern is optimising harm efficiently
- Testing and Compliance — the legal floor, which sits below the ethical question rather than answering it
- Research Ethics — the qualitative-research counterpart, where consent genuinely is obtainable and therefore expected
- Guardrail Metrics — the mechanism that limits harm during a test, and the reason user-experience guardrails belong in the standing set alongside commercial ones