Tags: ux experimentation concept
Usability Testing
Date: 2026-08-17
Watching someone attempt a real task with your interface, without helping them. It’s the highest-value research method per hour spent, and the discipline is almost entirely in staying quiet.
Usability testing gives a participant a task and observes them attempting it, recording where they hesitate, err, backtrack or fail.
The unit is the task, not the opinion. You are not asking whether they like it; you are watching whether they can do it.
Writing tasks
The task determines the quality of everything that follows.
BAD GOOD
"Have a look around the "You need a moisturiser
site" for sensitive skin,
→ no goal, no failure under £30. Find one
condition and add it to your basket"
"Use the filter to find "Find a product suitable
sensitive skin products" for sensitive skin"
→ tells them the method → lets them choose,
which is the finding
Never name the mechanism. If you say “use the filter”, you’ve tested whether they can operate a filter you pointed at — not whether they’d find it. The route they choose is the result.
Give them a scenario, not an instruction. A goal with a plausible motivation produces realistic behaviour; a command produces compliance.
Running one
1 set expectations
"we're testing the site, not you.
There are no wrong answers"
2 ask them to think aloud
3 give the task, then STOP TALKING
4 when they ask for help, deflect
"what would you do if I weren't here?"
5 after each task, probe what happened
6 never explain the design
Step 3 is the whole method. The instinct to rescue a struggling participant is strong and destroys the data — the struggle is the finding.
Think-aloud has a known cost: narrating changes behaviour slightly, usually making people more deliberate and slower. It’s accepted because the insight is worth more than the distortion, but it means task timings from a think-aloud session aren’t a clean measurement.
What you record
TASK SUCCESS completed · completed with
difficulty · failed ·
gave up
WHERE THEY LOOKED and where they didn't
ERRORS and whether they recovered
HESITATIONS the pause before a click
VERBATIM QUOTES the exact words
ASSUMPTIONS what they expected to happen
Hesitation is the most under-recorded signal. A two-second pause before clicking means the label didn’t do its job, even when the click is correct.
Five participants, and the caveats
The well-known guidance is that five participants surface most usability problems, on a model where each has an independent chance of hitting any given issue.
% of problems found, at a 31% per-user
detection rate
1 user 31%
3 users 67%
5 users 84%
8 users 95%
10 users 98%
Two caveats that get dropped:
- It’s five per distinct segment. Trade buyers and consumers are different populations; five of each, not five in total
- 31% is an assumption, not a law. At a 10% detection rate, five users find only 41% — and detection rates fall on complex, specialised or infrequently-used interfaces
The honest rule is saturation — keep going until sessions stop producing new findings — Sample Size in Qualitative Research.
Variants
- Moderated — a facilitator, probing live
- Unmoderated — recorded, at scale, no probing — Moderated vs Unmoderated Testing
- Guerrilla — short, informal, whoever’s available. Cheap and rough
- First-click — one screen, one question: where would you click?
- Five-second — show it briefly, ask what it was for. Tests comprehension
Where it goes wrong
- Testing with colleagues. They know the product, the vocabulary and what you want to hear
- Rescuing. The most common facilitator failure, and the most costly
- Treating small-sample rates as measurements. “3 of 5 failed, so 60%” is not a rate — Qualitative vs Quantitative Research
- Testing too late. A test after build produces findings nobody has budget to act on. Test the prototype
- Only testing the happy path. Out of stock, payment declined, wrong size — the failure paths are where usability problems concentrate and where they’re least often tested — Component States