Tags: ux experimentation concept

Usability Testing

Date: 2026-08-17


Watching someone attempt a real task with your interface, without helping them. It’s the highest-value research method per hour spent, and the discipline is almost entirely in staying quiet.


Usability testing gives a participant a task and observes them attempting it, recording where they hesitate, err, backtrack or fail.

The unit is the task, not the opinion. You are not asking whether they like it; you are watching whether they can do it.

Writing tasks

The task determines the quality of everything that follows.

BAD                         GOOD
"Have a look around the     "You need a moisturiser
 site"                       for sensitive skin,
 → no goal, no failure       under £30. Find one
   condition                 and add it to your basket"

"Use the filter to find     "Find a product suitable
 sensitive skin products"    for sensitive skin"
 → tells them the method     → lets them choose,
                               which is the finding

Never name the mechanism. If you say “use the filter”, you’ve tested whether they can operate a filter you pointed at — not whether they’d find it. The route they choose is the result.

Give them a scenario, not an instruction. A goal with a plausible motivation produces realistic behaviour; a command produces compliance.

Running one

1  set expectations
     "we're testing the site, not you.
      There are no wrong answers"
2  ask them to think aloud
3  give the task, then STOP TALKING
4  when they ask for help, deflect
     "what would you do if I weren't here?"
5  after each task, probe what happened
6  never explain the design

Step 3 is the whole method. The instinct to rescue a struggling participant is strong and destroys the data — the struggle is the finding.

Think-aloud has a known cost: narrating changes behaviour slightly, usually making people more deliberate and slower. It’s accepted because the insight is worth more than the distortion, but it means task timings from a think-aloud session aren’t a clean measurement.

What you record

TASK SUCCESS       completed · completed with
                   difficulty · failed ·
                   gave up
WHERE THEY LOOKED  and where they didn't
ERRORS             and whether they recovered
HESITATIONS        the pause before a click
VERBATIM QUOTES    the exact words
ASSUMPTIONS        what they expected to happen

Hesitation is the most under-recorded signal. A two-second pause before clicking means the label didn’t do its job, even when the click is correct.

Five participants, and the caveats

The well-known guidance is that five participants surface most usability problems, on a model where each has an independent chance of hitting any given issue.

% of problems found, at a 31% per-user
detection rate

 1 user    31%
 3 users   67%
 5 users   84%
 8 users   95%
10 users   98%

Two caveats that get dropped:

  • It’s five per distinct segment. Trade buyers and consumers are different populations; five of each, not five in total
  • 31% is an assumption, not a law. At a 10% detection rate, five users find only 41% — and detection rates fall on complex, specialised or infrequently-used interfaces

The honest rule is saturation — keep going until sessions stop producing new findings — Sample Size in Qualitative Research.

Variants

  • Moderated — a facilitator, probing live
  • Unmoderated — recorded, at scale, no probing — Moderated vs Unmoderated Testing
  • Guerrilla — short, informal, whoever’s available. Cheap and rough
  • First-click — one screen, one question: where would you click?
  • Five-second — show it briefly, ask what it was for. Tests comprehension

Where it goes wrong

  • Testing with colleagues. They know the product, the vocabulary and what you want to hear
  • Rescuing. The most common facilitator failure, and the most costly
  • Treating small-sample rates as measurements. “3 of 5 failed, so 60%” is not a rate — Qualitative vs Quantitative Research
  • Testing too late. A test after build produces findings nobody has budget to act on. Test the prototype
  • Only testing the happy path. Out of stock, payment declined, wrong size — the failure paths are where usability problems concentrate and where they’re least often tested — Component States