Tags: ux statistics concept

Sample Size in Qualitative Research

Date: 2026-08-17


How many participants is enough, and why the answer is a stopping rule rather than a number. “Five users” is a real finding from a specific model with assumptions people routinely drop — and the model says nothing about how common anything is.


Qualitative sample size is the number of participants needed to surface the issues that exist — not to estimate how often they occur. Those are different questions and only the first is answerable here.

The five-user model

The well-known guidance comes from a simple model: if each participant has an independent probability L of encountering any given problem, then the proportion of problems found after n participants is:

Worked, at the commonly-cited L = 31%:

 n     1 − (1 − 0.31)ⁿ
 1          31.0%
 2          52.4%
 3          67.1%
 5          84.4%    ← "five users"
 8          94.9%
10          97.6%
15          99.6%

In plain terms: each additional participant finds a share of what’s left, so the returns fall away sharply. The fifth participant adds a lot less than the second, which is the actual argument — not that five is magic.

Where the model breaks

L = 31% is an assumption, drawn from particular studies of particular interfaces. Change it and the recommendation changes completely:

             5 users   10 users
L = 10%        41%        65%
L = 20%        67%        89%
L = 31%        84%        98%
L = 50%        97%       100%

At L = 10%, five participants find fewer than half the problems. Detection rates fall on complex interfaces, specialised audiences, infrequent tasks, and anything where problems are subtle rather than blocking.

And the relationship runs the wrong way from intuition: a more usable interface needs more participants, because problems are rarer and each participant is less likely to hit any given one. The five-user guidance is least reliable on the interfaces that are already good — which is exactly when a programme is doing iterative refinement.

Three more caveats that get dropped:

  • It’s five per segment. Two distinct audiences means two studies. Trade and consumer buyers do not substitute for each other
  • It assumes independence. Problems that only appear in combination, or only for people with particular prior knowledge, violate it
  • It’s about finding problems, not ranking them. The model says nothing about which matter most

Saturation is the honest rule

Rather than a fixed number, stop when new sessions stop producing new findings.

session 1   9 new issues
session 2   6 new
session 3   4 new
session 4   2 new
session 5   1 new
session 6   0 new      ← approaching saturation
session 7   0 new      ← stop

Track it live. Logging new-versus-repeat findings per session turns “how many do we need” into an observation rather than a negotiation, and it justifies both stopping early and continuing.

Saturation is reached faster for narrow questions and slower for broad ones, which is another way of saying a tightly-scoped study needs fewer people.

What you cannot do with these numbers

The rule that matters most in practice:

VALID
  "participants did not see the delivery
   cost until checkout"
  "this is a real problem"

INVALID
  "3 of 8, so 37.5% of customers"
  "this is the most common issue"
  "60% task success rate"

A qualitative sample supports existence claims, not frequency claims. The participants weren’t randomly drawn from your customers, so no proportion computed from them estimates anything about your customers — Qualitative vs Quantitative Research, Populations and Samples.

In plain terms: finding a problem with five people tells you the problem is real. It tells you nothing about whether it affects 2% or 60% of visitors — and sizing it is analytics’ job, not research’s — Funnel Analysis.

Rough working numbers

Not rules, but the ranges that usually apply:

usability testing        5–8 per segment
user interviews          5–8 per segment,
                         to saturation
card sorting             15–30 (this one is
                         analysed statistically)
diary studies            8–15
surveys                  a genuinely
                         quantitative sample
                         — Sample Size Calculation

See: Sample Size Calculation

Card sorting is the exception in this list — it’s analysed with clustering across participants, so it needs enough people for the agreement patterns to be stable — Card Sorting and Tree Testing.