Tags: ux statistics concept
Usability Metrics
Date: 2026-08-17
Numbers describing how well people can use something. Most are collected from small non-random samples, which makes them directional rather than measurements — and treating a task success rate from eight people as a percentage of your customers is the standard error.
Usability metrics quantify effectiveness, efficiency and satisfaction — the three dimensions in the standard definition of usability.
EFFECTIVENESS can they complete it?
→ task success rate,
error rate
EFFICIENCY at what cost?
→ time on task, steps,
errors per task
SATISFACTION how did it feel?
→ SUS, ease-of-task
ratings
— System Usability Scale
The core measures
TASK SUCCESS completed / attempted
→ the primary measure
→ define "success"
before testing
TIME ON TASK for successful attempts
only — a fast failure
isn't fast
ERROR RATE errors per task, and
recovery rate
EFFICIENCY steps taken vs the
minimum
SEQ Single Ease Question:
"how difficult was
that task?" 1–7, asked
immediately after each
task
← cheap, and correlates
well with success
SEQ is the highest value-per-effort metric here — one question, asked after each task, and it surfaces tasks people completed but found hard, which task success alone hides.
The sample size problem
This is the part that gets misused.
8 participants, 5 succeeded
REPORTED AS "62.5% task success rate"
ACTUALLY 5 of 8 non-randomly
recruited people succeeded
the confidence interval
around that proportion is
enormous, and the sample
wasn't drawn from your
customers anyway
A proportion from a small qualitative sample is not an estimate of anything. Report the raw counts — “5 of 8 completed the task” — which is honest and just as actionable — Sample Size in Qualitative Research, Qualitative vs Quantitative Research.
In plain terms: the number tells you a problem exists. It does not tell you how many customers hit it. Sizing that is analytics’ job — Funnel Analysis.
Benchmarking is where it earns its keep
Usability metrics are far more useful as comparisons than as absolutes:
COMPARE
before vs after a redesign
variant A vs variant B
your site vs a competitor
this quarter vs last, same tasks
DON'T COMPARE
your 68 SUS against another
company's published 74
→ different tasks, participants,
context
Same tasks, same script, same recruitment criteria is what makes a comparison meaningful. Change any of them and the difference is uninterpretable.
Behavioural metrics from live data
Distinct from lab metrics, and genuinely quantitative because the sample is everyone:
- Task completion — funnel conversion — Funnel Analysis
- Error rates — form field errors — Form Analytics
- Rage clicks — repeated clicking on something unresponsive — Session Replay
- Search refinement — the first attempt failed
- Back-button use — from product to category
- Support contacts, by topic
These are the ones you can act on statistically, because they come from the whole population rather than a recruited handful — and they should be the sizing counterpart to lab findings.
Where it goes wrong
- Averaging time on task across successes and failures. A quick abandonment looks like efficiency
- Optimising time when time isn’t the goal. A considered purchase decision taking longer may be better
- Satisfaction without success. People rate attractive interfaces highly while failing tasks on them — Liking
- Metrics without observation. The number says something is wrong; only watching says what
- Tracking what’s easy rather than what matters. Time on task is easy to record and rarely the thing to improve
A workable minimum
PER STUDY
task success (as counts, not %)
SEQ after each task
observed problems, by severity
PER QUARTER
the same tasks, same script
→ a trend
CONTINUOUSLY
funnel completion, error rates,
support volume by topic
The continuous row is what makes the periodic row credible — lab findings that don’t correspond to anything in live behaviour usually mean the tasks weren’t representative.