Tags: ux statistics concept

Usability Metrics

Date: 2026-08-17


Numbers describing how well people can use something. Most are collected from small non-random samples, which makes them directional rather than measurements — and treating a task success rate from eight people as a percentage of your customers is the standard error.


Usability metrics quantify effectiveness, efficiency and satisfaction — the three dimensions in the standard definition of usability.

EFFECTIVENESS   can they complete it?
                → task success rate,
                  error rate

EFFICIENCY      at what cost?
                → time on task, steps,
                  errors per task

SATISFACTION    how did it feel?
                → SUS, ease-of-task
                  ratings
                — System Usability Scale

The core measures

TASK SUCCESS       completed / attempted
                   → the primary measure
                   → define "success"
                     before testing

TIME ON TASK       for successful attempts
                   only — a fast failure
                   isn't fast

ERROR RATE         errors per task, and
                   recovery rate

EFFICIENCY         steps taken vs the
                   minimum

SEQ                Single Ease Question:
                   "how difficult was
                   that task?" 1–7, asked
                   immediately after each
                   task
                   ← cheap, and correlates
                     well with success

SEQ is the highest value-per-effort metric here — one question, asked after each task, and it surfaces tasks people completed but found hard, which task success alone hides.

The sample size problem

This is the part that gets misused.

8 participants, 5 succeeded

REPORTED AS   "62.5% task success rate"

ACTUALLY      5 of 8 non-randomly
              recruited people succeeded

              the confidence interval
              around that proportion is
              enormous, and the sample
              wasn't drawn from your
              customers anyway

A proportion from a small qualitative sample is not an estimate of anything. Report the raw counts — “5 of 8 completed the task” — which is honest and just as actionable — Sample Size in Qualitative Research, Qualitative vs Quantitative Research.

In plain terms: the number tells you a problem exists. It does not tell you how many customers hit it. Sizing that is analytics’ job — Funnel Analysis.

Benchmarking is where it earns its keep

Usability metrics are far more useful as comparisons than as absolutes:

COMPARE
  before vs after a redesign
  variant A vs variant B
  your site vs a competitor
  this quarter vs last, same tasks

DON'T COMPARE
  your 68 SUS against another
  company's published 74
  → different tasks, participants,
    context

Same tasks, same script, same recruitment criteria is what makes a comparison meaningful. Change any of them and the difference is uninterpretable.

Behavioural metrics from live data

Distinct from lab metrics, and genuinely quantitative because the sample is everyone:

  • Task completion — funnel conversion — Funnel Analysis
  • Error rates — form field errors — Form Analytics
  • Rage clicks — repeated clicking on something unresponsive — Session Replay
  • Search refinement — the first attempt failed
  • Back-button use — from product to category
  • Support contacts, by topic

These are the ones you can act on statistically, because they come from the whole population rather than a recruited handful — and they should be the sizing counterpart to lab findings.

Where it goes wrong

  • Averaging time on task across successes and failures. A quick abandonment looks like efficiency
  • Optimising time when time isn’t the goal. A considered purchase decision taking longer may be better
  • Satisfaction without success. People rate attractive interfaces highly while failing tasks on them — Liking
  • Metrics without observation. The number says something is wrong; only watching says what
  • Tracking what’s easy rather than what matters. Time on task is easy to record and rarely the thing to improve

A workable minimum

PER STUDY
  task success (as counts, not %)
  SEQ after each task
  observed problems, by severity

PER QUARTER
  the same tasks, same script
  → a trend

CONTINUOUSLY
  funnel completion, error rates,
  support volume by topic

The continuous row is what makes the periodic row credible — lab findings that don’t correspond to anything in live behaviour usually mean the tasks weren’t representative.