Tags: ux statistics concept

System Usability Scale

Date: 2026-08-17


A ten-question survey producing a single 0–100 usability score. It’s decades old, extensively validated, and its score is not a percentage — which is the misreading that makes a 68 sound like a failure when it’s roughly average.


The System Usability Scale (SUS) is a standardised ten-item questionnaire, answered on a five-point agree/disagree scale, producing one score between 0 and 100.

Its value is standardisation. Because everyone uses the same ten questions, scores are comparable across studies and over time in a way bespoke satisfaction questions never are.

The instrument

Ten statements, alternating positive and negative, each rated 1 (strongly disagree) to 5 (strongly agree):

1  I think I would like to use this
   frequently
2  I found it unnecessarily complex
3  I thought it was easy to use
4  I would need technical support to
   use this
5  The functions were well integrated
6  There was too much inconsistency
7  Most people would learn this very
   quickly
8  It was very cumbersome to use
9  I felt very confident using it
10 I needed to learn a lot before I
   could get going

The alternation is deliberate — it disrupts acquiescence bias, where people agree with everything without reading.

Scoring it

ODD-NUMBERED items    score − 1
EVEN-NUMBERED items   5 − score
sum the 10 adjusted values  (0–40)
multiply by 2.5             (0–100)

Worked example, for a respondent answering 4,2,4,1,4,2,5,2,4,2:

odd  (1,3,5,7,9):   4−1 + 4−1 + 4−1 +
                    5−1 + 4−1  = 15
even (2,4,6,8,10):  5−2 + 5−1 + 5−2 +
                    5−2 + 5−2  = 16
                    ─────────────────
sum                            = 31
× 2.5                          = 77.5

A SUS score of 77.5.

It is not a percentage

The single most common misreading. A score of 68 is not “68% usable” — it’s a point on a 0–100 scale whose distribution is not uniform.

ROUGH INTERPRETATION

  <  50    poor
  ~  68    around average
  >  80    good
  >  85    excellent

Around 68 is the commonly-cited average across many studies, which means a score of 70 is slightly above average rather than a near-miss.

[CHECK: the current published benchmark average and percentile bands before quoting them — these come from specific dataset compilations and are periodically restated.]

What it does and doesn’t tell you

IT TELLS YOU
  a comparable overall impression
  whether a redesign improved or
    worsened perceived usability
  a trend over time

IT DOES NOT TELL YOU
  WHAT is wrong
  which task failed
  whether they could complete anything
  ← it's a thermometer, not a
    diagnosis

Always pair it with task-based observation. A SUS score with no accompanying findings gives you a number to report and nothing to fix — Usability Metrics, Usability Testing.

Practical use

  • Administer immediately after the session, before discussion
  • Don’t change the wording. The validation applies to these items; altering them breaks comparability, which is the whole reason to use it
  • “System” can be swapped for the product name — this is the one accepted modification
  • Ten or more respondents before treating a mean as stable; small samples produce noisy scores
  • Report the mean with the spread, not the mean alone
  • Compare against yourself, over time, with the same tasks — not against a published figure from a different product

Alternatives

SEQ                one question after each
                   task
                   → task-level rather than
                     whole-product
                   → cheaper, and more
                     diagnostic

UMUX-Lite          two items, correlates
                   well with SUS
                   → when ten is too many

NPS                a loyalty question, not
                   a usability one
                   → frequently misused for
                     this purpose
                   — Net Promoter Score

See: Net Promoter Score

SEQ is more useful for iterating; SUS is more useful for tracking. They answer different questions, and running both costs almost nothing extra.

The honest limitation

It measures perception, not performance. Participants can rate an attractive interface highly while failing tasks on it, and the aesthetic-usability effect means the two diverge systematically rather than randomly — Liking.

Which is why the score is a companion to observed task success, never a replacement for it.