Heuristic Evaluation
Date: 2026-09-28
Experts inspecting an interface against a checklist of principles, with no users involved. Cheap and fast, and it produces hypotheses rather than findings — any one evaluator misses most problems, and some of what they flag no user ever trips over.
Heuristic evaluation is an inspection method in which evaluators examine an interface against a set of usability principles (heuristics) and record each place it violates one. Introduced by Nielsen and Molich (1990); Nielsen’s ten heuristics (revised 1994) are the default set.
What it is not: usability testing. Nobody attempts a task. The output is a prediction of where users will struggle, and it stays a prediction until something else — the funnel, replay, a test session — corroborates it.
The procedure
1 BRIEF the scope (templates, devices), the user, the tasks that matter
↓
2 INSPECT each evaluator alone, two passes:
pass 1 walk the tasks as a user would — flow
pass 2 go screen by screen against the heuristics — detail
↓
3 RECORD one row per issue: where, which heuristic, what happens,
severity — with a screenshot
↓
4 AGGREGATE only now compare lists; merge duplicates, keep every singleton
↓
5 RATE severity agreed across evaluators, not taken from one
Independence is the whole method. Evaluators who see each other’s lists anchor on them, and the aggregate collapses to one person’s view with extra signatures.
Why one evaluator isn’t enough
Averaged across six of Nielsen’s projects, a single evaluator found about 35% of the usability problems; three to five working independently found around 75%. Hence the standard advice: five evaluators, never fewer than three.
The reason is the evaluator effect — different people evaluating the same interface with the same method find markedly different problems. Hertzum and Jacobsen’s review (2001) put agreement between any two evaluators at between 5% and 65%, and found the effect in novices and experts alike, and for severe problems as well as cosmetic ones.
In plain terms: a solo review isn’t a smaller version of a proper one. It’s a sample of about a third of the problems, and you can’t tell which third.
Practical consequence for a CRO audit, where it’s usually one person: treat a solo review as a hypothesis generator, and borrow a second pair of eyes for the critical path — checkout, the product page — even if it’s only an hour.
Severity
Nielsen’s scale, the one most teams use:
| Score | Meaning |
|---|---|
| 0 | Not a usability problem |
| 1 | Cosmetic — fix if time allows |
| 2 | Minor — low priority |
| 3 | Major — important to fix |
| 4 | Catastrophe — must fix before release |
Severity combines frequency (how many users hit it), impact (how hard it is to get past) and persistence (once, or every time). In a commercial audit, the funnel supplies frequency far better than an evaluator’s guess — so rate impact from the inspection, and take frequency from the data.
The heuristic sets
- Nielsen’s ten — general, durable, and so general that almost anything can be filed under “consistency and standards”. Good for coverage, weak for commerce specifics
- Research-derived commerce guidelines — Baymard Institute publishes checkout and product-page guidelines derived from its own usability testing, which makes them closer to evidence than to heuristics [CHECK: current scope and access model of Baymard’s guideline set]
- Persuasion frameworks — CRO-specific checklists covering value proposition, relevance, clarity, anxiety and distraction. They inspect motivation rather than usability, which is the half Nielsen’s set ignores
- Your own — the failure patterns this site has already produced. The best set, once it exists
Mixing a usability set with a persuasion set is the sensible default for a conversion audit: the first finds where people can’t proceed, the second where they don’t want to.
Where it goes wrong
- False positives. Evaluators flag deviations from principle that real users sail past. Without corroboration, a heuristic finding is an opinion with a citation — Triangulation
- Expert blindness. Evaluators know the conventions too well to be confused by them, and miss problems a first-time visitor has. It systematically under-finds comprehension and vocabulary issues, which is exactly what Usability Testing catches
- Checklist tunnel vision. Scoring every screen against every heuristic produces a long list of minor issues and misses the one flow-level problem nobody’s heuristic names. Pass 1 exists to stop this
- Severity by the loudest voice. Rated in a meeting rather than independently, severity follows seniority
- Testing its own conclusions. Evaluator experience is not a substitute for evidence about this site’s users — the same trap as a hypothesis built on a behavioural principle rather than on data — Hypothesis Design
Neighbours
- Cognitive walkthrough — also expert inspection, but narrower: step through one task asking at each step whether the user would know what to do. Better for first-use learnability; worse for coverage
- Usability Testing — the method it approximates. Five sessions find what users actually hit; a heuristic review predicts it for a fraction of the cost
- Accessibility Law in the UK and WCAG — an accessibility audit is a heuristic evaluation with a normative checklist and legal weight, and can’t be skipped on the grounds that the review found nothing