Card Sorting and Tree Testing
Date: 2026-08-17
Card sorting asks people how they’d group things; tree testing asks whether they can find something in the structure you built. They answer opposite halves of the same question, and running only the first is how teams ship a taxonomy nobody validated.
Card sorting gives participants a set of items and asks them to group and name the groups. Tree testing gives them a text-only hierarchy and a task, and records where they look.
CARD SORTING generative
how would YOU organise this?
→ informs the structure
TREE TESTING evaluative
can you FIND this in our structure?
→ validates the structure
Sort first, build, then tree test. Sorting alone tells you what people expect; only tree testing tells you whether what you built works.
Card sorting variants
OPEN participants create and name
their own groups
→ best for discovering mental
models and vocabulary
CLOSED groups are fixed; participants
place items into them
→ validates categories you've
already chosen
HYBRID fixed groups, plus the ability
to add their own
→ a reasonable middle
Open sorting is where the vocabulary comes from, and the names people give their groups are often more valuable than the groupings — they’re the words that belong on your navigation — Taxonomy and Labelling.
Reading a card sort
Output is an agreement matrix — how often each pair of items was grouped together.
Serum Cleanser Brush Mirror
Serum — 87% 12% 4%
Cleanser 87% — 9% 6%
Brush 12% 9% — 71%
Mirror 4% 6% 71% —
two clear clusters, and no ambiguity
about which items belong where
Look for the items with no strong home. An item sorted into four different groups by four people is the finding — it means the item’s purpose is unclear, or it genuinely belongs in two places, and both need a decision.
This is the one qualitative-feeling method that is analysed statistically, via clustering across participants, so it needs enough people for the agreement patterns to stabilise — 15–30 is the usual range — Sample Size in Qualitative Research.
Tree testing
Participants see the hierarchy as plain text — no visual design, no search, no images — and are given a task.
TASK "Where would you find a
moisturiser for sensitive skin?"
Home
├─ Skincare
│ ├─ Cleansers
│ ├─ Moisturisers ← correct
│ └─ Treatments
├─ Make-up
└─ Tools
Stripping the visuals is the point. It isolates the structure and the labels from the design, so a failure is unambiguously a taxonomy problem rather than a layout one.
Metrics:
SUCCESS found the right place
DIRECTNESS got there without backtracking
← the more useful of the two
TIME how long
FIRST CLICK where they went first
Directness matters more than success. Someone who eventually finds it after three wrong turns will not do that on a live site — they’ll leave — Search and Findability.
What each is bad at
- Card sorting doesn’t produce your navigation. It produces evidence about mental models. Merchandising priorities, SEO, and commercial structure all legitimately override it
- It ignores frequency. Participants treat every card as equally important; your traffic doesn’t. A category with 40% of revenue deserves prominence a card sort won’t suggest
- Tree testing can’t tell you what’s missing. It only tests the structure you gave it
- Neither sees the real page. Filters, search, images and merchandising all change findability once the design exists — Faceted Filtering
Practical notes
- Use real item names, not sanitised ones. “Hyaluronic Acid Serum 30ml” sorts differently from “Serum”
- Cap the deck at ~40 cards. Beyond that, fatigue produces arbitrary grouping
- Run tree testing unmoderated and at scale — it’s cheap, fast, and produces comparable numbers across variants — Moderated vs Unmoderated Testing
- Test the alternatives against each other, not just the current structure. Tree testing two candidate hierarchies is the cheapest way to settle a navigation argument
- Retest after launch. Real search and browse behaviour will show whether the structure held — Information Architecture