Tags: ux concept

Card Sorting and Tree Testing

Date: 2026-08-17


Card sorting asks people how they’d group things; tree testing asks whether they can find something in the structure you built. They answer opposite halves of the same question, and running only the first is how teams ship a taxonomy nobody validated.


Card sorting gives participants a set of items and asks them to group and name the groups. Tree testing gives them a text-only hierarchy and a task, and records where they look.

CARD SORTING        generative
  how would YOU organise this?
  → informs the structure

TREE TESTING        evaluative
  can you FIND this in our structure?
  → validates the structure

Sort first, build, then tree test. Sorting alone tells you what people expect; only tree testing tells you whether what you built works.

Card sorting variants

OPEN      participants create and name
          their own groups
          → best for discovering mental
            models and vocabulary

CLOSED    groups are fixed; participants
          place items into them
          → validates categories you've
            already chosen

HYBRID    fixed groups, plus the ability
          to add their own
          → a reasonable middle

Open sorting is where the vocabulary comes from, and the names people give their groups are often more valuable than the groupings — they’re the words that belong on your navigation — Taxonomy and Labelling.

Reading a card sort

Output is an agreement matrix — how often each pair of items was grouped together.

                Serum  Cleanser  Brush  Mirror
Serum             —      87%      12%     4%
Cleanser         87%      —       9%      6%
Brush            12%      9%      —      71%
Mirror            4%      6%     71%      —

two clear clusters, and no ambiguity
about which items belong where

Look for the items with no strong home. An item sorted into four different groups by four people is the finding — it means the item’s purpose is unclear, or it genuinely belongs in two places, and both need a decision.

This is the one qualitative-feeling method that is analysed statistically, via clustering across participants, so it needs enough people for the agreement patterns to stabilise — 15–30 is the usual range — Sample Size in Qualitative Research.

Tree testing

Participants see the hierarchy as plain text — no visual design, no search, no images — and are given a task.

TASK  "Where would you find a
       moisturiser for sensitive skin?"

Home
├─ Skincare
│   ├─ Cleansers
│   ├─ Moisturisers      ← correct
│   └─ Treatments
├─ Make-up
└─ Tools

Stripping the visuals is the point. It isolates the structure and the labels from the design, so a failure is unambiguously a taxonomy problem rather than a layout one.

Metrics:

SUCCESS       found the right place
DIRECTNESS    got there without backtracking
              ← the more useful of the two
TIME          how long
FIRST CLICK   where they went first

Directness matters more than success. Someone who eventually finds it after three wrong turns will not do that on a live site — they’ll leave — Search and Findability.

What each is bad at

  • Card sorting doesn’t produce your navigation. It produces evidence about mental models. Merchandising priorities, SEO, and commercial structure all legitimately override it
  • It ignores frequency. Participants treat every card as equally important; your traffic doesn’t. A category with 40% of revenue deserves prominence a card sort won’t suggest
  • Tree testing can’t tell you what’s missing. It only tests the structure you gave it
  • Neither sees the real page. Filters, search, images and merchandising all change findability once the design exists — Faceted Filtering

Practical notes

  • Use real item names, not sanitised ones. “Hyaluronic Acid Serum 30ml” sorts differently from “Serum”
  • Cap the deck at ~40 cards. Beyond that, fatigue produces arbitrary grouping
  • Run tree testing unmoderated and at scale — it’s cheap, fast, and produces comparable numbers across variants — Moderated vs Unmoderated Testing
  • Test the alternatives against each other, not just the current structure. Tree testing two candidate hierarchies is the cheapest way to settle a navigation argument
  • Retest after launch. Real search and browse behaviour will show whether the structure held — Information Architecture