Tags: ux concept

Qualitative Coding

Date: 2026-08-17


Systematically labelling qualitative data so patterns can be identified rather than remembered. Without it, analysis is whoever ran the sessions recalling what struck them — which reliably over-weights the vivid and the recent.


Qualitative coding assigns labels to segments of qualitative data — interview transcripts, open survey text, diary entries, support tickets — so themes can be identified across the whole set rather than from impression.

The problem it solves is memory. After eight interviews, the researcher recalls the articulate participant, the one who was angry, and the last one — which is not the same as what most people said.

Two approaches

INDUCTIVE (open)
  codes emerge from the data
  → read, label what's there, group
    the labels
  → for exploratory work, and for
    finding what you didn't expect

DEDUCTIVE (closed)
  a framework decided in advance
  → faster, comparable across studies
  → risks confirming what you expected
    and missing the rest

HYBRID
  a starting framework, extended as
  new things appear
  ← usually the practical choice

The process

1  FAMILIARISE      read everything once,
                    without coding

2  OPEN CODE        label segments
                    → "expected free
                      delivery"
                    → "didn't see the
                      filter"

3  GROUP            cluster codes into
                    themes
                    → "cost surprises"
                    → "findability"

4  REVIEW           do the themes hold
                    against the raw data?
                    → this step catches
                      the theme you wanted
                      to find

5  DEFINE           name each theme, state
                    what it does and
                    doesn't cover

6  REPORT           with evidence attached
                    — Research Repositories

See: Research Repositories

Step 1 matters more than it sounds. Coding while first reading anchors the scheme on the first transcript, and everything after gets forced into it.

Codes describe, themes interpret

The distinction that keeps the output honest:

CODE     "said delivery cost was a
          surprise"
         ← descriptive, close to the data

THEME    "cost transparency is the
          primary trust failure in
          checkout"
         ← interpretive, a claim about
           what it means

Keep codes close to what was said. A code that already contains the conclusion means the analysis happened before the coding.

Counting, carefully

Coding produces counts, and the counts are not rates.

LEGITIMATE
  "6 of 8 participants mentioned
   delivery cost"
  → tells you it recurred, not
    coincidence

NOT LEGITIMATE
  "75% of customers are concerned
   about delivery cost"
  → 8 non-randomly-recruited people
    estimate nothing about your
    customers

Report counts against the sample, never as percentages of a population — Sample Size in Qualitative Research, Qualitative vs Quantitative Research.

Frequency isn’t importance either. One participant blocked entirely matters more than five mildly irritated — severity and frequency are separate axes and should be reported separately.

Reliability

SINGLE CODER      fast, and carries that
                  person's interpretation

TWO CODERS        code a subset
                  independently, compare,
                  reconcile the scheme
                  → catches idiosyncratic
                    reading
                  ← worth it for anything
                    consequential

INTERCODER        a formal agreement
AGREEMENT         statistic
                  → academic contexts;
                    rarely proportionate
                    commercially

Two people coding the same two transcripts, then comparing, is the pragmatic version — it takes an hour and catches most of the drift.

Where it goes wrong

  • Cherry-picking the vivid quote. The most quotable participant is not the most representative
  • Coding to confirm. A scheme built around the hypothesis will find it
  • Too many codes. Sixty codes across eight interviews is a transcript with labels, not an analysis
  • Losing the link to evidence. A theme with no traceable quotes is an assertion — Research Repositories
  • Skipping it entirely, which is the most common failure. “I watched all the sessions and here’s what I think” is memory, presented as analysis

Beyond research sessions

The same technique applies to data you already have, in volume:

SUPPORT TICKETS       code by reason →
                      a ranked list of
                      what's broken
PRODUCT Q&A           objections your copy
                      doesn't answer
REVIEWS               vocabulary, and
                      recurring complaints
SURVEY OPEN TEXT      — Surveys
EXIT SURVEY RESPONSES — Voice of Customer Data

See: Surveys · Voice of Customer Data

Support tickets are the highest-volume qualitative dataset most retailers already own and never code, and unlike research they arrive continuously and for free.