Qualitative Coding
Date: 2026-08-17
Systematically labelling qualitative data so patterns can be identified rather than remembered. Without it, analysis is whoever ran the sessions recalling what struck them — which reliably over-weights the vivid and the recent.
Qualitative coding assigns labels to segments of qualitative data — interview transcripts, open survey text, diary entries, support tickets — so themes can be identified across the whole set rather than from impression.
The problem it solves is memory. After eight interviews, the researcher recalls the articulate participant, the one who was angry, and the last one — which is not the same as what most people said.
Two approaches
INDUCTIVE (open)
codes emerge from the data
→ read, label what's there, group
the labels
→ for exploratory work, and for
finding what you didn't expect
DEDUCTIVE (closed)
a framework decided in advance
→ faster, comparable across studies
→ risks confirming what you expected
and missing the rest
HYBRID
a starting framework, extended as
new things appear
← usually the practical choice
The process
1 FAMILIARISE read everything once,
without coding
2 OPEN CODE label segments
→ "expected free
delivery"
→ "didn't see the
filter"
3 GROUP cluster codes into
themes
→ "cost surprises"
→ "findability"
4 REVIEW do the themes hold
against the raw data?
→ this step catches
the theme you wanted
to find
5 DEFINE name each theme, state
what it does and
doesn't cover
6 REPORT with evidence attached
— Research Repositories
Step 1 matters more than it sounds. Coding while first reading anchors the scheme on the first transcript, and everything after gets forced into it.
Codes describe, themes interpret
The distinction that keeps the output honest:
CODE "said delivery cost was a
surprise"
← descriptive, close to the data
THEME "cost transparency is the
primary trust failure in
checkout"
← interpretive, a claim about
what it means
Keep codes close to what was said. A code that already contains the conclusion means the analysis happened before the coding.
Counting, carefully
Coding produces counts, and the counts are not rates.
LEGITIMATE
"6 of 8 participants mentioned
delivery cost"
→ tells you it recurred, not
coincidence
NOT LEGITIMATE
"75% of customers are concerned
about delivery cost"
→ 8 non-randomly-recruited people
estimate nothing about your
customers
Report counts against the sample, never as percentages of a population — Sample Size in Qualitative Research, Qualitative vs Quantitative Research.
Frequency isn’t importance either. One participant blocked entirely matters more than five mildly irritated — severity and frequency are separate axes and should be reported separately.
Reliability
SINGLE CODER fast, and carries that
person's interpretation
TWO CODERS code a subset
independently, compare,
reconcile the scheme
→ catches idiosyncratic
reading
← worth it for anything
consequential
INTERCODER a formal agreement
AGREEMENT statistic
→ academic contexts;
rarely proportionate
commercially
Two people coding the same two transcripts, then comparing, is the pragmatic version — it takes an hour and catches most of the drift.
Where it goes wrong
- Cherry-picking the vivid quote. The most quotable participant is not the most representative
- Coding to confirm. A scheme built around the hypothesis will find it
- Too many codes. Sixty codes across eight interviews is a transcript with labels, not an analysis
- Losing the link to evidence. A theme with no traceable quotes is an assertion — Research Repositories
- Skipping it entirely, which is the most common failure. “I watched all the sessions and here’s what I think” is memory, presented as analysis
Beyond research sessions
The same technique applies to data you already have, in volume:
SUPPORT TICKETS code by reason →
a ranked list of
what's broken
PRODUCT Q&A objections your copy
doesn't answer
REVIEWS vocabulary, and
recurring complaints
SURVEY OPEN TEXT — Surveys
EXIT SURVEY RESPONSES — Voice of Customer Data
See: Surveys · Voice of Customer Data
Support tickets are the highest-volume qualitative dataset most retailers already own and never code, and unlike research they arrive continuously and for free.