Tags: analytics commerce concept
Self-Reported Attribution
Date: 2026-09-27
Asking customers how they heard about you. It sees the channels no pixel can — podcasts, word of mouth, a video they watched on another device weeks ago — and it’s biased in its own ways. Its value is as a second opinion that disagrees with click-based attribution in informative places, not as the answer.
Self-reported attribution is crediting marketing channels from customers’ own answers to a question such as “How did you hear about us?”, usually asked at checkout or sign-up.
What each method sees
| Channel | Click-based attribution | Self-reported |
|---|---|---|
| Paid search | Sees it well — and over-credits it | Under-reported; people remember the reason they searched, not the ad |
| Paid social | Partly — view-throughs, cross-device gaps | Reported often — “Instagram” |
| Podcasts, radio, TV | Barely | Well |
| Word of mouth | Not at all — shows as direct or brand search | Well |
| Influencers | Only with a code or tracked link | Well |
| Well | Rarely mentioned — it’s not where they first heard |
The two methods are wrong in different directions. That’s why they’re worth comparing — Attribution Models, Direct Traffic and Lost Referrers.
Clicks answer “what did they touch last?” The survey answers “what do they remember?” Neither is “what caused it” — Incrementality Testing.
Its biases
- Recall. People remember vivid and recent sources, and the first time they noticed you, which may not have been the first time they saw you
- Option order and wording. The first options in a list get picked more. Randomise order, and keep the list short
- “Google” means anything. Most people who answer “Google” searched after hearing about you somewhere else
- Who answers. If the question is optional, respondents may differ from non-respondents
- Satisficing. Picking anything to get past the question — Surveys
Reading the numbers
Worked example. 1,200 customers answered this month; 180 said “podcast”.
share p = 180 ÷ 1,200 = 15.0%
standard error √(p(1−p) ÷ n)
√(0.15 × 0.85 ÷ 1,200) = 1.03 percentage points
95% interval 15.0% ± 1.96 × 1.03 = 13.0% to 17.0%
In plain terms: with 1,200 answers, “15% from podcasts” is accurate to about two points either way as a measure of what these customers said — Confidence Intervals.
What the interval doesn’t cover is the bias. The ±2 points is sampling error only. If podcast listeners are more likely to answer the question, the true share could be well outside that range. More responses shrink the interval; they don’t fix a biased question.
Watch trends, not levels. The absolute share is distorted by the biases above; the change from month to month, with the question unchanged, is much more reliable. A podcast share rising from 5% to 15% after a sponsorship is informative even if neither figure is exactly right.
Designing the question
- Single choice, randomised order, 6–10 options plus “Other (please say)”
- Name specific sources where you spend — “a podcast”, “a friend or family member”, “TikTok” — not “social media”
- Ask after the purchase, on the confirmation page or in the post-purchase email — never as a checkout field, where it costs conversions — Form Design
- Don’t change the options often — every change breaks the trend line — Annotation and Change Logs
- Code the free-text “Other” answers regularly; new channels show up there first — Qualitative Coding
Using it
- As a check on the attribution model. Where the survey says 15% podcast and the click model says 0%, the click model is missing something
- To find demand creation that gets credited to search — Demand Creation vs Demand Capture
- As one input alongside others, not a replacement — Marketing Mix Modelling, Geo Holdout Tests
- For lead-gen, a “how did you hear” field on the enquiry form, stored in the CRM, joins to deal value — Offline Conversion Imports