Tags: commerce experimentation concept
Creative Testing
Date: 2026-08-16
Testing the ad rather than the targeting. Since the platforms automated audience selection, creative is the main variable left under your control — and it’s usually tested with none of the discipline applied to on-site experiments.
What it is
Creative testing is systematically comparing ad executions — image, video, copy, format, offer — to find what performs and, ideally, why.
It matters more than it used to because targeting has been automated away. The platforms’ optimisers now decide who sees what; what you supply is the creative and the constraints.
Test the message, not the execution
The distinction that separates a learning programme from a treadmill.
EXECUTION TESTING MESSAGE TESTING
blue button vs green "free returns" vs "made in the UK"
image A vs image B social proof vs product demonstration
problem-first vs product-first
teaches: which one won teaches: which claim resonates
→ transfers to every future ad,
the landing page, and the site
A message finding transfers; an execution finding doesn’t. Testing five images of the same claim tells you which image; testing five claims tells you what the audience cares about, which is reusable across channels and informs Landing Page Strategy and product copy.
Structure tests as concept → variations within the winning concept, not as an undifferentiated pile of assets.
The statistics don’t change
Creative tests are experiments and inherit every requirement, usually unmet:
- Sample size. Comparing six creatives is five comparisons against control, so significance thresholds need adjusting — The Multiple Comparisons Problem
- Stopping early. Pausing the “losing” creative after two days is Peeking with a budget attached
- The platform’s optimiser interferes. It reallocates budget towards early winners, which means later data isn’t a fair comparison — the winner won partly because it got more impressions. This is the specific reason platform creative tests overstate differences
- Metric choice. Optimising on click-through finds attention-grabbing creative that may not sell — judge on cost per acquisition or contribution, not engagement
Where the platform offers a proper split-test mode that holds allocation constant, use it. Where it doesn’t, treat the results as directional.
Creative fatigue
Real, measurable, and distinct from a bad creative:
week 1 CPA £32 fresh
week 2 CPA £35
week 3 CPA £41 frequency rising
week 4 CPA £48 saturated
Rising cost per acquisition alongside rising frequency is fatigue, not decline in quality. The response is rotation, not optimisation — and it’s why a creative pipeline matters more than any single winner.
Fatigue is also why testing volume is the real constraint. A programme producing four concepts a month outperforms one producing one excellent concept a quarter.
What to record
Creative testing generates the same institutional-memory problem as on-site testing, and usually has no archive at all:
- The concept and claim, not just the asset ID
- The audience and placement
- Cost per acquisition and contribution, not engagement metrics
- The verdict, including losses. A programme that records only winners can’t compute its win rate — Experiment Archive
Where it connects
- Paid Social — where creative dominates performance most
- Landing Page Strategy — the claim that won the click has to be honoured on arrival, or the click is wasted
- Value Propositions — creative testing is the cheapest way to discover which proposition actually lands, and the findings belong back in your site copy
- Incrementality Testing — creative tests compare two ads against each other, both subject to the same over-attribution. They tell you which is better, never whether either is incremental