Tags: commerce experimentation concept

Creative Testing

Date: 2026-08-16


Testing the ad rather than the targeting. Since the platforms automated audience selection, creative is the main variable left under your control — and it’s usually tested with none of the discipline applied to on-site experiments.


What it is

Creative testing is systematically comparing ad executions — image, video, copy, format, offer — to find what performs and, ideally, why.

It matters more than it used to because targeting has been automated away. The platforms’ optimisers now decide who sees what; what you supply is the creative and the constraints.

Test the message, not the execution

The distinction that separates a learning programme from a treadmill.

EXECUTION TESTING          MESSAGE TESTING
blue button vs green       "free returns" vs "made in the UK"
image A vs image B         social proof vs product demonstration
                           problem-first vs product-first

teaches: which one won     teaches: which claim resonates
                           → transfers to every future ad,
                             the landing page, and the site

A message finding transfers; an execution finding doesn’t. Testing five images of the same claim tells you which image; testing five claims tells you what the audience cares about, which is reusable across channels and informs Landing Page Strategy and product copy.

Structure tests as concept → variations within the winning concept, not as an undifferentiated pile of assets.

The statistics don’t change

Creative tests are experiments and inherit every requirement, usually unmet:

  • Sample size. Comparing six creatives is five comparisons against control, so significance thresholds need adjusting — The Multiple Comparisons Problem
  • Stopping early. Pausing the “losing” creative after two days is Peeking with a budget attached
  • The platform’s optimiser interferes. It reallocates budget towards early winners, which means later data isn’t a fair comparison — the winner won partly because it got more impressions. This is the specific reason platform creative tests overstate differences
  • Metric choice. Optimising on click-through finds attention-grabbing creative that may not sell — judge on cost per acquisition or contribution, not engagement

Where the platform offers a proper split-test mode that holds allocation constant, use it. Where it doesn’t, treat the results as directional.

Creative fatigue

Real, measurable, and distinct from a bad creative:

week 1    CPA £32     fresh
week 2    CPA £35
week 3    CPA £41     frequency rising
week 4    CPA £48     saturated

Rising cost per acquisition alongside rising frequency is fatigue, not decline in quality. The response is rotation, not optimisation — and it’s why a creative pipeline matters more than any single winner.

Fatigue is also why testing volume is the real constraint. A programme producing four concepts a month outperforms one producing one excellent concept a quarter.

What to record

Creative testing generates the same institutional-memory problem as on-site testing, and usually has no archive at all:

  • The concept and claim, not just the asset ID
  • The audience and placement
  • Cost per acquisition and contribution, not engagement metrics
  • The verdict, including losses. A programme that records only winners can’t compute its win rate — Experiment Archive

Where it connects

  • Paid Social — where creative dominates performance most
  • Landing Page Strategy — the claim that won the click has to be honoured on arrival, or the click is wasted
  • Value Propositions — creative testing is the cheapest way to discover which proposition actually lands, and the findings belong back in your site copy
  • Incrementality Testing — creative tests compare two ads against each other, both subject to the same over-attribution. They tell you which is better, never whether either is incremental