
Creative Testing
Part of Creative testing programmes
Choosing a stable comparison for a creative experiment
Choose a relevant control, keep campaign conditions steady and record changes that limit a creative experiment.
A stable comparison is an approved creative that can run alongside the proposed version for the same intended audience and outcome, with other material settings held as steady as the platform allows. Choose it for relevance to the decision, not because it had the best result in an unrelated period.
Select a control that answers the question
If the decision is whether a new headline improves an existing campaign, use the current eligible ad with its existing headline as the control. Keep the offer, destination, format and conversion action the same. If the current ad has a broken page or an expired offer, repair that problem before the experiment; an invalid baseline will not provide a useful comparison.
Record the exact control asset and proposed change. A historical benchmark can provide context, but it is not a control running at the same time: demand, media prices, targeting and the offer may have shifted since it ran. Where a live split is unavailable, describe the exercise as a comparison over time and keep the resulting claim modest.
Control vs. Treatment: Key differences to maintain for valid results
- Control (current ad)
- Existing headline, offer, destination, format, conversion action
- Treatment (proposed ad)
- Only the new headline tested; all other elements unchanged
- Invalid baseline example
- Ad with broken landing page or expired offer
Hold campaign conditions steady
Before launch, check audience eligibility, geography, placements, schedule, budget approach, bidding, destination, tracking and approval status. Set the success measure and observation window in advance. Check that both arms serve, then inspect delivery rather than assuming a planned split produced equal spend or impressions.
The split method matters, so check how the chosen experiment type assigns users before relying on the comparison.
Changing a campaign while an experiment is running can make results harder to interpret. Experiment setups vary, so check the chosen type rather than assuming one procedure applies to all.
Keep a run log of budget edits, targeting changes, disapprovals, stock changes, landing-page revisions and tracking interruptions, with dates and affected arms. If a change could explain the observed gap, the test may need to be repeated. A favourable result does not erase a disturbance.
State what the comparison supports
A well-run creative split supports a claim about the compared versions, for the tested audience, placement and period, using the chosen outcome. It does not automatically establish a universal winning message. If the control and treatment differ in copy, offer and landing page, call the result a package comparison; the effect of each component remains unknown.
Finish with a decision statement: keep the control, adopt the treatment, run a more precise follow-up, or treat the result as inconclusive. Give the reason and note any delivery imbalance. That record makes the control useful to the next test as well as the current one.
Critical factors affecting experiment validity
- Valid control must run simultaneously
- Not a historical benchmark
- Split method affects result interpretation
- Check user assignment before relying on data
- Favourable result does not override disturbances
- Changes during test may invalidate findings



