Which ad creative elements drive performance
You cannot A/B test twenty elements. How to read which ad creative elements drive performance across the ads you already ran, without inventing winners.
You cannot A/B test twenty elements. How to read which ad creative elements drive performance across the ads you already ran, without inventing winners.
Three hundred and forty ads over fourteen months, and the question at the quarterly review was fair enough: so what works for us? The answer came out of memory. UGC with a price overlay.
I went back and counted. Nine ads matched. Seven ran in November, on a warm audience, at four times the daily budget of anything else in the account. The overlay was not why they won. It was just present while they did.
Short answer: You cannot A/B test twenty creative elements, so element-level reading has to be retrospective: tag every ad you have already run with its attributes, then compare performance across the whole library. That comparison shows correlation and cannot prove cause, and scanning many attributes at once manufactures winners by chance unless you correct for it.
The takeaways
Because the list is longer than your conversion volume can pay for. Take eight attributes with two settings each and you already have 256 possible ads. Meta's own documentation puts the learning phase at roughly 50 optimisation events per ad set, so a single clean comparison of two variants needs weeks of spend before it means anything.
Run that sequentially through hook type, format, spokesperson, overlay, offer framing, length and colour, and the testing calendar is measured in years. The offer will have changed twice by then.
So the elements never get tested. They get shipped mixed together, inside ads built to win rather than to isolate a variable. That is the right call for the account, and it leaves you with a pile of results and no design behind them. Reading that pile afterwards is a different discipline, with weaker conclusions at the end of it.
The schema, and it has to exist before launch. Decide the fields once: format, hook type, offer framing, spokesperson or none, overlay type, opening shot, dominant colour, length. Then every ad gets tagged the day it ships, by the person who briefed it.
Tag retroactively and you are guessing. Six months later nobody remembers whether the discount sticker was on the 4:5 or only the 9:16, and the ads you liked get generous labels while the flops get filed as "generic". The bias runs one way.
Two constraints. Keep the field values closed ("hook type: question / pain / claim / demo", never free text) or you end up with forty categories holding one ad each. And record campaign, audience and month alongside the attributes, because you will need them later to work out what else changed. Anything you did test deliberately still needs its own result written down.
Because the element was not running alone. In my November example the overlay came bundled with a warm audience, a seasonal buying peak and four times the budget, and any one of those explains the lift more comfortably than a price badge does. Copy the badge onto a cold-traffic ad in February and nothing happens, which is exactly what should happen.
The partial fix is to demand that a pattern repeats across contexts before you believe it. An attribute that only looks good inside one campaign, one audience or one quarter is describing that campaign. Compare within a campaign where you can, so audience and period are held roughly still, and require the same direction across three or four separate campaigns before the attribute earns a name. Even then it stays an association, and retrospective data cannot tell you which way the arrow points.
More than most reports admit. Twenty attributes read against three or four metrics is sixty to eighty comparisons, and at a conventional 95% threshold roughly one in twenty comes back positive with nothing behind it. Three or four confident, entirely fictional winners fall out of a clean dataset before anything real shows up.
This is the multiple-comparisons problem, and it is why a creative-insights dashboard can hand you a fresh "learning" every week. The correction is to control the false discovery rate across the whole family of comparisons instead of judging each on its own, the same machinery that keeps a batch creative test honest.
Without that correction, an element-level report is a ranked list of things that happened, and the top of it is suggestive at best.
A shortlist, ranked by how much spend sits behind it. Retrospective element analysis should produce three or four hypotheses worth a real test next month, phrased as testable statements: "question hooks hold attention longer than claim hooks on cold traffic". A production rule needs more evidence than this can give you.
That reframing changes the job. The analysis stops deciding next quarter's creative direction and starts deciding what you bother testing, which is smaller and far more defensible. Most accounts have eight or ten test slots a year. Filling them from patterns that repeated across your own spend beats filling them from a swipe file.
The tagging half is mechanical, so I automated it. Ads pulled from the three ad libraries get decoded by vision AI into around twenty structured fields plus design, copy and strategy scores: the same schema discipline, at a scale nobody does by hand. On the pattern side, mining across many creatives runs under false-discovery-rate control (Benjamini-Hochberg, q=0.10), because ranking dozens of candidates is where chance winners appear. Thompson sampling then proposes which surviving hypothesis is worth the next test slot. That is the loop ad intelligence runs.
What none of it does is hand you an element-level ROAS report for your own account and call it causal. Nothing can. The output is a shortlist, and the test is still the only thing that settles it.
This is the thinking behind Adscalr.
See the product →