All articles
Ad Intelligence5 min read

What to do when an ad test has no clear winner

When an ad test has no clear winner, the tie is still information. How to separate an underpowered test from a real one, and what each is worth.

Four creatives, €1,600, two weeks on a cold audience. Best cost per purchase came back at €38, worst at €43, the other two somewhere in between. Forty-one purchases across the whole test. The client call is Thursday and the deck has a slide called "Winner".

I have filled in that slide. The €38 ad became the control for the next round, then the round after that. By month three a chunk of the account rested on a five euro gap nobody could reproduce.

Short answer: An ad test with no clear winner is usually a test whose variants were too similar to separate at the volume you bought. Before you rerun it, work out what size of difference that spend could have detected. Then pick one of three: a bigger creative swing, more volume on the same comparison, or dropping the question.

The takeaways

  • A tie has two causes that look identical on the dashboard. The variants may never have been far enough apart to move the number, or the test may have stopped short of the conversions needed to show a real gap. Only the second is worth rerunning as it stands.
  • Counting noise shrinks with the square root of your conversions. At ten purchases per ad, each cost-per-purchase figure carries a swing of roughly thirty percent. A 13% spread sits comfortably inside that.
  • Promoting a coin flip is expensive later. The accidental winner becomes next round's control, and every brief after it inherits an assumption nobody tested.

Why did the test end in a tie?

Two different things produce the same flat result, and the dashboard prints the same picture for both. Either your variants were never far enough apart to move the number, or the test never accumulated the conversions to resolve a difference that was sitting right there. One is a verdict on the creative. The other is a verdict on the sample size, and the ads are innocent.

Most of the advice written for this moment comes from website testing, where the standard move is to segment the traffic and hunt for the group where the variant won. On a paid social test with forty conversions, that hunt is the most expensive thing you can do. Slice a tie four ways and one slice will always look like a win.

Could the test have separated the variants at all?

Answer this before you look at the ranking again. Conversion counts carry noise that shrinks with the square root of the count, so ten purchases per ad leaves each cost-per-purchase figure with a swing of roughly thirty percent in either direction. A 13% gap between best and worst sits inside the range you would get by running one ad against a copy of itself.

Meta's own learning-phase guidance puts a stable read at around 50 optimization events per week per ad set, and while the learning phase and a readable comparison are separate problems, that number is a useful floor. A test that never came near it did not resolve anything. The arithmetic behind that read is in what statistical significance means for ad tests.

How different were those four ads?

Go back to the briefs before you look at the results again. If all four opened on the same promise with a different crop, a different colour grade and a different stock face, the test asked whether small things matter. At this budget they usually don't, and another two weeks will not change the answer.

This is the part nobody wants to hear on the Thursday call. The fix is a bigger swing, not a longer run. Two ads built on opposed promises, say price-led against risk-led, will separate on far less volume than four cousins of the same idea, because a large question comes with a large answer.

A tie between near-identical creatives is the test telling you it was never worth running.

What does it cost to leave the test open?

More than the media. Every week that ad set stays alive holding an unresolved comparison, the next question does not get asked, and most accounts only have so many clean reads available in a quarter. That queue is the real budget.

So price the answer before you extend. If the winning hook would move maybe €200 of monthly spend, it is not worth €900 of extra media to settle at a confidence you would trust. Put the read somewhere with more room in it.

I have watched an account keep a tie running for six weeks because closing it felt like admitting failure. The six weeks of spend was the cheap part. The four tests queued behind it were not.

Should you just pick the higher number?

Sometimes, yes. If the campaign has to run and one of them has to be in it, pick one and move on. The damage comes later, from what gets written down.

A tie recorded as a win turns into a rule. Six months on, someone tells a new designer that this account performs better with the founder on camera, and the belief traces back to a five euro gap on forty-one purchases. Record it as unresolved, with the spend and the conversion count next to it, so the next person sees an underpowered test rather than a settled question. Keeping that log is the step most accounts skip: how to track creative test results.

So what do you test instead?

Three honest options, and naming one out loud beats leaving the test open. Rerun the same comparison with enough budget to resolve it, if the answer earns that. Swap the variants for two that disagree. Or drop the question and spend the read on a layer you have no evidence about at all.

The thin-data trap is the one part of this I built into the tooling. Ad scores get pulled toward what the format normally does before anything is ranked, so a creative with ten conversions and one lucky week cannot climb over one with a long record. Composite scoring runs on six metrics whose weights you set per project, and the next-test recommendation is reproducible per calendar week. That is the whole read behind ad intelligence.

None of it turns a tie into a winner. Nothing does. The useful thing to say on Thursday is that the comparison did not resolve, which of the three options you picked, and when the next test starts.

This is the thinking behind Adscalr.

See the product