All articles
Ad Intelligence6 min read

How to track ad creative test results

Every test produces a winner and no memory. How to track ad creative test results so month six does not start where month one did.

60-something creatives shipped since January. The account is doing fine. Then somebody new joins, asks what we have learned so far, and the room produces two answers that contradict each other: UGC beats studio, and UGC only works with the founder on camera. Nobody can point at a number for either.

The evidence exists. It is sitting in Ads Manager under names like hook_v2_FINAL and buyback_9x16_final_v3, which is another way of saying it does not exist.

Short answer: Tracking ad creative test results starts before launch. Decide the fields that describe every creative (angle, format, hook type, proof type, offer framing, who is on camera), store them as columns instead of inside filenames, then log each result against them. One test answers one comparison. 40 tagged creatives let you ask which attribute keeps winning.

The takeaways

  • Tags chosen after the result are rationalisations. Once you know the buyback ad won, "bold offer" writes itself. Fix the field list while you still have no idea which row wins.
  • 40 creatives across 8 attributes is dozens of comparisons at once. At a 5% threshold, roughly 1 in 20 clears on chance alone, so the library hands you a confident house rule per quarter whether or not one exists.
  • The record is what makes the next test cheaper. One test buys one comparison. A log is the only thing that turns 12 months of single comparisons into a question worth asking.

Why does month six start where month one did?

Because the finding lives in one person's head and the filename carries none of it. A creative test produces two artifacts: a number in Ads Manager, and an explanation somebody gives in a Slack thread on the Friday. The number persists. The explanation is gone in a fortnight, and it leaves with whoever ran the test.

So the account keeps buying the same information twice. I have re-run a hook style in September that I had already killed in April, and only caught it because an old export happened to be open in another tab.

None of that is carelessness inside the test. Each round was designed and read properly, the way a clean creative test is supposed to work. The failure sits between tests, in the part nobody owns.

What should you record about each creative?

The attributes it is made of, and the comparison it sat in. Angle, format, hook type, proof type, offer framing, who is on camera, whether a price appears on screen. One column each. Then the outcome: spend, the metric you judged on, and which test the creative belonged to.

Teams try to store this in the filename, and it fails for a boring reason: you cannot pivot a string. ugc_founder_pricehook_4x5_v2 contains the information and gives you no way to ask it a question. Move the same six facts into Google Sheets columns and one filter answers "every creative with a price on screen".

Keep the schema small. 6 to 8 fields you will fill in every time beat 20 you abandon in week three. Write the permitted values down as a closed list too, so "founder", "Founder" and "owner on camera" do not quietly become three different attributes.

Why do the tags have to be decided before launch?

Because after the result lands you are describing the winner, not predicting it. Once the buyback ad has won, the story assembles itself: bold offer, direct hook, no music. Each label is true. None of them were selected in advance, so none of them count as evidence.

Choose the fields while every row is still equally likely to win. It is the cheapest piece of rigour in a creative program and it costs one meeting.

This is also where the current wave of AI tagging tools works against you. They watch every video and assign tags across dozens of dimensions after the fact, which sounds like leverage and mostly buys extra ways to find a pattern in noise. More dimensions, more comparisons, more coincidences dressed up as insight. Pick your fields first, then let a tool fill them in.

When does a tagged library start lying to you?

The moment you mine it. Take 40 creatives and 8 attributes with 3 values each: that is 24 groups you can hold against the rest, and at a 5% threshold roughly 1 in 20 comparisons clears the bar on chance alone. A library that size produces a confident-looking house rule every quarter, real or not.

The fake rule always sounds good. "Statics with a price on screen beat everything." It gets promoted to a brief, and the next 6 creatives are built on a coin flip.

Two things fix it. Correct for how many comparisons you ran, which is what false-discovery-rate control does. And treat whatever survives as a hypothesis, so the pattern earns its promotion by winning a fresh test built to check it. I went through what makes a difference real separately.

Adscalr does the first of those inside its pattern mining: Benjamini-Hochberg at q=0.10, applied when ranking many creatives at once, precisely because attribute mining generates candidate rules faster than an account generates evidence.

Can a log separate the attribute from the offer?

No. And that is worth saying before someone builds a quarter on it. Suppose your UGC creatives keep winning. Maybe UGC works. Or you handed UGC your strongest offer, because that is what everyone does with the format they believe in, and the log stores both facts in separate columns without knowing which one moved the number.

No amount of tagging breaks that tie. A test does: same offer, both formats, everything else held still.

A second limit is worth naming. A new creative's early score is noisy, and feeding raw numbers into the library teaches it that unusual formats are brilliant. Adscalr's scoring pulls a new ad toward what its format normally does through Bayesian shrinkage with format-specific priors, so one lucky week never becomes a permanent entry.

Turning the log into the next test

The point of a record is that it changes what you launch. A well-kept sheet turns "what should we test next" from an opinion into a shortlist: the attribute with the widest spread and the thinnest evidence goes first.

Two honest notes. Adscalr scores creatives on a blend of six metrics whose weights you set per project and funnel stage, then recommends the next test worth running through Thompson sampling, reproducible per calendar week. It does not tag your uploads: the vision model that decodes ads into roughly 20 structured fields runs on competitor ads pulled from the three ad libraries. The ad-intelligence page covers how that scoring holds up on small samples.

Your schema stays yours either way. 6 columns in a sheet, filled in before launch, beat any amount of labelling done afterwards.

This is the thinking behind Adscalr.

See the product