Incrementality testing for Facebook ads
Incrementality testing for Facebook ads on a normal budget: three holdout designs you can afford, how to size one, and what the read cannot tell you.
Incrementality testing for Facebook ads on a normal budget: three holdout designs you can afford, how to size one, and what the read cannot tell you.
The account looks fine. Ads Manager reports 4.1 on the month, the creative tests are behaving, frequency is under control. Then somebody asks what happens to revenue if the channel goes dark for two weeks, and the only answer available comes from the same dashboard.
That is the thread shape I keep reading. Nobody posts that their ROAS fell. They post that they no longer believe the number, and some of them have run paid social for a decade.
Short answer: Incrementality testing for Facebook ads means withholding spend from part of your market and comparing your own total revenue in the exposed and withheld halves. The platform cannot run that comparison for you, because every conversion it reports is one it already decided to claim credit for. Sizing the test honestly is the hard part.
The takeaways
Because attribution assigns credit, and credit is a weaker claim than cause. Every conversion in the report is a purchase that happened. The platform decides which of those purchases it touched on the way, and a buyer who was going to order anyway looks identical to one the ad created. Nothing in the interface separates them.
That is a different problem from the count being wrong, which it also is, in both directions at once. Signal loss removes conversions the platform never sees, modelling adds estimates back, and the two errors do not cancel on any schedule you can predict. I went through that mismatch in why Meta conversions don't match your sales.
Incrementality is the harder question underneath it. A perfectly counted conversion still tells you nothing about whether the ad caused it. The only way to find out is to withhold the ad from somebody and watch what they do.
Comparable regions, and enough of them. You switch ads off in one set of regions, leave them running in a matched set, and compare total revenue between the two groups over a fixed window. That holds only if those regions were already tracking each other before you touched anything.
So build the split from your own pre-period data. Pull weekly revenue by region for the two months before the test, pair regions whose lines moved together, then check the two combined groups sit on top of each other. A map is not a matching method. Two cities of similar size can have completely different seasonality.
Then freeze everything else for the window: no creative refresh, no budget change, no promo in the withheld regions. Every guide I read while writing this quietly assumes a channel spending five figures a month, and that assumption does more work than the methodology sections admit.
Run the channel on and off in time instead of in space. Pick a stretch where your spend line has been boring for a month, turn Meta off for a defined block, and compare total revenue against the blocks either side. Cheapest design to execute, easiest to misread, because your revenue moves with the calendar anyway.
Bracket it. An off block sandwiched between two on blocks beats a single before-and-after, and alternating two weeks on, two weeks off beats that. Never run it across a payday, a holiday, or a promotion somebody scheduled three months ago.
A third option is a step change instead of a switch: halve the budget in one matched half of the market, hold it in the other. That measures the margin. It answers whether your last increment of spend is doing anything and says less about the channel as a whole.
Big enough in orders, and the calendar has little to do with it. If the withheld side of your split produces a dozen purchases across the window, a 20% difference between the halves is invisible inside ordinary week-to-week noise. Four more weeks at that volume does not repair it. More events per week would.
This is the arithmetic behind every ad test, moved up to channel level. I worked through it for creative comparisons in statistical significance for Facebook ads, and the conclusion travels: resolution comes from events.
Meta's own Conversion Lift study exists for this and carries an eligibility floor, which plenty of accounts discover by being told they sit under it. Size the test yourself first. If your account can only ask whether the channel does anything at all, ask that instead of faking precision.
It costs the revenue the withheld group would have made, so price that before you start. A two-week regional holdout on a healthy channel is a deliberate hole in the quarter. Decide in advance which decision the result changes.
What you get back is channel-level: whether Meta produced incremental revenue at your current spend, in this season, at this offer. It cannot tell you which campaign, audience or creative did the work. It also has an expiry date, because changing your spend materially retires the answer you bought.
Say both possible results out loud before you run it. If a flat read would not change anything you do on Monday, the test is expensive theatre.
Adscalr does not run incrementality tests, and I would rather write that plainly than let the word do sales work. A holdout is manual, occasional and expensive. It belongs in your calendar.
What the ad intelligence side does is narrower: a composite score from six metrics including ROAS and revenue per install, with Bayesian shrinkage against format-specific priors, so a new ad's wild first days get pulled toward what that format normally does. That stops a lucky week getting promoted. It says nothing about whether the channel is incremental, and nothing reading platform data can.
Run the holdout once a year. Believe the dashboard a little less in between.
This is the thinking behind Adscalr.
See the product →