All articles
Ad Intelligence6 min read

Which catalog ad overlay drives the best ROAS?

No report breaks catalog ad ROAS down by overlay type, and here is the structural reason plus the comparison you can build yourself instead.

Fourteen hundred SKUs, three overlay templates in the render tool, and a client asking which one makes the most money. You open Ads Manager and there is nowhere to look. Age, placement, platform, region. Nothing that says price badge or rating star. Every page a search returns is a vendor page ending in a ROAS lift from somebody else's account.

Short answer: No report can break catalog ad ROAS down by overlay type, because the overlay is a render-time template binding rather than an ad entity, so nothing in the account ever recorded which label an impression carried. To compare overlays you have to build the split into your account structure first, and read it as a setup comparison.

The takeaways

  • The overlay is not an entity, so no tool holds the data. Meta keys revenue to an ad and a product, your render vendor holds the template, and nothing joins them per impression.
  • A template comparison compares two setups. Delivery picks which product each person sees, so each arm carries its own product mix and the winner expires with the assortment.
  • Three labels means three arms. Each needs its own purchase volume and its own learning phase, which is why most catalogs afford one question a quarter.

Why is there no overlay column in Ads Manager?

Because the overlay never becomes something the reporting system counts. A catalog ad is one ad object pointing at a product set, and the label is styled once and stamped in as the ad renders. Fourteen hundred products and three template variants still collapse into one ad ID in the report.

Search Engine Land's writeup of Meta's dynamic overlay launch lists 4 label options (current price, a struck-through sale price, percentage off, free shipping) plus an auto-select mode where Meta picks based on performance signals. Auto-select makes it worse: if the label varies between impressions, there is no stable treatment to attribute a purchase to.

Breakdowns exist for dimensions the delivery system stores against an impression. The overlay is not one, and no amount of column-picking makes it appear.

Can any tool report ROAS by overlay type?

Not from the data you already have. Revenue sits with Meta, keyed to an ad ID and a product ID. Template identity sits with whatever rendered the image, keyed to a render job and a feed row. No shared key exists per impression, so a dashboard promising this breakdown is reading a split you built yourself or filling the gap with an assumption.

One join does work, and it carries a catch. Meta reports results per product ID. If your render tool stamped template A on one part of the catalog and template B on another, you can map product ID back to template and assemble the breakdown by hand.

What you get is the template tangled up with the assortment. The shoes ran the price badge, the jackets ran the rating star, and the winning label is the winning category.

What does a template comparison measure?

Two setups, not two creatives. In a catalog campaign, delivery chooses which product each person sees, so when you split the catalog into two product sets and run one template on each, every arm's ROAS carries whatever the system decided to serve underneath it. Holding audience, placement, budget and window steady still leaves that part moving.

There is a cheap fix most people skip because it feels wrong: split the catalog at random instead of by category. Half the shoes in each arm, half the jackets. It offends the merchandiser, and it is the only version where the two product mixes start out alike.

Even then the answer is directional and dated. New season, new assortment, and you run it again.

What does a three-way overlay test cost?

Three arms of purchase volume and three learning phases. Meta's documentation puts the learning phase at around 50 optimisation events per ad set in the week after a significant edit, so three ad sets means clearing that bar three times before any of them delivers normally. Only then does the reading window open.

I ran the two-arm version on a catalog spending about €900 a day, and it was six weeks before I would put a number in front of a client. The three-way version splits that spend three ways over the same six weeks. A thinner answer, sold internally as the thorough option.

So rank your questions. Run the one label you would rebuild the feed over, and leave the rest for next quarter.

How do you read a thin arm without fooling yourself?

By comparing it to the norm before you compare it to the other arm. A product set sitting on 20 purchases produces a ROAS that moves on two refunds, and the week-two winner is usually the arm that caught a good fortnight. Pull the number toward what your catalog normally returns, then look at the gap again.

That correction is the piece of my own scoring that transfers here. The composite runs on 6 metrics with Bayesian shrinkage against format-specific priors, so a wild early score gets dragged toward what its format usually does before anything is ranked.

The blunt test: would you rebuild 1,400 feed rows on this gap? If it does not survive that question twice, you do not have a finding.

What if the test is not worth running?

Then choose on the argument and write down why, so the next person inherits a decision instead of a habit. The three labels do different jobs. A current price qualifies the click and filters out people who were never paying it. A star rating borrows trust from buyers who already took the risk. A discount sticker pulls hardest and brings the shopper who came for the discount.

Pick the one that answers the objection your product has. If the objection is "is this worth the money", the price helps. If it is "will this be any good", the rating does more.

Whether the price belongs on the image at all is a separate decision, and it is the one to settle first.

Adscalr does not score overlays and never sees your catalog templates: the 6-metric composite runs per ad, and the Vision AI read of ~20 structured fields plus design, copy and strategy scores is pointed at competitor statics from the ad libraries. What it does is keep the reading honest once numbers exist, which is the ad intelligence job. Whether an arm had enough behind it to read comes off the product-level export, which has a trap of its own.

This is the thinking behind Adscalr.

See the product