← Blog

Can You Tell Which Thumbnail Won Without an A/B Test?

Inspect performance differences by creative attribute while distinguishing association from causation.

Try the analysis with sample data. Free, with no sign-in or file needed.

Everyone knows A/B testing is the right answer. The problem is that you cannot run one every time — the budget is small, there are too few creatives, or the network refuses to split delivery evenly.

There is a second-best option for those situations: turn the attributes of creatives you already served into data and compare them.

Setup: make attributes into columns

One creative is one row.

creative_id thumb_person title_has_number length_sec CTR
cr_001 1 0 15 0.021
cr_002 0 1 30 0.014

Attributes go in as 0/1 or as numbers. What matters here is recording facts, not judgements. "Well-made thumbnail" cannot be a column. "Has a person in it" can.

Why a simple average comparison is not enough

Suppose thumbnails with a person have a higher average CTR. If those thumbnails were mostly used on short videos, that gap may be the person or may be the length.

Putting several attributes in at once shows the difference once the other attributes are held at the same level. Content element analysis runs that calculation, and creative fatigue analysis separately requires daily dates, creative IDs, exposure, and outcome metrics. The attribute table above cannot establish a fatigue timeline.

Do not miss combination effects

One trap deserves attention. Attribute by attribute, you get as far as "person thumbnails are better" and "short videos are better." In reality, person thumbnails may only work on short videos.

That is a combination effect, and it shows up in a cross-tab of the two axes with performance per cell. When one exists, you cannot pick each axis separately — you have to pick the combination.

An illustrative cross-tab makes this concrete:

Short video Long video
Person present 2.4% 1.1%
No person 1.5% 1.4%

The direction differs by video length. Check support and uncertainty in all four cells before choosing an interaction to test.

Moving results into production guidelines

  • Do not use an attribute whose interval crosses its no-difference reference (0 for differences, 1 for odds or rate ratios). Not even the direction is settled.
  • Drop attributes with thin samples. An attribute present in three creatives just reflects the other characteristics of those three.
  • Send only the top one or two to an experiment. Turning all of them into guidelines stacks up unverified rules.

The job of observational analysis is deciding what to test next. Narrow the candidates, hand them to experiment analysis, and the question becomes one a small budget can settle.

To validate the candidates this produces, design a test with ad creative testing; for a wider view of which elements contribute, read ad creative performance analysis.

Try this today

  1. Open your last 20–30 creatives and add just three attribute columns — the three you actually argue about in retros. Three real columns beat fifteen aspirational ones, and you can fill three from the assets themselves without relying on memory.

  2. Before reading any coefficient, count how many creatives carry each attribute. Anything appearing in fewer than about five is describing those specific creatives, not the attribute. Mark those as unreadable rather than reading them anyway.

Limits of this approach

This is observational data, and the delivery algorithm chose which creatives got volume. It gave impressions to what it predicted would perform, so high-performing attributes are partly a record of what the algorithm liked, not only what audiences liked. That selection bias cannot be removed by adding more columns.

Which is why the output here is a shortlist, not a conclusion. Narrow to the top one or two candidates, hand them to experiment analysis, and let a small controlled test settle what the regression only suggested.

The honest framing to bring to a retro: "these two attributes are worth testing next," not "person thumbnails perform better."

How are you looking at content elements?

Review

Reviewed by Growth Opt Playbook

Frequently asked questions

How much should I trust a result produced without an experiment?
Use it only to narrow direction. When the delivery algorithm concentrates impressions on creatives that respond well, the attributes of those creatives look better than they are. Confirmation has to come from an experiment.
How many creatives do I need before this analysis works?
Meaningfully more than the number of attributes you are comparing. With four attributes you need at least dozens of creatives, and each attribute needs enough creatives both with and without it before the coefficients stabilise.