← Blog

Reading Net Lift From a Holdout Test

Read incremental outcomes from the difference between exposed and holdout groups.

Try the analysis with sample data. Free, with no sign-in or file needed.

A Brand Search campaign has a great CPA. Retargeting collects many conversions. But if total revenue or new customers do not rise with spend, the report may include conversions that would have happened without advertising.

The question is not “how many conversions did this ad receive?” It is “how many conversions would disappear without this ad?” That difference is advertising uplift, or net incremental outcome.

Conversion-rate difference between an exposed group and a holdout group

Uplift is the difference between two worlds

If an exposed group converts at 8% and a holdout group converts at 5%, absolute uplift is 3 percentage points while relative uplift against the holdout is 60%. They are different units; label pp and % explicitly. That gap is the candidate outcome the ad added.

Uplift = exposed-group conversion rate − holdout-group conversion rate

The critical assumption is that the groups are alike except for ad exposure. That is why a randomized holdout is strongest. Regional or pre/post comparisons can help, but need more caution for seasonality, promotions, and price changes.

Why great CPA or ROAS is not enough

CPA and ROAS are calculated from people who encountered the ad. If an ad was the last touch for someone already likely to buy, the conversion still appears as advertising performance. Start with these campaigns:

  • Brand Search: paid ads may capture demand already looking for the brand.
  • Retargeting: the audience often has high purchase intent already.
  • Late promotion spend: the discount, rather than the ad, may drive conversion.
  • Products with strong organic demand: paid media may re-buy organic conversions.

A good CPA can be an operating-efficiency signal, but it is not automatically a scaling case. Read uplift and incremental ROAS alongside marginal efficiency: the average effect at current spend does not guarantee the return on additional spend.

How to design a holdout test

1. Choose one decision

Choose a real decision: keep Brand Search on, or increase retargeting budget by 20%. Lock one primary outcome such as install, signup, or purchase. Changing the goal midway invites favorable interpretation.

2. Split exposed and holdout groups

Use a platform Conversion Lift product or user-level random holdout when possible. Otherwise keep a regional, audience, or campaign-level comparison group. Check that the groups followed a similar trend before the test.

Do not choose the holdout share from a rule of thumb. Derive it from the minimum detectable effect, baseline conversion rate, available duration, and opportunity cost. Leave enough observation time for delayed conversions to mature.

3. Write success criteria before launch

Item Example
Primary metric New-purchase conversion rate
Minimum expected effect +1.0pp uplift
Observation window Two weeks
Next action Consider a small increase only if design and business criteria pass; otherwise hold and plan further testing

4. Read rate difference and absolute lift together

A +1pp uplift means about 100 extra conversions in a 10,000-person audience and about 10,000 in a million-person audience. Pair the rate with absolute incremental outcomes and cost, then calculate iCPA and iROAS from incremental rather than observed conversions.

How to read the result

Result Interpretation Next step
Confidence interval is positive Evidence that advertising added outcomes Scale in a small step, then re-check
Positive estimate but not significant Not proof of no effect; power may be low Hold judgment; preplan sample, window, and stopping rules for further testing
Near zero or negative Weak incremental evidence in this design Pause scale-up and review targeting or channel role

Do not translate “not significant” as “no effect.” But do not make a large scale decision from a p-value alone either. Absolute lift and iROAS must clear the business threshold.

When a holdout is not feasible

For a new campaign, compare a new ON period with a similar group left off. For an existing campaign, consider a short OFF period with difference-in-differences. These are weaker than random holdouts, so disclose comparison trends and external changes rather than claiming certainty from before/after alone.

Marketing Response Analysis can prioritize hypotheses from observational data. Its contribution estimates are hypotheses; a holdout should validate a large budget move.

For why uplift differs from the conversions in your report, incrementality measurement comes first; the decision rules — sample size, early stopping — are shared with A/B testing.

Wrap-up

Uplift is not a metric for cutting ads. It identifies the campaigns that truly create outcomes so you can invest with confidence. Before scaling a campaign because CPA looks good, ask how much disappears without it.

With exposed and holdout numbers, new-ON data, or shutdown data, use Incrementality Analysis to calculate the appropriate method entirely in the browser.

Try this today

  1. List your campaigns by reported CPA and look at the best one. Ask what share of its conversions would have happened anyway. Brand search and retargeting sit at the top of most reports precisely because they harvest existing intent — the best-looking campaign is often the least incremental.

  2. Pick one campaign and calculate the smallest holdout that can detect the effect you care about within the available window. Write the success criterion down before it runs. A deliberately sized holdout answers a question that attribution analysis cannot.

Limits of this approach

Uplift is estimated with uncertainty, and uncertainty is usually wide. A holdout that returns "positive but not significant" has not shown the ad works, and it has not shown it fails either — limited power is one possible explanation, but nonsignificance alone does not prove it. Reporting that as "no effect" is the most common way an incrementality result gets misused.

Nor does one measurement generalise. Uplift shifts with season, competitive pressure, and how saturated the channel already is, so a number measured in a peak month does not describe a quiet one. For campaigns carrying a large share of budget, re-measure periodically rather than treating the first result as settled.

How can you check the ad effect?

Sources and review2 references

Reviewed by Growth Opt Playbook

Frequently asked questions

Are uplift and incrementality the same?
In marketing practice they are often used for the outcome advertising truly added. Absolute uplift is a conversion-rate difference in percentage points; relative uplift divides that difference by the holdout rate and is expressed as a percentage.
Can I skip uplift testing when CPA is good?
No. Brand Search and retargeting can show strong CPA among people likely to convert anyway. Use a holdout when the decision is whether more spend will create additional business outcome.