Holdout test
A holdout test measures a marketing activity's incremental effect by randomly withholding it from a portion of the audience - the holdout group - and comparing their behavior to the treated group that received it. Ecommerce operators use it to prove whether an email flow or ad set actually causes extra sales or just reaches people who would have bought anyway.
| Term | What it covers |
|---|---|
| Treated group | Customers who received the activity - the email sent, the ad shown, the flow enrolled. |
| Holdout group | A random, representative slice deliberately excluded, standing in for "what happens with nothing". |
| Incremental outcome | Treated result minus holdout result, measured per customer so the two groups compare fairly. |
| Significance | Whether the gap is larger than random noise, given how many customers are in each group. |
Lift % = (treated rate − holdout rate) ÷ holdout rate. The holdout is the counterfactual platform reporting never shows you.
Worked example
Example numbers. The flow's send cost is near zero, so almost all of the A$270,000 is incremental margin - but only a sixth of what the platform claimed.
What is a good holdout test?
A good holdout test is one that can actually detect the effect it is looking for and then answers an economic question, not a vanity one. There is no benchmark lift percentage - a small but statistically significant lift on a near-free email flow is excellent, while a large lift that fails significance tells you nothing you can bank. The two checks that matter: is the holdout big enough that the gap is real rather than noise, and does the incremental revenue exceed the cost of running the activity. On a costly ad set the bar is higher than on an owned email, so "good" is defined by your margin and the activity's cost, never by the raw lift number alone.
Holdout test vs related metrics
| Metric | What it measures | How it differs from a holdout test |
|---|---|---|
| Geo-lift test | Incremental lift from changing spend by region. | Randomizes by geography; a holdout randomizes by user - the tool when you can withhold from individuals. |
| Incrementality | The extra sales an activity actually causes. | The quantity a holdout is the cleanest way to measure for targetable channels. |
| Marketing mix modeling (MMM) | Channel contributions modeled from history. | MMM covers every channel at once from the past; a holdout measures one activity directly, now. |
| Attribution model | Credit assigned to tracked touchpoints. | A holdout tells you what that credit is really worth, by showing what happens with the touch removed. |
Common mistakes
- An underpowered holdout. Too small a group cannot detect a modest lift, so a real effect reads as "no difference" and a working activity gets killed.
- A non-random split. If the holdout skews toward lower-value or dormant customers, the gap measures the split, not the activity.
- Contamination. Holdout customers still reached through another channel are no longer a clean control, and the measured lift shrinks toward zero.
- Peeking and stopping early. Ending the test the moment the gap looks good inflates false positives; decide the duration up front.
- Measuring on revenue, not contribution. A discount-heavy flow can lift revenue while adding no margin - judge the lift on contribution margin.
Holdout test FAQ
Related
Blufire S13 Experiments runs email and ad-set holdouts (and geo-lift tests) on your own audiences, so the incremental revenue behind each flow and campaign is measured rather than assumed.
Updated July 2026