The incrementality test calculator above compares an exposed test group with an unexposed holdout and returns the conversions your advertising actually caused — not the conversions that happened to follow an ad impression. It also returns the cost of those causal conversions, the relative lift between the two groups, and a significance check so you know whether the difference is large enough to act on.
Arb Digital runs holdout tests for clients whose platform-reported returns look strong but whose overall revenue is flat, because that combination almost always means the ad account is claiming credit for demand that already existed. This calculator does the arithmetic behind that diagnosis, and the sections below explain how to read a result that will usually be less flattering than your dashboard.
What This Incrementality Test Calculator Does
Enter the size and conversion count of your exposed group and your holdout, and the tool calculates the conversion rate in each, then applies the holdout's rate to the test group to estimate how many conversions would have happened anyway. Subtract that baseline from the observed test conversions and what remains is incremental — the conversions attributable to the advertising.
Add the media spend on the test group and you get incremental cost per acquisition alongside the reported cost per acquisition your platform would show, so the difference between the two is visible in a single view. The tool also reports relative lift, the proportion of conversions that were genuinely incremental, and a two-proportion significance test with the p-value, so a small apparent lift on a small sample is not mistaken for a real effect.
How to Use It
- Enter your test group size. Use the count of people, households or geographic units eligible to be exposed, not the count that actually saw an ad.
- Enter test conversions. All conversions from that group during the test window, from your own analytics or order data rather than platform-attributed conversions.
- Enter your holdout group size and conversions. The groups do not need to be equal — the calculator normalises by rate, not by count.
- Enter the spend used on the test group during the same window so the incremental cost per acquisition can be calculated.
- Click Calculate and check the p-value before drawing any conclusion from the lift figure.
The Formula / How It's Calculated
The core logic is a counterfactual. Conversion rate in each group is conversions divided by group size. The expected baseline for the test group is holdout rate × test group size — what the exposed group would have produced with no advertising, assuming the two groups are otherwise comparable. Incremental conversions are then test conversions − expected baseline.
With the default figures, the test group converts at 4,200 ÷ 500,000 = 0.84% and the holdout at 3,400 ÷ 500,000 = 0.68%. The expected baseline is 0.68% × 500,000 = 3,400, so incremental conversions are 4,200 − 3,400 = 800. Relative lift is (test rate − holdout rate) ÷ holdout rate, or 23.5%. Incremental CPA is spend ÷ incremental conversions = $60,000 ÷ 800 = $75.00, against a reported CPA of $60,000 ÷ 4,200 = $14.29. The incremental share of conversions is 800 ÷ 4,200 = 19%.
Significance uses the standard two-proportion z-test: the difference in rates divided by the pooled standard error, z = (p₁ − p₂) ÷ √[p(1−p)(1/n₁ + 1/n₂)], where p is the pooled conversion rate across both groups. The two-tailed p-value comes from the normal approximation, which is appropriate at these sample sizes. The NIST/SEMATECH e-Handbook of Statistical Methods documents this test and its assumptions in detail.
Why Platform-Reported Conversions Overstate Contribution
An ad platform records a conversion when someone who saw or clicked an ad later converts within the attribution window. That is a correlation, and for a meaningful share of those conversions the person would have bought anyway. Branded search is the clearest case: someone who already decided to buy searches your company name, clicks the ad sitting above your own organic result, and converts. The platform books a conversion. Almost none of that revenue was created by the ad.
Retargeting has the same structure. It shows ads to people who already visited your site and demonstrated intent, which makes attributed performance look outstanding and incremental performance look modest. This is not a flaw in the platforms' reporting so much as a limitation of what attribution can see — which is also why the platforms offer separate lift study products alongside standard attribution reporting, documented in the Meta Business Help Center. Only withholding advertising from a comparable group reveals the counterfactual, which is why the gap between the reported CPA and the incremental CPA in this calculator is the entire point of running the test.
Reading a Disappointing Incremental CPA
A first incrementality test usually delivers an uncomfortable number. A reported CPA of $14 turning into an incremental CPA of $75 looks like the channel has failed. Before cutting it, check what the $75 is being compared against. If your contribution margin per conversion is $110, a $75 incremental acquisition cost is still profitable — it is just far less profitable than the dashboard implied. The correct response is to reprice your expectations, not to switch the channel off.
It also matters which part of the account was tested. Testing a whole platform at once produces a blended answer that hides enormous variation: brand search and retargeting typically show low incrementality, while prospecting campaigns aimed at cold audiences typically show high incrementality. A blended result can lead you to cut the campaigns that were doing the real work. Test at the level you intend to make decisions at, and run the resulting numbers through the CPA calculator and the ROAS calculator against your actual margin.
What Makes a Holdout Valid
The whole calculation rests on one assumption: that the only systematic difference between the two groups is exposure to advertising. Contamination breaks that assumption, and it is easy to introduce accidentally. A user-level holdout on one platform does not stop that person seeing your ads on another platform, or receiving your email, or walking past a billboard. Household sharing, shared devices and logged-out browsing all leak exposure across the boundary.
Geographic tests avoid most of that, at the cost of a much smaller effective sample — you are comparing regions, not people, so the unit of randomisation is the region. If you use geo tests, match regions on prior conversion volume and seasonality rather than on population, and pick enough of them that one unusual city cannot swing the result. Whichever design you use, define the groups before the test starts and never reassign anyone mid-flight, because reassignment silently reintroduces the selection bias the test exists to remove.
How Long to Run and When to Stop
Incrementality tests need more data than conversion-rate A/B tests, because you are measuring the difference between two already-small rates rather than a difference in a rate you can amplify with traffic. The practical constraint is usually that the effect being measured is a fraction of an already low baseline, so the standard error is large relative to the difference you hope to detect.
Fix the test window in advance and do not stop early because the numbers look good, since peeking repeatedly at accumulating data inflates the false-positive rate substantially. The window should also cover at least one full purchase cycle: if your typical customer takes three weeks to decide, a two-week test measures the delay in purchases, not the increase in them. Size the test before launching with the sample size calculator, and if you are running a straightforward creative or landing page comparison instead, the A/B test calculator is the right tool.
Arb Digital designs and runs holdout tests across paid search and paid social, then rebuilds budget allocation around incremental performance rather than platform-attributed performance.
Paid Advertising Services Google Ads & PPC ServicesCommon Mistakes to Avoid
- Using platform-attributed conversions for both groups — the holdout has no attributed conversions by definition, so both counts must come from your own analytics or order data.
- Stopping the test when the result looks good — repeated peeking at accumulating data makes a false positive far more likely than the stated significance level suggests.
- Testing an entire platform at once — brand, retargeting and prospecting have very different incrementality, and the blended answer hides all of it.
- Ignoring cross-channel contamination — a holdout that still receives your emails and organic touches is not an unexposed group.
- Running a window shorter than the purchase cycle — a short test measures deferred purchases rather than additional ones.
Related Free Tools From Arb Digital
Size your test before launch with the sample size calculator, evaluate simpler creative tests with the A/B test calculator, and check the baseline you are measuring against with the conversion rate calculator. Convert the result into business terms with the CPA calculator or the marketing ROI calculator, and browse the full free online tools hub.
Frequently Asked Questions
Incrementality is the portion of conversions that would not have happened without the advertising. It is measured by comparing an exposed group against a comparable unexposed holdout, rather than by counting conversions that followed an ad impression.
Reported cost per acquisition divides spend by every conversion the platform attributes to the ads, including conversions that would have happened anyway. Incremental cost per acquisition divides the same spend by only the conversions the advertising caused, so it is almost always the higher figure.
Large enough that a difference worth acting on would be statistically detectable, which depends on your baseline conversion rate and the size of the lift you expect. Small baseline rates require much larger groups, so size the test in advance rather than after launch.
Many teams use 0.05 as the threshold, meaning a result this extreme would occur by chance about one time in twenty if there were no real effect. Set the threshold before the test runs and interpret a marginal result as inconclusive rather than as a small win.
Not with the same reliability. Alternatives such as before-and-after comparisons or media mix modelling estimate incrementality without withholding advertising, but they rely on far stronger assumptions because no genuinely unexposed comparison group exists.
Branded search and retargeting typically show the lowest, because both concentrate on people who have already demonstrated intent. Prospecting campaigns aimed at audiences with no prior relationship generally show the highest, which is why blended platform-level tests can be misleading.
Not automatically. Low incrementality means the channel creates less demand than its reported figures suggest, but the incremental cost per acquisition can still sit comfortably below your contribution margin. Compare the incremental figure against your own economics before making a decision.
Results produced by this tool are statistical estimates only — validity depends on the comparability of your test and holdout groups, and on the test design rather than on the arithmetic.