Definition
An incrementality test answers one question that attribution cannot: what would have happened if we had not run this at all?
Attribution takes the conversions that happened and divides credit between the touches it recorded. It assumes every recorded touch contributed something. Incrementality testing makes no such assumption. It removes the channel from part of your market, watches what happens, and reports the difference. That difference is the only number that describes cause rather than correlation.
The three methods
[table]
Method | How it works | Best for | Main limitation
Geo holdout | Turn the channel off in matched regions, keep it on elsewhere, compare pipeline | Any channel, any budget | Needs enough regions with comparable volume
Audience holdout | Platform withholds ads from a random share of your target audience | LinkedIn and Google conversion lift studies | Usually needs high spend to qualify
Time-based holdout | Pause the channel for a set period, compare against forecast | Small accounts with one market | Weakest method, seasonality contaminates it easily
[/table]
Geo holdout is the workhorse for most B2B SaaS companies. It needs no platform minimum, it works on modest budgets, and you control the design. Split your markets into matched pairs by historical pipeline, turn the channel off in one side of each pair, and compare.
What incrementality testing is really for
It exists to catch the two mistakes that attribution makes in opposite directions.
Over-crediting branded search. This is the big one. Someone hears about you on a podcast, reads your posts for two months, then searches your company name and converts. Last click gives all the credit to branded search. Multi-touch attribution gives most of it to branded search. Neither is wrong about what happened, and both are wrong about what caused it. A geo holdout on brand terms usually shows that a meaningful share of those conversions would have arrived anyway through organic, because the person was searching for you by name.
Under-crediting demand creation. LinkedIn campaigns built to create demand generate very few clicks and almost no direct conversions. In any attribution model they look weak. In a holdout test, the regions where you turned LinkedIn off often show fewer branded searches, fewer direct visits, and a longer sales cycle 60 days later. That is the channel proving itself in a way no attribution model can show.
Those two corrections usually point in the same direction, and it is the opposite of what the dashboard says.
Why most B2B SaaS incrementality tests fail
Two failure modes, and both are about patience.
The test is too short for the sales cycle. A B2B SaaS deal takes 45 to 90 days. A four week holdout measures the effect on form fills and tells you almost nothing about the effect on pipeline. Tests need to run for at least one full sales cycle, and then you need to wait another cycle to read the result properly. Eight to twelve weeks end to end is realistic.
The regions were never comparable. Matching on population or on spend is not enough. Match on historical pipeline, because that is the outcome you are measuring. Two regions with identical spend and very different win rates will produce a result that looks like a lift and is actually a sampling artefact.
There is a third issue worth naming. Most B2B SaaS companies simply do not have the conversion volume to detect a small effect. If you generate 30 opportunities a month across all channels, a holdout will not reliably detect a 10 percent lift. It will detect a 40 percent one. Design the test to answer a question your volume can actually support, or do not run it.
Incrementality testing at a glance
- Measures what a channel caused, by comparing exposed and unexposed groups.
- Three methods: geo holdout, audience holdout, time-based holdout.
- Geo holdout is the most practical for B2B SaaS at normal budget levels.
- Usually shows branded search is over-credited and demand creation is under-credited.
- Needs at least one full sales cycle running, plus another to read the result.
- Low conversion volume limits how small an effect you can reliably detect.
The rule for B2B SaaS
Run one geo holdout a year on the channel you are least sure about, and let it run long enough to matter.
The candidate is almost always brand bidding or LinkedIn. Brand bidding because it is the line item that looks best in every report and is the most likely to be partly redundant. LinkedIn because it looks weakest in every report and is the most likely to be carrying the rest of the account.
Design it simply. Pair your markets by historical pipeline rather than by spend. Turn the channel off in one side of each pair. Leave everything else untouched, including budgets in the live regions. Run it for a full sales cycle and read it a cycle later. Compare opportunities created and pipeline value, not clicks or form fills.
Then pair the result with the two things it cannot do. Data-driven attribution gives you direction between tests, and self-reported attribution on your demo form captures the podcast, the community mention and the peer recommendation that no test and no model will ever see.
Run the model for direction, the survey for the invisible touches, and the holdout for cause. Any one of the three on its own will give you a confident answer that is wrong in its own particular way.
Common questions about incrementality testing
What is incrementality testing?
A measurement method that compares a group exposed to your advertising against a matched group that was not, in order to measure how many conversions the channel actually caused rather than how many it was credited with.
What is the difference between incrementality and attribution?
Attribution divides credit among the touchpoints it recorded and assumes each contributed. Incrementality removes the channel from part of your market and measures what changed. Attribution describes correlation, incrementality measures cause, and the two often disagree sharply on branded search.
How long should an incrementality test run in B2B?
At least one full sales cycle while the test is live, which for most B2B SaaS companies is 45 to 90 days, and then a further cycle before reading the result. Eight to twelve weeks end to end is a realistic plan. Shorter tests measure form fills rather than pipeline.
What is a geo holdout test?
A test where you split your markets into pairs matched on historical pipeline, turn a channel off in one region of each pair, leave it running in the other, and compare opportunities created. It is the most practical incrementality method for B2B SaaS because it needs no platform spend minimum.
Which channel should you test first?
Usually branded search or LinkedIn. Branded search because it looks strongest in attribution and is the most likely to be capturing demand something else created. LinkedIn because demand creation looks weakest in attribution and is the most likely to be under-credited.
Book a Free Google Ads Audit