Incrementality Testing

Incrementality testing measures how many conversions a channel actually caused, by comparing a group exposed to your ads against a matched group that was not.

On this page
What Is Centralized Notification Infrastructure?
The Default State: The Big Notification Mess
Building Workflows Without Engineering Dependency
Easy Debugging with Logs 
Embedded & Branded In-App Inbox
Multi-Tenant Aware Preference Center

Definition

An incrementality test answers one question that attribution cannot: what would have happened if we had not run this at all?

Attribution takes the conversions that happened and divides credit between the touches it recorded. It assumes every recorded touch contributed something. Incrementality testing makes no such assumption. It removes the channel from part of your market, watches what happens, and reports the difference. That difference is the only number that describes cause rather than correlation.

The three methods

[table]
Method | How it works | Best for | Main limitation
Geo holdout | Turn the channel off in matched regions, keep it on elsewhere, compare pipeline | Any channel, any budget | Needs enough regions with comparable volume
Audience holdout | Platform withholds ads from a random share of your target audience | LinkedIn and Google conversion lift studies | Usually needs high spend to qualify
Time-based holdout | Pause the channel for a set period, compare against forecast | Small accounts with one market | Weakest method, seasonality contaminates it easily
[/table]

Geo holdout is the workhorse for most B2B SaaS companies. It needs no platform minimum, it works on modest budgets, and you control the design. Split your markets into matched pairs by historical pipeline, turn the channel off in one side of each pair, and compare.

What incrementality testing is really for

It exists to catch the two mistakes that attribution makes in opposite directions.

Over-crediting branded search. This is the big one. Someone hears about you on a podcast, reads your posts for two months, then searches your company name and converts. Last click gives all the credit to branded search. Multi-touch attribution gives most of it to branded search. Neither is wrong about what happened, and both are wrong about what caused it. A geo holdout on brand terms usually shows that a meaningful share of those conversions would have arrived anyway through organic, because the person was searching for you by name.

Under-crediting demand creation. LinkedIn campaigns built to create demand generate very few clicks and almost no direct conversions. In any attribution model they look weak. In a holdout test, the regions where you turned LinkedIn off often show fewer branded searches, fewer direct visits, and a longer sales cycle 60 days later. That is the channel proving itself in a way no attribution model can show.

Those two corrections usually point in the same direction, and it is the opposite of what the dashboard says.

Why most B2B SaaS incrementality tests fail

Two failure modes, and both are about patience.

The test is too short for the sales cycle. A B2B SaaS deal takes 45 to 90 days. A four week holdout measures the effect on form fills and tells you almost nothing about the effect on pipeline. Tests need to run for at least one full sales cycle, and then you need to wait another cycle to read the result properly. Eight to twelve weeks end to end is realistic.

The regions were never comparable. Matching on population or on spend is not enough. Match on historical pipeline, because that is the outcome you are measuring. Two regions with identical spend and very different win rates will produce a result that looks like a lift and is actually a sampling artefact.

There is a third issue worth naming. Most B2B SaaS companies simply do not have the conversion volume to detect a small effect. If you generate 30 opportunities a month across all channels, a holdout will not reliably detect a 10 percent lift. It will detect a 40 percent one. Design the test to answer a question your volume can actually support, or do not run it.

Incrementality testing at a glance

  • Measures what a channel caused, by comparing exposed and unexposed groups.
  • Three methods: geo holdout, audience holdout, time-based holdout.
  • Geo holdout is the most practical for B2B SaaS at normal budget levels.
  • Usually shows branded search is over-credited and demand creation is under-credited.
  • Needs at least one full sales cycle running, plus another to read the result.
  • Low conversion volume limits how small an effect you can reliably detect.

The rule for B2B SaaS

Run one geo holdout a year on the channel you are least sure about, and let it run long enough to matter.

The candidate is almost always brand bidding or LinkedIn. Brand bidding because it is the line item that looks best in every report and is the most likely to be partly redundant. LinkedIn because it looks weakest in every report and is the most likely to be carrying the rest of the account.

Design it simply. Pair your markets by historical pipeline rather than by spend. Turn the channel off in one side of each pair. Leave everything else untouched, including budgets in the live regions. Run it for a full sales cycle and read it a cycle later. Compare opportunities created and pipeline value, not clicks or form fills.

Then pair the result with the two things it cannot do. Data-driven attribution gives you direction between tests, and self-reported attribution on your demo form captures the podcast, the community mention and the peer recommendation that no test and no model will ever see.

Run the model for direction, the survey for the invisible touches, and the holdout for cause. Any one of the three on its own will give you a confident answer that is wrong in its own particular way.

Common questions about incrementality testing

What is incrementality testing?

A measurement method that compares a group exposed to your advertising against a matched group that was not, in order to measure how many conversions the channel actually caused rather than how many it was credited with.

What is the difference between incrementality and attribution?

Attribution divides credit among the touchpoints it recorded and assumes each contributed. Incrementality removes the channel from part of your market and measures what changed. Attribution describes correlation, incrementality measures cause, and the two often disagree sharply on branded search.

How long should an incrementality test run in B2B?

At least one full sales cycle while the test is live, which for most B2B SaaS companies is 45 to 90 days, and then a further cycle before reading the result. Eight to twelve weeks end to end is a realistic plan. Shorter tests measure form fills rather than pipeline.

What is a geo holdout test?

A test where you split your markets into pairs matched on historical pipeline, turn a channel off in one region of each pair, leave it running in the other, and compare opportunities created. It is the most practical incrementality method for B2B SaaS because it needs no platform spend minimum.

Which channel should you test first?

Usually branded search or LinkedIn. Branded search because it looks strongest in attribution and is the most likely to be capturing demand something else created. LinkedIn because demand creation looks weakest in attribution and is the most likely to be under-credited.

Book a Free Google Ads Audit

Frequently asked
‍questions

What is ScalixAI?

ScalixAI is a performance-driven Google Ads agency specializing in helping high-growth, AI-first companies scale with predictable, profitable customer acquisition. Founded by an ex-Googler with 9 years of insider advertising experience, we manage the entire Google Ads lifecycle—from campaign strategy and account setup to conversion tracking, analytics, and ongoing optimization. Our data-centric, AI-powered approach ensures you know exactly which campaigns are working, why they’re working, and what to do next to outpace your competitors.ScalixAI is a performance-driven Google Ads agency specializing in helping high-growth, AI-first companies scale with predictable, profitable customer acquisition. Founded by an ex-Googler with 9 years of insider advertising experience, we manage the entire Google Ads lifecycle—from campaign strategy and account setup to conversion tracking, analytics, and ongoing optimization.

How fast can I expect results?

Most clients see performance stabilize by month three. Google Ads isn’t a slot machine—it takes time to compound.

Do you require long-term contracts?

No. We work month-to-month. All we ask is that you give us three months to prove the results.

Do you only run Google Ads?

While Google Ads is our entry point, we also support LinkedIn Ads, Reddit, and X campaigns when needed.

What’s included in your CRO audit, and what’s expected from our side?

The CRO audit covers your landing pages, CTAs, forms, and overall user flow. We’ll flag what’s holding back conversions and recommend fixes. If changes require design or dev resources, we’ll hand over clear action steps for your team, so you know exactly what to adjust.

How do you work with internal teams?

We integrate directly. Whether it’s syncing with your PMM on messaging, your design team on creative assets, or RevOps on tracking, we plug into existing workflows so we’re aligned and moving fast.

How do you handle Google rep recommendations that don’t fit our goals?

As an ex-Googler, I know which recommendations are useful, and which are just there to hit Google’s internal targets. We’ll filter their advice for you, implementing only what actually helps us hit revenue goals.

What changes in your approach to ads in B2B vs. B2C?

For B2B, I focus on lead quality, longer sales cycles, and nurturing conversions across the funnel. For B2C, speed and volume matter more, so I optimize for quick wins and scalable growth. Either way, the playbook adapts to your model.

What do your weekly reports include, and how do you define “good” vs. “scalable”?

Weekly reports show spend, conversions, CPL/CPA, and how we’re tracking against projections. “Good” means campaigns are meeting efficiency targets. “Scalable” means we can push budget and expect the same or better efficiency without breaking ROI.

Book a Call