Attribution

Short answer: A geo holdout test switches a channel off, or changes its spend, in some regions while leaving matched regions unchanged, then compares total sales or pipeline between them. On a small budget, test one channel, choose a few large, clean regions, collect daily pre-test data, run a power analysis first, and accept that you can only detect fairly large effects.

How to run a geo holdout test on a small budget (including India state-level design), cover

In my piece on MMM versus multi-touch attribution I argued that a geo holdout test is the most practical incrementality method for growth-stage companies, and in the ROAS article I covered Meta's built-in Conversion Lift. What I hear back from founders is: that sounds great at large budgets, but can we do it with modest spend? Often yes, if you design it carefully and are honest about what a small test can and cannot detect. This is the step-by-step version, including the India-specific details that trip people up.

What a geo holdout test measures, and why it suits small teams

A geo holdout test splits your market into regions. In the test regions you change one thing, usually switching a channel off or significantly increasing its spend. In the control regions, nothing changes. You then compare the business outcome, such as orders, revenue, qualified leads or pipeline, between the two groups against how they behaved before the test. The difference, after accounting for the pre-test relationship, is the incremental effect of that channel. The appeal for growth-stage companies is that it needs no user-level tracking, no cookies, and no consent to individual-level data joins. It only needs your own outcome data broken down by location and date, which most CRMs, Shopify stores and order systems already hold. It also measures what platform attribution cannot: whether the conversions a channel claims would have happened anyway. Two open-source tools make the analysis accessible. Meta's GeoLift package uses synthetic control methods, and Google's GeoX, part of its Meridian project, is designed for geo experiments too. You do not need either to start, but both document sensible design rules, and I lean on them below rather than on rules of thumb.

Design: one channel, one question, enough history

On a small budget, the biggest mistake is testing too much at once. Pick one channel and one question, for example: does our Meta prospecting spend drive incremental orders, or does branded search capture demand that would arrive anyway? The cleanest design for small budgets is usually a go-dark holdout: keep spend running everywhere except the control regions, where you switch the channel off. It costs nothing extra; you simply forgo some spend in a few regions. Next, data. You need the outcome metric by region and by day. Google's GeoX documentation asks for at least three times the test length in daily pre-test data, and a year or more if your business is strongly seasonal, and it accepts only daily data, not weekly. Meta's GeoLift walkthrough says every experiment needs at least a location, a date and the outcome metric, and that missing dates or locations get dropped, so clean the data first. Test length: GeoLift advises that the test cover at least one full purchase cycle, and GeoX suggests extending, for example from four weeks to six or eight, when daily data is volatile. For long B2B cycles, measure a leading outcome like qualified opportunities rather than closed revenue.

Power: what a small budget can and cannot detect

Every geo test should start with a power analysis, which estimates the smallest effect your design could reliably detect given your data's noise. Meta's GeoLift documentation calls running a prospective power analysis fundamental before executing a test, and its walkthrough shows how to simulate different effect sizes, test lengths and numbers of test markets. Here is the honest implication for small budgets. If the channel you are testing drives a small share of total sales, switching it off in a few regions produces a small dip, and that dip may be smaller than the normal day-to-day noise in those regions. The test will then come back inconclusive, which is not the same as showing the channel does nothing. Three ways to improve the odds without more money: choose regions with high, stable volume, since noisy small regions swamp the signal; run longer; and test a bigger change, such as fully off rather than a 20% cut. If the power analysis says you can only detect an effect far larger than you expect the channel to have, do not run the test yet. Spend the time fixing data quality or consolidating regions instead. An underpowered test wastes weeks and produces a result people will over-interpret.

India state-level design: the details that matter

States are the natural geo unit in India: Google Ads and Meta both support targeting at region level, and most order and CRM data includes a state or pincode. But states are very uneven, so a few rules. Do not pair states by population. Pair them by how similar their historical sales curves are; synthetic control methods do this matching for you. Handle metro spillover. The National Capital Region spans Delhi, Haryana and Uttar Pradesh, so Gurugram and Noida buyers often live, work and shop across state lines. Treat NCR as one unit, or exclude it, rather than putting Delhi in test and Haryana in control. Similar care applies to other cross-border commuter belts. Watch regional calendars. Festival peaks differ by state, for example Onam in Kerala, Durga Puja in West Bengal and Pongal in Tamil Nadu, and a festival falling in your test window in only some regions will distort the result. Check the location setting in Google Ads. Google's default option, presence or interest, can show ads to people outside a region who have shown interest in it; for a geo test, use the presence option so ads reach people physically in or regularly in the targeted area. Finally, define location from delivery address or billing state, consistently, before the test starts.

Running it and reading the result

During the test, change nothing else in the test or control regions: no new promotions in one group only, no pricing changes, no other channel shifts. Keep a log of anything unavoidable, such as a stock-out in one state, so you can account for it. When the test ends, compare actual outcomes in the test regions with the counterfactual the model built from control regions, and look at the confidence interval, not just the point estimate. Convert the lift into cost per incremental outcome: spend withheld or added, divided by incremental orders, leads or pipeline. Then compare that number with what the platform reported for the same period. If the platform claimed far more conversions than the test found incremental, you have learned that part of the channel's reported performance was demand you would have captured anyway, and you can recalibrate how you read platform ROAS. If the result is inconclusive, report it as inconclusive. Re-run with a stronger design rather than picking the reading you prefer. Once you have one clean result, schedule the next test on a different channel. A small library of geo tests, repeated yearly, is the most credible incrementality evidence a growth-stage company can show its board or investors.

Sources

Meta GeoLift documentation, Walkthrough (data requirements, purchase cycle guidance, power analysis): https://facebookincubator.github.io/GeoLift/docs/GettingStarted/Walkthrough/ Google Meridian GeoX, Prepare your pretest data (at least 3x test length of daily data, daily data required, extending tests when data is volatile): https://developers.google.com/meridian/geox/prepare-your-pretest-data Google Ads Help, About advanced location options (Presence vs Presence or interest): https://support.google.com/google-ads/answer/1722038

FAQ

A go-dark holdout needs no extra budget; you switch a channel off in control regions. What matters more is volume and noise in your outcome data. Run a power analysis first to see the smallest effect you could detect.

At least one full purchase cycle, per Meta's GeoLift guidance. Google's GeoX suggests extending, for example from four weeks to six or eight, when daily data is volatile. Collect at least three times the test length of daily pre-test data.

Yes, but match states on historical sales patterns rather than population, treat cross-border metros like NCR as one unit, avoid region-specific festival periods, and use the Presence location option in Google Ads.

Report it as inconclusive, not as zero effect. Improve the design with higher-volume regions, a longer test or a bigger spend change, then re-run.

Read this article on your favourite platform

Ready to build the system?

If this describes your funnel, a 30-minute call will find where the constraint sits in your own numbers and what it would take to fix it.

It starts with a 30-minute call. Pick a time below.

Choose a time

Prefer email?