A whiteboard test that needs 80,000 users a week is not a strategy. It is a wish. When weekly actives sit closer to eleven thousand — a number we see often among regional marketplaces — you have three honest options: change the unit, lengthen the window, or do not run the test.

The dishonest option is to run it anyway, call a noisy week-2 wiggle a win, and hope finance does not ask for the power calculation. We have watched that sequence. It ends with a feature that cannot be killed because it was never truly tested.

Change the unit before you shrink the holdout

If contamination lives at the store, randomise stores. If search ranking mixes users, randomise listing clusters. Shrinking a user-level holdout to 5% to “save sample” usually just guarantees a result nobody should believe. A larger cell on a coarser unit is often the adult move.

Coarser units come with uneven sizes. Upcountry provinces will not match Bangkok clusters. Document the imbalance. Do not inverse-weight it into a fake national average unless you pre-registered that estimator and can explain it without a footnote.

The sentence you owe finance

“We do not have enough independent units to detect a 2-point move in four weeks. We can detect a 6-point move, or we can wait twelve weeks, or we can skip this idea.” That sentence is unfashionable in a growth meeting. It is the one we grade on the Experiment Bench.

Holdout Atelier sessions exist for teams who already shipped into this mess and need a stop rule after the fact. That is more expensive than designing the cell first. It is still cheaper than another quarter of contaminated CRM.

Module 03 of Holdout Design for Product Retention is built around thin samples. Bring your actual weekly actives, not the ones on the pitch deck.