When to stop an A/B test without guessing
Most product teams do not struggle with launching variants. They struggle with knowing when the experiment has said enough. Closing too early turns noise into a roadmap item. Waiting forever freezes shipping.
A practical stopping rule starts before traffic flows. Agree on the primary metric, the minimum lift worth chasing, and the calendar window that matches one full usage cycle for your app. For a weekly grocery app that might be seven days; for a payroll product it might be a full pay period.
Peeking at results mid-test is fine if you treat mid-window numbers as directional only. The decision meeting should use the pre-agreed window, not the first day the chart looked exciting. If leadership needs interim updates, present confidence bands and remaining sample, not a premature ship/kill vote.
When the window ends, ask three questions: Did we reach the planned exposure? Did guardrail metrics stay acceptable? Is the lift large enough to justify the engineering and support cost of shipping the winner? A statistically detectable change that is too small to matter in revenue or retention can still be a clean stop — just stop by keeping the control.
If results conflict across segments, resist averaging them into a single story. Document which audiences saw which outcome, then decide whether a targeted rollout is safer than a global change. That decision belongs in product judgment, informed by the numbers — not replaced by them.
A/B testing stopping rules