Multiple comparisons: More comparisons, more problems
Tue Sep 29 2026
The multiple comparisons problem is one of the most common and easily overlooked statistical pitfalls in A/B testing.
This problem can happen when you have multiple metrics you’re looking at in one batch of testing data, more than two variants, or resulted sliced by segment. If left unaddressed, it increases the likelihood of at least one false positive (Type I error) across a set of comparisons. This can lead teams to make decisions based on noise rather than genuine impact.
So, what exactly is the Multiple Comparisons Problem, why does it matter, and how does Statsig guard against it?
What is the multiple comparison problem?
When we run an A/B experiment, we often use a 5% significance level, commonly described as a 95% confidence threshold. Let’s say we run a sufficiently powered experiment to full sample and observe a statistically significant 3% lift in our primary KPI. A 5% significance level means that a single no-effect comparison will still look like a win about once in twenty. We accept that trade-off for a real shot at a 3% lift.
However, that 5% false-positive risk applies only to a single comparison, i.e., the primary metric, Control vs. Treatment. For testing teams, experiments are often launched with multiple important metrics and several audience segments to analyze, such as desktop versus mobile. This is where the multiple comparisons problem comes in: When we test multiple hypotheses and treat any significant result as evidence of an effect, we create more opportunities for false positives. Each individual test can still have a 5% false-positive rate, while the probability of at least one false positive across the set is much higher.
The math behind this is straightforward, and you can explore it with the calculator below! The underlying formula is 1-(1-α)ⁿ, where α is your alpha (significance level) and n is the number of comparisons.
Multiple comparisons calculator
Multiple comparisons calculator
More comparisons mean more opportunities for a false positive. See how the chance adds up across a test.
False-positive chance
Calculated as 1 − (1 − α)n. Assumes independent comparisons and no real effect in any comparison (all null hypotheses are true). Correlated metrics can change the overall chance. This is the chance of at least one false positive, not the probability that a particular result is false.
Try using a significance level of .05 for just one comparison. You’ll see the result is .05, or 5%. However, if you use that same alpha for 10 comparisons, the false-positive rate jumps from 5% to 40%.
This means that in an experiment with 10 independent comparisons, there is roughly a 40% chance that at least one will appear statistically significant without correction.
Why is that a problem?
When we don’t correct for multiple comparisons, the probability of a Type I error increases dramatically. The Type I error rate is the false positive rate, or the detection of an effect that isn’t actually there.
In the example above, a Type I error is relatively harmless. We can intuitively tell that even though the fire alarm is going off, there isn’t a fire, so we’re not in danger. But in experimentation, we often don’t have an easy way to check whether an apparent “win” was a false alarm. What happens when we make business decisions based on metrics that didn’t actually move? When we tell our stakeholders we improved revenue by 3%, but we actually didn’t?
The consequences can extend beyond a single inaccurate readout. Teams feature, prioritize the wrong follow-up experiments, and report gains that fail to materialize. For experimentation programs with multiple launches per week, repeated false wins can distort the program’s understanding of what works.
What does Statsig do to control it?
Luckily, Statsig offers two approaches to address multiple comparisons: the Bonferroni Correction and the Benjamini-Hochberg procedure. They control different types of error, so understanding the distinction helps teams choose an approach that fits their goals.
A Bonferroni Correction is a method for controlling the family-wise error rate (FWER), while the Benjamini-Hochberg procedure controls the false discovery rate (FDR).
Bonferroni asks, “What’s the chance I called anything a win that wasn’t?”
Benjamini-Hochberg asks, “Of the comparisons I called wins, what share could be noise?”
Bonferroni is often described as swatting a fly with a sledgehammer because it takes a very conservative approach to minimizing false positives. Bonferroni is one division: you take your preset alpha and divide it by the total number of comparisons.
In the example above, we would take .05 / 10, which returns 0.005, corresponding to a 99.5% confidence level. A result would need a p-value at or below 0.005 to be declared significant. Every comparison faces the same adjusted threshold. As more comparisons are added, that threshold becomes stricter, making real effects harder to detect. Bonferroni is useful when teams want strong protection against even one false positive within the family of comparisons. Statsig lets you adjust the conservatism of a Bonferroni correction by enabling a Preferential Bonferroni, which allows you to give your primary KPIs a larger share of alpha and split the rest across the secondary metrics.
Contrary to Bonferroni, Benjamini-Hochberg does not use a single bar for every comparison. It takes all p-values in the experiment, sorts them from smallest to largest, and assigns each rank a threshold. The Benjamini-Hochberg threshold = (rank / total number of comparisons) × target FDR. Note: Target FDR is often set to 0.05.
It then finds the largest rank whose p-value meets the threshold. That result and all results with smaller-ranked p-values are declared significant. When the numerical error targets are the same, the first BH threshold roughly equals the Bonferroni threshold. Each subsequent threshold is less restrictive. Statsig also allows you to apply Benjamini-Hochberg to only the primary metrics and leave the rest of the scorecard, including guardrails, uncorrected. This can be beneficial, as we’d want to retain sensitivity, like a real dip on a guardrail metric to still show up.
Generally, a Benjamini-Hochberg correction can help teams detect more real effects while controlling the expected proportion of false discoveries. That flexibility comes from controlling FDR rather than the stricter FWER.
For this example, assume these five metrics represent the entire family of comparisons, with α = 0.05 for the uncorrected and Bonferroni tests and a target FDR of 0.05 for BH.
Without correction, four results are significant. Bonferroni retains one, while Benjamini-Hochberg retains two. For Benjamini-Hochberg, rank two is the highest rank that meets its threshold, so the first two results are declared significant.
Corrected conclusions
When configuring multiple-comparison corrections in Statsig, consider both the error you want to control and which comparisons the selected setting covers. Use this decision table to help you understand which to apply. Ensure you choose your approach before reviewing results: