To the glossary
AnalyticsTerm

Statistical significance

statistical significance · significance of A/B test · p-value · confidence interval

Statistical significance is the confidence that the difference between the A/B test options is real and not random. Standard: p < 0.05 (95% confidence).

Statistical significance is the probability that the observed result of an A/B test is not explained by random fluctuations in traffic. Standard threshold in marketing: p-value < 0.05, that is, the probability of mistakenly recognizing the result as real is less than 5%. This is called the “95% confidence level”.

Practically: if the test showed +12% conversion for option B with 95% significance, there is a 5% chance that this is just noise and there is no real effect. This is an acceptable risk for marketing decisions. For medical research, the threshold is stricter - p < 0.01 or p < 0.001.

The most common mistake in A/B testing is stopping the test too early. Example from practice: the test was launched on Monday, by Wednesday option B shows +25% conversion with 91% confidence, the test is stopped and the winner is announced. By Friday it turns out that the audience at the beginning of the week and at the end are different, and there is no real effect. Rule: Determine the minimum sample size before running the test (calculators are available), and do not touch the test until this sample is reached.

The second common sin is testing too many hypotheses at once. With 20 parallel tests with a 95% threshold, one “win” will turn out to be statistical noise simply according to the laws of probability. I stick to the rule: no more than 3-4 active tests simultaneously on the same traffic segment, and always with a Bonferroni correction for multiple testing.

Frequently asked questions about Statistical significance

What is statistical significance?+
This is confidence that the difference between the A/B test options is real and not random noise. The standard threshold in marketing p-value is less than 0.05, that is, the probability of mistakenly recognizing the result as real is less than 5%. This is called a 95% confidence level.
What does p < 0.05 mean?+
That the probability of getting such a result by chance, in the absence of a real effect, is less than 5%. For marketing decisions, this is an acceptable risk. In medicine, the threshold is stricter, p < 0.01 or p < 0.001.
Why can't you stop an A/B test early?+
Because in a small sample, the early advantage often turns out to be noise: the audience is different on different days of the week, and the effect disappears. You need to calculate the minimum sample size in advance and not stop the test until it is reached.
Is it possible to run many A/B tests at once?+
Be careful. With 20 parallel tests with a 95% threshold, one false victory is almost guaranteed according to the laws of probability. I keep no more than 3-4 active tests on one traffic segment and apply a Bonferroni correction for multiple testing.

Related terms

Where is it understood in practice?

Need to set this up on your project?

I analyze metrics, calculate unit economics and collect funnels on real budgets. 30 minutes on call - free.