Statistical Significance in Plain English: When an A/B Test Is Real

Statistical significance is one of those phrases people either nod along to or tune out completely, and both reactions end up costing small business owners money. I’m not a statistician. I run small businesses and got tired of being fooled by my own numbers, so I’ve stripped this idea down to the version that actually changes decisions. No equations, just the intuition you need to look at a test result and know whether it’s real or just noise wearing a costume.

Start with a coin

Imagine I hand you a coin and claim it’s rigged to land heads. You flip it four times and get three heads. Is it rigged? Obviously not. Three out of four happens constantly with a perfectly fair coin. Now flip it four hundred times and get three hundred heads. That’s the same 75%, but it’s a completely different feeling, because a fair coin almost never produces that ratio at that volume. Statistical significance is just a precise way of describing that gut feeling about how surprised you should be.

The question significance actually answers

Every significance test answers one specific question: if there were truly no difference between your two versions, how likely is it that chance alone would produce a gap at least as big as the one you’re seeing? If the answer is “very unlikely,” you call the result significant, meaning chance is a poor explanation and a real effect is a better one. That’s the whole idea. It isn’t proof. It’s a statement about how easily luck could have faked what you’re looking at.

What a p-value actually is

The number that captures this is the p-value, and it’s badly named and widely misunderstood. A p-value of 5% does not mean there’s a 5% chance your result is wrong. It means that if there were genuinely no difference, you’d see a gap this big or bigger about 5% of the time from chance alone. A small p-value means chance is a weak explanation, so you lean toward a real effect. It’s a measure of how surprising your data would be in a world where nothing was actually happening.

The 95% line is a convention, not a law

People treat 95% confidence like a law of nature. It’s a convention, a line someone drew that says a result counts as significant when chance would only produce it 5% of the time or less. There’s nothing sacred about 95%. It’s a reasonable default for balancing two risks, but for a low-stakes tweak you might accept more uncertainty, and for a decision you can’t reverse you might want much more. The number is a choice about how much luck you’re willing to be fooled by, not a truth handed down from above.

Why sample size drives everything

Significance depends enormously on how many people were in your test, which is exactly why small operators get burned. With a handful of conversions on each side, even a large percentage gap can easily be chance, because tiny groups swing wildly. The more people flow through each version, the harder it is for luck to manufacture a big difference, and the more a real gap can separate itself from the noise. It’s the coin again: four flips tell you almost nothing, four hundred tell you a lot, and there’s no shortcut around needing the volume.

The difference that’s real but tiny

Here’s a trap that catches bigger operations too. With enough traffic, almost any difference becomes statistically significant, including ones so small they don’t matter. A test can prove, at high confidence, that a change lifts your conversion rate by a fraction of a fraction of a percent. Significant, and useless. Significance is only half the question. The other half is the size of the effect: whether the difference is big enough to be worth the trouble of shipping and maintaining. Real and worthwhile are two separate checks.

Practical significance vs. statistical significance

That gives you two questions to ask of every result, not one. First: is this difference likely real, or could chance easily explain it? That’s statistical significance. Second: is this difference big enough to change what I do? That’s practical significance. A change can pass one test and fail the other, in either direction. Small operators should honestly care more about the second question than they usually do, because a real but microscopic lift isn’t worth rebuilding your product for, no matter how confident the test is.

The multiple comparisons trap

Remember the coin. If a fair coin gives a surprising streak rarely, then flipping twenty different coins makes a surprising streak somewhere almost certain. The same thing happens when you test twenty changes, or slice one result twenty different ways. Some slice will look significant by pure chance, and it means nothing. If you go hunting through enough numbers, you’re guaranteed to find a flattering one. This is why deciding your one metric before the test even starts matters so much. It stops you from fishing until luck hands you a false winner.

Confidence intervals tell you more

A p-value gives a yes-or-no flavor, but a confidence interval is often more honest and more useful. Instead of a single number, it gives a range the true effect probably lives in. A test might say the real lift is somewhere between 1% and 9%. That range tells you both that the effect is probably real, since it doesn’t include zero, and how much you still don’t know, since it’s wide. A range respects your uncertainty in a way a single verdict never does, and it keeps you humble about how precise your answer really is.

What significance does not promise

Be clear about the boundaries. A significant result doesn’t mean the effect is large, it doesn’t mean it will hold forever, and it doesn’t mean your test was designed well. Significance only speaks to the role of chance in the specific numbers you collected. A beautifully significant result from a broken or biased test is just confidently wrong. The statistics sit on top of the quality of the experiment underneath, and no amount of significance can rescue a test that measured the wrong thing or split its groups unfairly.

A simple rule for small tests

Here’s the practical rule I actually use. If the test is small, distrust it, no matter how good the gap looks, because small samples lie easily. If the difference is tiny, ignore it even when it’s significant, because it won’t move your business. And if you tested many things at once, treat every exciting result as a hypothesis to confirm with a fresh test, not a conclusion. Those three habits catch most of the ways a solo operator gets fooled, without a single formula.

Bigger effects need less data

There’s a comforting flip side to the sample size problem. The larger the true difference between your two versions, the less data you need to be confident it’s real. A change that doubles your conversion rate will announce itself clearly even on modest traffic, because chance struggles to fake a gap that big. It’s the small, subtle improvements that demand mountains of visitors to prove. If a change is genuinely dramatic, don’t agonize over a small sample. If a change is marginal, accept that you may simply never have the traffic to settle it.

The false positive you will hit

If you run tests all year at 95% confidence, you’re accepting that roughly one in twenty times, chance alone will hand you a significant result for a change that does nothing. That’s not a flaw you can eliminate. It’s the price of the whole method. The defense is memory and repetition. A genuinely real effect tends to show up again when you retest it or watch it live, while a false positive quietly fails to reappear. Treating one surprising win as a reason to look again, rather than a settled fact, catches most of these.

Regression to the mean

Here’s a trap that fools people constantly. Whenever you act right after an unusually good or unusually bad stretch, the next stretch tends to drift back toward normal on its own, with or without your intervention. Change something after a terrible week and things improve, and you’ll credit the change, when a lot of the recovery was just the week being unusually bad and the average reasserting itself. A proper test with a parallel control group protects you from this, because the control group drifts back too, and you compare against it instead of against your own bad week.

Stopping rules in plain English

The cleanest way to avoid peeking yourself into a false result is a stopping rule you set in advance. Decide, before launch, either how long the test runs or how many people go through each side, and don’t look at the verdict until you get there. The reason is simple: the more often you check and reserve the right to stop, the more chances luck gets to briefly cross your finish line. One honest look at a predetermined end beats a hundred anxious peeks, every single time, and it costs you nothing but patience.

Confidence is not certainty

The language around these tests oversells them, so it helps to translate as you go. Ninety-five percent confident does not mean 95% certain the effect is real. It describes how often this whole method would avoid being fooled by chance if you ran it many times over, not a probability stamped onto your single result. Every individual test still lives with genuine uncertainty underneath it. The honest posture is to hold your conclusion firmly enough to act on it, but loosely enough to change your mind, because the number describes a long-run property of the procedure, not a guarantee about the one answer sitting in front of you today.

Putting it together

Statistical significance asks how easily pure chance could have faked the difference you’re seeing, and a small p-value means chance is a weak explanation. Sample size drives all of it, so small tests deserve deep suspicion. Always ask the second question too, whether the effect is big enough to matter, and stay alert to the trap of finding false winners when you test or slice many things at once. No promises, no guaranteed lifts. Just the idea explained until it’s yours to use.

If you want more breakdowns like this, plain-English explanations of the metrics and methods that actually run a small business, you can find the rest of them on the Data Research Analysis Collection home page.

Get new guides and videos first — join the Telegram channel.