In the early 1920s, Muriel Bristol, a biologist visiting an English research facility, refused a cup of tea because the milk had been poured in before the tea. The statistician standing next to her, Ronald Fisher, was skeptical that she could taste the difference between cups made with milk-first versus tea-first; she o began setting up an experiment. He placed eight cups of tea in front of her, each prepared in a different way, and asked her to identify which preparation method was which. Her choices were random; other than tasting each cup, Bristol had no information that could lead her to the correct answer. She got all eight cups right. Fisher later described that experiment and his reasoning for it in his landmark 1935 book The Design of Experiments, using it as the foundational example of the modern controlled randomized experiment with clearly defined hypotheses tested and a predetermined rule for deciding if your result was real or a false positive.
It’s also the entire foundation for what we now know as A/B testing. Show one group of randomly assigned people one version of something and another group the other version, then measure which one actually performs better using math instead of instinct to decide.
A/B testing is not a nice-to-have capability you should pursue if your company grows big enough to have budgets. It’s the only way to know whether the changes you make are actually helping, because our intuition about what will and will not work is incorrect more often than most care to believe.
That last sentence isn’t an opinion. Multiple studies have repeatedly demonstrated that humans are not very good at determining what will or will not be successful. Often not even close. In experiments, Ronny Kohavi ran while working at companies including Amazon, Microsoft, and Airbnb, and has published on the subject extensively, and has seen click-through rates improve by 30% or more from successful testing. But he has also reported that at Microsoft, two-thirds of the ideas they tested actually decreased the metric they were hoping to improve. At Airbnb, 250 ideas were put through controlled experiments. Only 20 of them drove improvement. More than 90% of ideas thought promising enough to test had no measurable effect.
At companies of all sizes. Industry-wide. Ideas that feel perfect probably don’t work.
Again, that’s not my opinion. It’s one of the best-studied and documented results across the field.
Why does this matter? If some of the most experienced teams running experimentation programs at the most sophisticated tech companies in the world are wrong about nine out of ten changes they believe will have a positive impact, how can you trust yourself to get it right? Confidence isn’t just slightly more likely to be wrong than testing. It’s always wrong.
The counterargument people sometimes make is “Sure, but my team is smart. We have a good instinct for what will work.” To which I reply, bright, experienced people fooled by their own optimism work at Dropbox. Not saying that to brag. I’m saying it so you know shouting, “but we’re smarter than them!” isn’t going to get you very far.
A/B tests start with a hypothesis and a number. A specific number you calculated in advance. Not a number you adjust as you go. Before you launch, you need to know how many people need to see each version of your experience before you can trust the result. You need to know that number based on how your metric is performing right now and how big a change you’re realistically trying to measure. “Let’s see how it goes” is amateur experimentation.
The second most common mistake is called peeking. Looking at your results before the test has finished and stopping early if you see something you like. It feels harmless. But research studying exactly this behavior has shown that if you check in on results five times during an experiment, you could be raising your false-positive rate to above 14%, more than triple what’s generally accepted as an absolute maximum (5%). Check in on results daily for a month,h and you’re essentially flipping a coin to see whether you declare a winner. The solution? Make sure you decide your sample size before you begin and don’t look at your results until you have enough data. Not once. You look at it when the test is over.
A third mistake is failing to account for what statisticians call sample ratio mismatch. What happens when your two groups being compared aren’t actually the same size, usually by a small amount due to a technical error causing some visitors to not be properly assigned? It’s common; Kohavi’s research found it happened in around 8% of experiments at Microsoft. Often goes unnoticed. And quietly destroys any test it affects. Always double-check that your control and variant received roughly the same number of visitors. If they don’t, don’t trust your results until you know why.
Another pitfall is failing to act on results, which can happen for any number of reasons. But one of the most commonly cited is something known as the HIPPO problem. Highest Paid Person’s Opinion. Sometimes leadership genuinely has better information than what is available from testing. Sometimes they don’t. Bing tried adding a new feature that testing showed would not improve their metric. They launched it anyway because leadership thought it was strategically important. And had to spend roughly the equivalent of 100 engineers reminiscing about what could have been when they were forced to roll back the change a year later, after additional experiments still failed to show value. The point isn’t that HIPPOs are never right. The point is that delaying or ignoring a negative result burns real time and costs real money.
Know your sample size and your success metric before launching, and write both down. Somewhere you can’t feel tempted to change them later.
Don’t peek. Ever. Let the test reach its predetermined size before making decisions based on its results.
Try to only test one meaningful change at a time, so if something changes, you know what caused it.
Expect to fail most of the time. Your ideas aren’t bad just because 8/10 don’t work. According to industry data, they’re probably still better than average. Even at some of the best companies,s experiments fail more often than they succeed.
Accept negative results. Again. If you’ve tested a negative result correctly, it is valuable data. The whole point of running a test is to learn that your idea doesn’t work before you’ve invested a ton of time and money building ten more ideas that assume it does.
Fisher was having a silly argument over tea. But he was trying to figure out, definitively, whether there really was a difference between how his colleague prepared cups versus how Bristol felt she should. He randomized, measured, and insisted on not fooling himself about the result.
Do the same, and you too can know, truly know, what works and what doesn’t.