TLDR: You did the hard work to trust your data and learn to read it. But analysis only ever explains what already happened, and no single method gives you the truth on its own. You get there by triangulating three things: what your numbers say, what your customers say, and what a controlled test can actually prove. Most businesses run on the first two and skip the third. This series is about what that missing leg is worth, and what it costs to leave it out.
The last series was about earning trust in your data. Getting the tracking right, mapping how customers actually move, defining your metrics so a number means what you think it means. If you did that work, you can now look at your reporting and believe it. That is a real milestone, and most businesses never reach it.
But trusting your data buys you less than it feels like it should. Everything your analytics can tell you is about the past. What happened, to whom, in what order. On a good day it can tell you what tends to go with what. What it cannot tell you is the one thing you actually want to know the moment you are about to spend money. If I change this, will it work.
The argument you have already heard
There are two loud camps in this, and each will tell you the other is a fool.
On one side, the data-driven zealots. If you did not test it, it did not happen. Test everything, trust nothing else, the number is the only truth. On the other, the sceptics, and some are blunt about it. One team calls A/B testing bullshit, a waste of time that promotes bad design. They have real points. You cannot test your way to a bold new product, only to a slightly better version of the one you already have, so taste and conviction have to lead. And on a smaller site the honest statistics are unkind, a single test can run for months before it tells you anything you can trust, by which point the world has moved on.
I do not sit in either camp. The zealots are wrong that you should test everything. The refuseniks are wrong that you can therefore skip it. What both miss is that no single method tells you the truth on its own.
You reach a decision you can trust by triangulating three things. Your quantitative data tells you what is happening. Your qualitative research, the talking to customers and watching sessions, tells you why. And an experiment tells you whether the change you make off the back of those two actually causes the result you expected. Ronny Kohavi, who built experimentation at Microsoft, Amazon and Airbnb, is blunt that a controlled experiment is the only practical way to establish that cause, everything else is an informed guess. Jonny Longden, who ran experimentation at Sky and Visa, comes at it from the other side and argues that analytics and experimentation get wrongly split into separate teams when they are “just parts of the same endeavour”. Any one of the three alone will mislead you. Together they check each other.

The triangulation of truth: quantitative data tells you what, qualitative research tells you why, and an experiment tells you whether your change causes it, meeting at decisions you can trust
So no, I am not telling you to test everything. I am telling you to test the things that matter, in a way you can actually read, as one leg of that triangle. Most businesses already do two of the three. They have the numbers, they have some feel for their customers, and then they jump straight from insight to shipping with no test in between. That missing leg is the whole subject of this series.
We are much worse at picking winners than we think
Here is the part that should make all of us humble. The teams with the most data, the best analysts and the most reason to be good at this still cannot reliably pick a winner.
At Microsoft’s experimentation platform, only about a third of the ideas that got built and tested actually improved the metric they were designed to improve. Two out of three did nothing or made things worse. In the same write-up Kohavi collects similar admissions from others: at Google roughly one in ten experiments led to a change worth keeping, and Netflix has said it considers ninety percent of what it tries to be wrong.
These are not junior teams guessing in the dark. They are the best resourced product organisations on the planet, and they are wrong most of the time. So when a redesign gets signed off because it obviously looks better, or a feature ships because everyone in the room liked it, be honest about the odds. You are not more reliable at this than Netflix. Nobody is. The difference is not talent, it is priorities. Netflix succeeds because it treats learning which decisions drive its outcomes as the actual work. Most businesses treat the outputs as the work, the shipped features and the launched redesigns, and never stop to learn which of them changed anything.
The most expensive version of this is the big rebuild
You have probably seen this one, maybe lived it. A business decides its website or its product needs a refresh. It spends months and a serious amount of money redesigning everything at once, the navigation, the layout, the checkout, the copy, and ships the new version in a single release. Then the numbers come in flat. Or down.
Now what. You changed a hundred things together, so you have no way to tell which of them helped, which hurt, and which did nothing. The good decisions and the bad ones are baked into the same release and you cannot pull them apart. You cannot roll back the mistake because you cannot find the mistake. You spent all that money and what you bought, sitting on top of a new website, was a mystery.
The rebuild feels like the bold, decisive move. It is actually the least decisive thing you can do, because it quietly destroys your ability to know what worked.
Guessing costs you even when nothing looks broken
It is tempting to think this only bites when things go wrong. It does not. When the business is growing you can afford not to know why, the same way you can ignore an odd noise in a car that still drives fine. The bill just arrives later. It arrives the quarter the bottom line slips under a stack of shipped changes and nobody can say which one did it. You can see the drop. You cannot see the cause. That is the real cost of never testing, and it is where the next article starts.
Experimentation is bigger than A/B testing
It is easy to hear experimentation and picture an A/B test, two versions of a page, a big enough audience, a statistics engine calling a winner. That is the most rigorous form of it, and where you have the traffic it is worth reaching for. But it is not the thing itself.
Experimentation is a way of learning. It is the scientific method wearing work clothes. You form a view about what will happen and why. You change one thing on purpose, watch what it does and keep what you learn whether the change wins or loses. Strip away the tooling and that is the whole habit: change less at a time, decide in advance what you expect and keep something to compare against, so that when a number moves you know what moved it.
That is why the small-business version still works, and why your qualitative and quantitative work can be experiments in their own right, not just inputs to one. You might never have the traffic for a clean A/B test, and that is fine. Approach a user study in a fair and structured way, with a view of what you expect before you watch and an honest account of its limits, and can be as much an experiment as a split test. So is a careful before-and-after on your own analytics, as long as you are straight about what else moved at the same time. The tool is not what makes it an experiment. The discipline is: a clear hypothesis, a fair test of it and honesty about the caveats.
What you are really choosing between these methods is risk appetite. A powered A/B test buys you the most certainty and the fewest caveats. A user study or a rough before-and-after buys you less certainty and more caveats, but it is often the only thing you can actually run, and a careful version of it beats a confident guess every time. You give up some statistical certainty. You keep the part that matters, a legible line between what you did and what changed.
That habit is the leg of the triangle most businesses are missing, and the thread that runs through everything that follows.
What this series is about
This is not a how-to. I am not going to teach you to run an A/B test. It is about what experimentation is worth to a business, and what quietly leaks out of one that never does it. The wrong bets that fail invisibly. The upside nobody thought to look for. The decisions handed to the loudest person in the room. The learning you pay for once and then throw away.
And it lands somewhere you can actually use. Not a programme you need to hire a team to run, but the smallest version of this that still pays, the kind you can start part time on the business you already have.
First, the big rebuild that taught you nothing, told on themselves by a company famous for building on conviction rather than tests.
Ready to learn what works?
If your reporting is finally trustworthy and you are about to spend real money acting on it, that is exactly the moment to make sure you can tell whether it worked. I help businesses find where they are guessing and build the habit of testing before they bet.