Fourteen orders. That is what sat underneath the nine percent lift I announced to myself on day six of a free shipping banner test, and it is the number I did not write down at the time because writing it down would have ended the celebration early.
I had run the banner against the old header for less than a week. Eight orders one side, six the other. Nine percent sounds like a result. Eight against six is a Tuesday.
What makes it worth writing up is that I got it wrong while doing everything the guides tell you to do. The split was random. Both versions ran at the same time. I had one metric picked out in advance. Every part of the setup was clean, and I still ended up with a conclusion made of nothing, because I had never decided when I was allowed to look.
Thirty mornings is thirty chances
A test that runs a month and gets checked every morning works out as thirty small tests drawing on one shared pool of data, and you stop at whichever of them comes back the way you wanted.
That framing is the whole idea, and once you have it, the rest is obvious. A running conversion rate wobbles constantly, because early on there is barely anything holding it in place. One extra order moves it a few points. So on any given morning, one version is ahead of the other, and which version that is changes more often than you would guess.
Now give yourself thirty opportunities to declare the leader a winner. You will find one. You would find one if the two versions were byte for byte identical, because the wobble does not care whether there is a real difference underneath it.
The significance number your tool reports assumes you looked once, at a moment fixed before the data existed. I have written separately about what significance actually measures and what a p value does and does not tell you, and both of those pieces make the same assumption without ever saying it out loud. Look thirty times and the number on the screen is answering a question you stopped asking on day one.
A rule that can only finish one way
Read your stopping rule back to yourself in plain words.
“I will stop when the difference looks big enough to trust.” That sentence says the test ends when you see a good result, so the only available ending is a good result. It has no branch where you conclude nothing happened, because you never gave it one.
Compare it to “I will stop on the eighteenth, or at two hundred orders per side, whichever comes second.” That sentence has three endings. One version wins, the other wins, or neither does. It is a worse sentence to live with and a much better one to have written.
Peeking sideways
The version I see most in small businesses has nothing to do with dates.
You check the conversion rate on the new signup form and it is flat. So you check the number of people who reached step two. Also flat. Then average order value, then repeat purchase rate, then the same figures for mobile visitors on their own.
Somewhere down that list, something is up. It is up for the same reason a leader exists on any given morning: you asked eight questions and accepted the first flattering answer.
The defence is dull and it works. Name the single number the change was supposed to move, write it above the test, and let that number decide. Everything you look at afterwards is background for the next idea, and it does not get a vote on this one.
Extending a flat test is the same mistake wearing a suit
Your test reaches its date. It comes back flat. And you think: it is so close, another ten days would settle it.
That is peeking. You saw the result and used it to change the rules. The direction happens to be more data rather than less, which makes it feel careful and responsible, and it is neither.
Run a flat test long enough with permission to keep extending and it will eventually stop being flat, in one direction or the other. You have simply put the finish line on wheels.
Two numbers and a date, written before anything happens
Here is the practical order, most useful first.
Decide the size and the ending before you launch. How many orders, signups or replies go through each side, and the calendar date on which you make the call regardless. Both, because a count on its own leaves you six weeks in and still short with no idea what to do, and a date on its own leaves you calling a test that never gathered enough to say anything.
Sample sizing has its own write up and I will not repeat it here. For stopping purposes the crude version does the job. Pick a per side count that would embarrass you if it were smaller, and let the thing run to it.
The reason this works has nothing to do with the arithmetic. It works because it moves the decision to the only moment when you have no preference. Before launch, you do not yet know which version you are rooting for. On day six, you do, and everything you think from then on is downstream of that.
If you must look, look for damage
Most people cannot leave a test alone, and I include myself. So change the question.
Is the new page loading properly on a phone. Is the payment step firing on both variants. Are the two groups receiving roughly equal traffic. Did a discount code go out on Wednesday to a list that only touches one of them.
All of those are worth checking every day. None of them requires you to know who is ahead. If your tool can hide the conversion column during a run, hide it. Mine cannot, so the running totals go in a file I do not open. It sounds like a child’s trick and it beats willpower comfortably.
Sometimes the honest answer is that you cannot run the test
A shop with sixty checkout visits a week converting at three percent produces about two orders a week, split across two versions. No amount of patience turns that into an answer about a headline.
I have covered elsewhere what it costs to run experiments when you have few users, and the conclusion there applies here in a harder form. Below a certain volume, the correct move is to stop calling it a test at all. Make the change because you think it is better, write down in one line why you thought so, and watch the trend over the next two months with no claim of proof attached.
That is a judgement call, and a judgement call you have labelled as one is worth more than a test you peeked your way through, because you know exactly what you are holding.
The three early stops that are real
None of this makes early stopping forbidden. There are reasons, and when they show up you act the same hour.
Something is broken. The variant errors in one browser, the form will not submit, tracking fires on one side only. Stop, fix, restart from zero, and throw the collected data out rather than trying to rescue a portion of it.
A version is hurting customers. If the new checkout is confusing people into ordering three of something, or the new email is generating complaints, pull it. No finding is worth annoying people who already pay you.
The context moved. A mention somewhere sent a flood of unusual traffic. A supplier failed mid week. A competitor ran a large sale. The test is now measuring a fortnight that describes nothing, and finishing it will not repair that.
Notice what those share. Not one of them is about who is winning. You could act on every one of them with the results hidden from you entirely, which gives you the test for whether any early stop is legitimate: would you still be stopping if you could not see the numbers? If no, you are peeking, whatever you have decided to call it.
Inconclusive beats a peeked winner
People push back on this, so I will put it flatly. A test you ran properly that came back inconclusive is worth more than a winner you arrived at by looking.
The inconclusive one told you something true. Whatever effect exists is smaller than you can detect with the customers you have, which means you should stop spending weeks on that class of question and go and try something larger. That is a real finding with a real consequence.
The peeked winner told you nothing and left a change on your site that you now believe in. You will defend it in six months. You will build the next decision on top of it. Wrong knowledge is more expensive than no knowledge because you keep acting on it.
The one I extended three times
I ran a two week test on a newsletter signup form, one version asking for a first name and one asking only for an email.
Flat at two weeks. I gave it another week, because it felt like it was nearly there. Flat again. Another week. On day twenty six the shorter form pulled ahead and I called it, and the piece I did not think about was that day twenty six followed a weekend where I had posted the link somewhere that sends unusually casual traffic.
I never repeated it, which is the part I am least happy about. I acted on it, put the short form everywhere, and I still do not know whether it was right. What I know is that the process which produced it could not have told me either way.
What I actually do now
The two numbers and the date go at the top of the same note where the results go, before a single visitor is counted. So every time I open the file, the first thing I read is a promise made when I knew nothing.
I break it about one test in five, which is a long way from where I started, and on the ones I break I can at least watch myself doing it. More on the numbers a small business actually runs on is on the home page.
Get new guides and videos first — join the Telegram channel.