A/B testing for solo founders: how to run one without fooling yourself

Most advice about A/B testing is written for teams with millions of visitors, and it quietly assumes you have traffic you do not have. A solo founder running a test on a few hundred people a week faces a different problem. It is not how to run the test. It is how to keep from lying to yourself about the result.

I run these on my own businesses, and I have called plenty of them wrong over the years, usually in the same few ways. This is less a lecture on the theory and more the checklist I wish someone had handed me before I shipped a change because a test looked good for two days. None of it needs software. It needs a little discipline about what you decide before you start, and the patience not to peek and pounce.

What an A/B test really is

An A/B test is a simple idea. You show one version of something to one group and a different version to another group at the same time, and you compare how each group behaves. The two versions are usually called A and B, or control and variant. The power of it comes from that phrase “at the same time,” because it means the only systematic difference between the groups is the thing you changed. Any difference in behavior can be pinned on that change rather than on the weather or the day of the week.

Why timing matters so much

People often test the lazy way: run version A for a week and version B the next week, then compare. That is not an A/B test. It is a before and after, and it is badly contaminated. The second week might have had a holiday, a viral mention, a payday, a different mix of traffic. By running both versions in parallel and splitting visitors randomly between them, you cancel all of that out. Same days, same traffic, same everything, except the one thing you’re testing.

Decide the metric before you start

The single most important habit is choosing your success metric before you launch, and writing it down. If you decide after the fact what counts as winning, you will always find some number that went up, because in any test something moves. Pick the one metric that actually matters, the purchase rate, the signup rate, whatever the change was meant to improve, and commit to judging the test on that. Everything else you look at afterward is context, not the verdict.

Pick one change at a time

The temptation is to redesign five things at once because you only have the traffic to run one test this month. Resist it. If you change the headline, the button color, the price, and the photo all together and the variant wins, you’ve learned that some combination of those helped, but not which one, and you can’t carry the lesson forward. A clean test changes one thing so that when it moves the number, you know exactly what did it. It’s slower, but you actually accumulate knowledge instead of noise.

The sample size problem

Here’s the hard truth for small operators. With low traffic, you need the test to run longer than feels comfortable, because small groups swing wildly on chance alone. If 40 people see each version, a difference of a few conversions looks dramatic and means almost nothing. You need enough people through each side that a real effect can separate itself from random luck. I’d rather run one honest test for three weeks than three rushed tests in the same time that each fooled me.

Don’t peek and pounce

The most seductive mistake is watching the test hourly and stopping the moment the variant pulls ahead. Early in any test the numbers bounce around enormously, and if you stop the instant you see the result you were hoping for, you’ll declare winners that are pure noise. Decide your stopping point in advance, either a fixed run time or a target number of people per side, and hold to it. The discipline isn’t glamorous, but it’s the entire difference between testing and fooling yourself.

The winner’s curse

Even a properly run test tends to flatter the winner. The version that happened to get a lucky bounce during the test is more likely to be the one you crown, which means the real world lift after you ship it is usually smaller than the test suggested. This isn’t a reason to distrust testing. It’s a reason to hold your expectations loosely. Expect the improvement to shrink a little once it’s live, and you’ll be pleasantly calibrated rather than disappointed.

When a flat result is a real result

People treat a test that shows no difference as a failure, and it’s the opposite. Learning that a change you were sure about did nothing is genuinely valuable, because it stops you rebuilding your whole site around a hunch that was wrong. A lot of confident redesigns would have been quietly killed by a test that said “these two are the same.” No difference is data. It saves you the effort you were about to spend, and that saved effort is a real return.

Segment the result carefully, later

It’s tempting to slice the result afterward: say the variant lost overall but won for mobile users, so let’s ship it just for them. Be very careful here. The more ways you slice a result, the more likely you are to find a flattering slice by pure chance, the same way flipping enough coins guarantees a streak somewhere. Segment to generate a new idea to test next, never to rescue a verdict from a test that didn’t support it. A slice found after the fact is a hypothesis, not a conclusion.

Not everything is worth testing

A/B testing is a tool for decisions where the stakes and the traffic both justify the wait. Testing the color of a button on a page four people see a day is a waste of the weeks it takes. Some choices are better made with judgment, a quick customer conversation, or plain common sense, and saved for when a test can actually resolve them. Part of using this well is knowing when the honest answer is just to decide and move on, because the test would take longer than the decision is worth.

Keep a record of what you tested

The compounding value of testing comes from memory. Keep a simple log: one row per test, what you changed, what you predicted, what happened. Over a year that log becomes the most honest document about your own business you own, a record of what actually moved your numbers versus what you were certain would. It also stops you re-running tests you already ran and forgot, which every solo operator does at least once. The log is where scattered tests turn into real accumulated knowledge.

What a test can’t tell you

And the honest limit. A test tells you which version performed better for the people in it, over the window you ran it, on the metric you chose. It doesn’t tell you why, it doesn’t promise the effect lasts forever, and it says nothing about the customers you never reached. Treat a win as evidence to act on, not proof of a permanent law. The map is useful precisely because you keep redrawing it, not because any single test settled the question for good.

Randomize properly

The split between your two groups has to be genuinely random, or the whole thing quietly breaks. If version A happens to get shown mostly in the morning and version B mostly at night, you’re no longer testing the change, you’re testing the time of day. The same goes for showing one version to logged-in users and the other to strangers. A good test assigns each visitor to a side by something with no relationship to who they are or when they arrived, so the only real difference between the two groups is the thing you set out to test.

Run across a full cycle

Give the test enough time to cover the natural rhythm of your business, not just enough people. Most small businesses behave differently on weekends than weekdays, and differently at the start of a month than the end. A test that runs only Tuesday to Thursday can miss an effect that only shows up when weekend buyers behave their own way. I try to run any test across at least one full week, and often two, so both versions get exposed to the same complete cycle of how my customers actually behave over time.

Guard against interference

Watch for things that quietly contaminate a test while it runs. A big email blast, a press mention, an influencer post, or a sudden sale can flood one part of your funnel and swamp the difference you’re trying to measure. If something unusual happens mid-test, note it, and be ready to throw the test out rather than trust a result that got hit by an outside shock. The cleanest tests run during ordinary weeks, when nothing dramatic is happening, precisely because ordinary is when your real baseline behavior shows through.

Weigh the cost of being wrong

Before you run a test at all, ask what a wrong answer would cost you. If shipping the losing version by mistake would barely dent anything, you can accept a looser, faster test and move on. If the change touches your pricing or your checkout, where a wrong call quietly bleeds money for months, you want far more evidence before you trust the result. Matching how careful you are to how expensive a mistake would be is the judgment that separates testing as theater from testing as a real decision tool.

Beware the novelty effect

One specific way a variant quietly fools you: when your regulars see something new, some of them click on it just because it’s different, not because it’s actually better, and that curiosity fades fast. A bold redesign can win its very first week on novelty alone and then settle right back to ordinary once the newness has worn off. If a change is aimed at people who see your site again and again, give it long enough for the shine to come off before you trust the lift. The effect you truly want is the one that survives after the novelty is gone, not the temporary spike that came purely from it being unfamiliar.

Recap

An A/B test runs two versions at the same time to isolate one change, and its honesty depends on what you decide before you start. Choose the metric up front, change one thing, give it enough people, and never stop the moment it looks good. Respect the winner’s curse, treat a flat result as real information, and be suspicious of slices you found after the fact. None of this needs to be complicated. It just needs to be done in the same order, every time, whether the result flatters you or not.

For more plain-English breakdowns of the metrics and methods behind running a business on your own numbers, head back to the Data Research Analysis Collection homepage.

Get new guides and videos first — join the Telegram channel.