How to Run an A/B Test That Actually Tells You Something
The statistical and operational discipline that separates a test result you can act on from a coin flip that got a headline written about it.
Somewhere between 60 and 80% of A/B tests run by marketing teams produce results that wouldn’t survive a rerun. Not because the tools are broken, but because the test was set up to produce an answer, not to produce the truth — stopped early when the numbers looked good, run on too little traffic to mean anything, or measuring a change too small to matter even if it were real.
Decide what a win looks like before you look at any data
The test needs a pre-registered hypothesis, a primary metric, and a minimum detectable effect written down before launch — not decided retroactively once results start rolling in. Write a single sentence: “We believe changing X will increase Y by at least Z%, because [reason].” That sentence does three jobs at once:
- Forces you to pick one primary metric instead of scanning five metrics after the fact and reporting whichever one moved
- Sets the minimum effect size worth caring about, which determines how much traffic you actually need
- Gives you a documented hypothesis to check against reality, rather than a story invented after seeing which variant won
Without this step, teams fall into “metric shopping” — running a test, checking conversion rate, checking average order value, checking bounce rate, checking time on page, and declaring victory on whichever one happened to move in the right direction. With five metrics checked at 95% confidence, you have roughly a 23% chance one of them shows “significance” by pure chance, even with a completely ineffective change.
Calculate sample size before launch, not during
The most common test-invalidating mistake is peeking at results daily and stopping the moment the variant pulls ahead. Statistical significance calculated mid-test, on a partial sample, with the intent to stop as soon as a threshold is crossed, inflates false positive rates dramatically — a test “monitored” this way and stopped at the first sign of significance can have a 25-30% false positive rate instead of the intended 5%.
Fix it with two decisions made before launch:
- Calculate required sample size upfront, using your baseline conversion rate and the minimum effect size from your hypothesis. Free calculators exist for this — the point isn’t the tool, it’s committing to a number before you start.
- Pick a fixed test duration and stick to it, covering at least one full business cycle (usually 2+ weeks) to account for day-of-week variation, and don’t call the test early even if it looks decisive on day four.
If your traffic volume can’t reach the required sample size in a reasonable window, that’s real information too — it means the test as designed isn’t going to produce a trustworthy answer, and you should either test a bigger change (larger effects need less traffic to detect) or test a page with more volume instead.
A worked example: sizing a test from scratch
Say your product page converts at 3.2% and you want to detect a 15% relative lift (moving to 3.68%) at 95% confidence with 80% power — the standard defaults for a marketing test. Plugging those into a standard two-proportion sample size calculation gets you to roughly 17,000 visitors per variant, or about 34,000 total. If the page gets 4,000 visitors a week, that’s a little over 8 weeks to reach significance — probably too long to be practical, which is exactly the kind of thing you want to know on day one rather than day thirty when someone asks why the test still isn’t done.
Now change one input: instead of trying to detect a 15% relative lift, you only care about detecting a 30% lift (a much bigger, riskier change — say, replacing the entire page layout rather than swapping a headline). The required sample per variant drops to roughly 4,300, or about 8,600 total — a 2-week test at the same traffic level. This is the practical lesson buried in the math: small changes need much bigger samples to validate than big changes do, which is why testing “should this button be orange or blue” on a page with modest traffic is often a waste of a testing slot that could have gone to a change big enough to actually move the needle in a reasonable timeframe.
Run this calculation before you pick what to test, not after — it tells you which ideas on your roadmap are actually testable within your traffic constraints and which ones need to wait for more volume or get redesigned as bigger swings.
Sample ratio mismatch: the failure mode nobody checks for
Even a well-designed test can be silently invalidated by a bug in how traffic gets split. If you intended a 50/50 split and your test ran with 5,200 visitors in variant A and 4,650 in variant B, that’s not necessarily meaningless — but a large enough deviation from the intended ratio (a sample ratio mismatch, or SRM) is a strong signal that something in your randomization, redirect logic, or bot filtering is broken, and any result from that test should be thrown out regardless of how clean the conversion lift looks.
Common causes: a caching layer serving one variant to repeat visitors inconsistently, a redirect that adds latency and causes slower-connection users to drop out of one variant disproportionately, or bot traffic that gets bucketed into one variant more than the other because of how your targeting rules are set up. Run a simple chi-square test on your traffic split before you even look at the conversion numbers — most testing platforms surface this automatically, and if they don’t, it’s a five-minute manual check that catches a category of false result that no amount of sample size or test duration will fix, because the problem isn’t statistical power, it’s that the two groups were never actually comparable to begin with.
Segment your traffic sources before you trust an aggregate result
An aggregate “variant B wins” conclusion can hide the fact that variant B won because of one traffic source and lost or was flat everywhere else. This matters enormously for a business running multiple acquisition channels, because a page redesign that helps paid search traffic can simultaneously hurt organic traffic with different intent and expectations.
Before declaring a winner, check the result holds across:
- Your two or three largest traffic sources independently
- New visitors versus returning visitors
- Mobile versus desktop, since layout changes frequently affect these very differently
A result that only holds in the aggregate, and reverses or disappears in every meaningful segment, isn’t a real result — it’s Simpson’s paradox waiting to embarrass you in a follow-up test.
Test one variable, or be explicit that you’re not
Multivariate creep is the second-biggest killer of usable test results. A team sets out to “test the new hero section” and along the way also changes the headline, the CTA color, and the social proof placement. If it wins, you have no idea which change drove it — and worse, if you try to “keep the winning parts” and ship a version that combines elements from both the control and the variant, you’ve shipped something that was never actually tested.
If you genuinely want to test multiple variables at once, that’s a legitimate multivariate test design — but it requires meaningfully more traffic to reach significance on each variable’s individual contribution, and it needs to be set up as such deliberately, not stumbled into because the design team made three changes in one PR.
Watch novelty effects before trusting week-one results
A new page layout, a new CTA, a new interaction pattern — all of these can produce an artificial lift in the first days of a test simply because it’s different and draws attention, independent of whether it’s actually better. This effect fades, typically within one to two weeks, and a test that looked like a 12% winner on day three can flatten to a 2% winner — or a loser — by day fourteen.
The practical fix is the same fixed-duration discipline from earlier: don’t call a winner before your pre-committed test duration ends, and if the change is a major UX shift rather than a minor copy or color tweak, consider running it slightly longer than the calculated minimum specifically to let novelty effects wash out.
Prioritizing which tests to run first
Most teams have a backlog of ten or fifteen test ideas and no principled way to decide order. A simple scoring framework fixes this: for each idea, rate potential impact (how much of the funnel does this touch, and how far from optimal does it currently look), confidence (do you have data — heatmaps, session recordings, support tickets, past test results — actually suggesting this will move the metric, or is it a hunch), and ease (how much design and engineering work does it take to ship a valid version). Multiply or average the three on a 1-10 scale and rank the backlog by the composite score.
The point isn’t the precision of the score — it’s forcing an explicit conversation about why a checkout-flow test that touches every paying customer is sitting behind a homepage hero test that touches a much smaller, earlier-funnel audience with less commercial intent. Teams that skip this step tend to test whatever the loudest stakeholder wants tested that quarter, which is a fine way to keep people happy and a poor way to build a testing program that compounds.
A related sequencing rule: don’t run two tests that touch the same page or the same part of the funnel simultaneously unless your platform supports mutually exclusive test groups. Overlapping tests contaminate each other’s results in ways that are hard to detect after the fact — a pricing page test and a checkout button test running at the same time can each attribute the other’s effect to themselves.
Validate the win after you ship it
A test result at 95% confidence still has up to a 5% chance of being a false positive by design, and real-world implementation rarely matches the test environment exactly — a variant tested to 100% of traffic for two weeks can behave differently once it’s permanently live for six months across seasons, promotions, and audience mix shifts that the original test window didn’t capture. Before fully retiring the control, hold back a small percentage of traffic — 5 to 10% — on the old version for another two to four weeks after rollout and confirm the lift persists at a similar magnitude.
This holdback step catches two things a single test can miss: a result that was a false positive despite passing your significance threshold, and a result that was real but decayed once novelty wore off completely or once a downstream process (support team scripts, sales follow-up, email sequences) hadn’t yet adjusted to the new experience. Skipping this step means your test log accumulates “wins” that were never actually re-confirmed against reality, which quietly erodes trust in the whole program once someone eventually notices a metric that was supposed to be up isn’t.
Build a test log, not just a results deck
Every finished test should get logged somewhere durable — hypothesis, sample size, duration, result, and a plain-language note on what you’d do differently. Most teams skip this and lose the institutional memory within two quarters, which means the same “let’s just try making the CTA button bigger” idea gets re-tested by a new hire eighteen months later with no memory that it was already tried and produced a null result.
A simple log with five columns solves this:
- What was tested and why
- Primary metric and result (with confidence interval, not just a point estimate)
- Whether it shipped, and to what percentage of traffic
- What surprised you
- What it implies for the next test
Over a year, this log becomes more valuable than any individual test result — it’s the closest thing your team has to an accumulated understanding of what actually moves your specific audience, on your specific product, at your specific price point.
Treat a null result as data, not a failure
A test that shows no significant difference isn’t a wasted quarter — it’s confirmation that the variable you tested isn’t the lever you thought it was, which redirects effort toward hypotheses more likely to matter. Teams that only report “wins” to leadership create an incentive to keep testing trivial changes that are likely to show some effect, rather than the harder, higher-variance tests that could produce a real step change in conversion. Report null results with the same rigor as wins, and the whole testing program gets more honest — and more useful — over time.
