Reading A/B Test Results Before You Call Them

Most disputed test results are disputed because the decision was taken before the data was ready. A short checklist applied before announcing a winner prevents the argument and, more importantly, prevents acting on noise.
Check that the test ran long enough
Duration should cover at least one full business cycle, including weekend behaviour and the payment cycle if the site has one. A five-day test on a business-to-business site stops before the people who actually buy have returned to work.
Check the traffic split held
Confirm that the two variants received a comparable share of traffic and a comparable mix of sources and devices. An unbalanced split means the result reflects the difference in the audiences rather than the difference in the design.
Look for the practical effect, not only the statistical one
Statistical significance says a difference is unlikely to be chance. It says nothing about whether the difference is large enough to matter commercially. A two per cent lift on a large volume is worth having; a two per cent lift on a low-volume step may not justify the maintenance cost of the change.
Check the guardrail metrics
Revenue per visitor, refund rate and support contacts should be checked alongside the primary metric. A change that improves completion while quietly increasing returns has not improved the business.
Decide in advance what counts as conclusive
Agree the minimum detectable effect and the sample size before the test starts. Decisions made after seeing the data tend to drift towards whatever the audience already wanted.