Opens in a new tab
3 4 5 A B C D E F G H I J K L M N O P Q R S T U V W X Y Z

What is A/B Testing

A/B TestingDefinition: A/B testing is a controlled experiment that compares two versions of an element by randomly assigning equivalent units, such as users, sessions, or accounts. Version A normally serves as the control, while version B contains the change being evaluated. Both are observed during the same period with a metric defined in advance to estimate whether detected differences can be attributed to the change rather than ordinary traffic variation.

How A/B Testing Works

The experiment divides an eligible population into two groups. Each unit should retain its assigned version for the relevant period so that the same person does not alternate between A and B. The comparison can apply to a landing page, form, message, price, purchase flow, or digital product feature.

Before the test starts, a hypothesis connects a change to a measurable effect. For example, simplifying a form might increase the proportion of completed submissions. The objective is not to decide which design people like more, but to estimate the effect of an intervention on a specific metric.

A test commonly includes these components:

  • Control and variant: the current reference and the alternative compared with it.
  • Assignment unit: the user, account, session, device, or other entity distributed between groups.
  • Primary metric: the outcome used to evaluate the hypothesis, such as a conversion rate, revenue per user, or task completion.
  • Guardrail metrics: indicators that should not deteriorate, such as errors, cancellations, returns, or loading time.
  • Population and exposure: the conditions that determine who enters the experiment and when a unit is considered to have seen the variant.

Randomization aims to balance prior differences between groups. It does not eliminate instrumentation errors, incomparable traffic, or external effects, so assignment and data collection must also be checked.

Experiment Design

The required sample size is not determined by a universal number of visits. It depends on the baseline rate, variability, minimum effect of interest, statistical power, and acceptable error level. A small change usually needs more observations than a large one. Duration should also cover relevant business cycles without being extended arbitrarily.

The plan is defined before results are observed. It should specify:

  • The hypothesis and the decision that could change.
  • The primary metric and its formula.
  • The assignment unit and included population.
  • The minimum detectable effect and planned sample size.
  • The duration, stopping rules, and exclusions.
  • The guardrail metrics and quality checks.

Repeatedly inspecting results and stopping as soon as one version appears to win raises the risk of false conclusions unless the method supports sequential analysis. Testing many metrics, variants, or segments and reporting only the favorable combination can also distort interpretation. The GOV.UK guidance on A/B comparative studies recommends starting with a hypothesis, deciding the sample size when necessary, and comparing an intervention with a control.

Process for Running and Validating a Test

A reproducible procedure separates preparation, execution, and interpretation:

  1. Define the problem: document current behavior and the evidence that justifies an experiment.
  2. State the hypothesis: specify the change, population, expected effect, and metric that will represent it.
  3. Prepare the versions: alter only what is necessary so the result can be linked to an identifiable difference.
  4. Instrument and test: validate assignment, exposure, events, formulas, devices, and journeys before using the data.
  5. Run according to plan: maintain the planned rules and monitor failures, harmful effects, and imbalances between groups.
  6. Analyze the outcome: estimate effect size and uncertainty, review guardrail metrics, and check sample quality.
  7. Decide and document: adopt, reject, or repeat the variant according to the evidence, preserving dates, configuration, results, and limitations.

An unexpected difference in the proportion of assigned units may reveal failures in identification, redirects, exclusions, or processing. A study of sample ratio mismatch documents causes across assignment, execution, processing, and analysis.

Interpretation, Types, and Limitations

Statistical significance alone does not show that a change matters to the business. The absolute and relative difference, uncertainty interval, implementation cost, and possible adverse effects should also be examined. An inconclusive result does not prove that both versions are identical: the sample may be unable to detect the effect of interest.

An A/B test compares two experiences within one experiment. A multivariate test combines changes to several components to study their effects and interactions, so it generally requires more traffic. A split URL test may serve two separate pages, while a before and after comparison lacks a concurrent control and is more exposed to seasonality, campaigns, or other changes.

Results can be affected by novelty, learning, contamination between groups, users with multiple devices, interference among participants, or an unrepresentative sample. Segment analyses should be justified and retain adequate size; splitting the data after observing the outcome can produce chance patterns.

Tests affecting indexable pages also require checks of canonicals, redirects, and duration. Google’s website testing guidelines explain how to run variants without using cloaking or turning a temporary test into a permanent state.

A/B testing supports conversion rate optimization when it connects a hypothesis to a verifiable decision. It does not guarantee that one variant will perform the same way for every audience or at every time. Its value depends on experimental design, instrumentation, sample quality, and transparent interpretation and retention of the results.