Methods Lab · Causal inference and experimentation

A/B testing

A/B testing is a method within causal inference and experimentation. Experimental and quasi-experimental methods designed to estimate what changed because of an intervention rather than alongside it. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Direct definition

A/B testing is a method within causal inference and experimentation. Experimental and quasi-experimental methods designed to estimate what changed because of an intervention rather than alongside it. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

When A/B testing is used

Experimental and quasi-experimental methods designed to estimate what changed because of an intervention rather than alongside it. It is most useful when the decision owner can state what different findings would cause the organisation to do differently.

Business and research questions it can answer

Experimental and quasi-experimental methods designed to estimate what changed because of an intervention rather than alongside it.

Translate the question into an observable measure, comparison or decision rule before collecting evidence.

What assumptions must hold?

Translate the question into an observable measure, comparison or decision rule before collecting evidence.

How will the result be validated?

Translate the question into an observable measure, comparison or decision rule before collecting evidence.

What decision changes if the result moves?

Translate the question into an observable measure, comparison or decision rule before collecting evidence.

How the design should work

Define the outcome, predictors or inputs, scale of measurement, dependence structure, missingness, sampling process and intended inference. Pre-specify transformations, tuning, validation and the decision rule where the analysis is confirmatory.

Frame the estimand or decision

Define the population, unit, outcome, alternatives, time horizon and consequence of error.

Specify the evidence

Choose sources, measures and comparisons that can distinguish the competing explanations.

Protect quality

Predefine recruitment, exclusions, missing-data rules, assumptions, validation and audit trail.

Translate the result

Report effect, uncertainty, limitations and the action that follows each plausible result.

Sample-size and data considerations

Adequacy depends on model complexity, outcome prevalence, number of parameters, signal-to-noise ratio and validation strategy. Prefer simulation, power analysis or stability testing over a universal observations-per-variable rule.

Analysis and interpretation

Report assumptions, preprocessing, model specification, uncertainty, validation and sensitivity checks. Separate in-sample fit from out-of-sample performance and statistical visibility from decision relevance.

Practical example

A team considering a/b testing would begin with a clearly defined outcome and a held-out validation plan. It would compare the method with a simpler benchmark, inspect errors by meaningful subgroup and translate the result into a decision threshold rather than presenting a software output as proof.

Advantages

  • Creates a reproducible analytical structure when assumptions are explicit.
  • Can reveal patterns that are difficult to see in raw tables.
  • Supports sensitivity analysis and comparison with simpler benchmarks.

Limitations

  • Results depend on model assumptions and data-generating conditions.
  • Complexity can create false confidence when validation is weak.
  • A technically good model may still be irrelevant to the business decision.

Common mistakes

  • Choosing the technique because it is sophisticated rather than because it matches the estimand and data structure.
  • Ignoring assumptions, preprocessing choices or dependence in the data.
  • Reporting fit or significance without uncertainty, validation and practical effect size.
  • Interpreting association or prediction as a causal effect without a causal design.

Practical checklist

Frequently asked questions

What is A/B testing in simple terms?

A/B testing is a method within causal inference and experimentation. Experimental and quasi-experimental methods designed to estimate what changed because of an intervention rather than alongside it. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

How much data is needed?

Adequacy depends on model complexity, outcome prevalence, number of parameters, signal-to-noise ratio and validation strategy. Prefer simulation, power analysis or stability testing over a universal observations-per-variable rule.

What is the biggest interpretation risk?

Choosing the technique because it is sophisticated rather than because it matches the estimand and data structure.

Can Ninth Atlas apply this to a live study?

Yes. The engagement would begin with the business decision and evidence gap, then specify the method, sample, quality controls, analysis and decision output.

Related terms

Editorial and methodology note

This reference is written by Ninth Atlas as a decision-oriented explainer. It separates definition, design, analysis and limitations so a method is not mistaken for an answer. Final study specifications should be reviewed against the actual population, evidence, risk and regulatory context.

Sources and further reading

Need the method applied to a live decision?

Share the market, customer, product or investment question that needs a defensible evidence design.

Brief Ninth Atlas