Start with a decision, not a p value
This workflow is for product, growth, marketing, research, and operations teams that need a defensible A/B readout. It helps define a metric, size evidence, select a test, and review the result. It does not allocate traffic, validate telemetry, monitor guardrails, or decide what the business should ship.
Plan and test
Record the baseline, minimum useful effect, alpha or confidence level, target power, traffic, and stopping rule before the final readout. Match the test to the data: two independent conversion rates use a two-proportion Z test; matched measurements use a paired test; independent continuous groups with unequal variances use Welch.
Review and accept
Read absolute effect, relative lift, confidence interval, p value, sample size, and warnings together. Compare the interval with the practical threshold and check data quality and guardrails. Statistical significance describes compatibility with a null model; it does not establish size, safety, profitability, or business success.