# A/B Test Significance Calculator

Compute conversion-rate or mean differences, 95%/99% confidence intervals, two-sided p-values (Z-test for proportions, Welch t for continuous metrics), lift, sample size for 80%/90% power, recommended experiment days, plus sequential-testing and Bonferroni multiple-comparison warnings.

> Canonical page: https://elysiatools.com/en/tools/ab-test-significance-calculator

- **Category:** Data Analysis

- **Keywords:** a/b test, ab test, significance, p-value, confidence interval, z-test, welch t-test, conversion rate, sample size, statistical power, bonferroni, sequential testing, lift

## Overview

Decide whether your A/B test result is real signal or noise — with the statistics done correctly.

**Two test types.**
- **Proportion / conversion rate.** Enter successes + visitors for control and treatment. The tool computes conversion rates, relative lift, a pooled Z-test, and the 95% / 99% confidence interval for the difference.
- **Continuous metric (mean & SD).** Enter mean, standard deviation and sample size for each group. The tool runs a Welch t-test (which does not assume equal variance) with Welch–Satterthwaite degrees of freedom, and reports the difference, relative change and confidence interval.

**Beyond the p-value.** A single p-value is not enough to act on. This tool also gives you:
- **Lift** — the relative change, the number stakeholders actually care about.
- **Confidence interval** — the plausible range of the true effect; if it crosses zero the result is not significant.
- **Sample size** — how many users per variant you would need to reliably detect an effect this size at the chosen power (80% / 90%).
- **Recommended experiment days** — based on your daily traffic and split.
- **Bonferroni correction** — when you run multiple comparisons, the effective α shrinks; the tool tells you if the result survives the correction.

**Sequential-testing warning.** Peeking at the p-value mid-experiment and stopping early inflates false positives. The tool warns against this. If you must look repeatedly, use an always-valid sequential method (e.g. mSPRT) or pre-register your sample size and stick to it.

**Reading the verdict.** "Significant" means the data is incompatible with there being no difference at α = 0.05 — it does *not* mean the lift is big or important. Always pair significance with effect size (the lift and CI) before shipping.

## Inputs

- **Test Type** (select)
- **Control: successes** (number): e.g. 1200
- **Control: visitors** (number): e.g. 20000
- **Treatment: successes** (number): e.g. 1320
- **Treatment: visitors** (number): e.g. 20000
- **Control: mean** (number): e.g. 42.5
- **Control: std dev** (number): e.g. 12.0
- **Control: sample size** (number): e.g. 500
- **Treatment: mean** (number): e.g. 45.8
- **Treatment: std dev** (number): e.g. 12.5
- **Treatment: sample size** (number): e.g. 500
- **Significance level α** (select)
- **Statistical power** (select)
- **Baseline conversion (%)** (number): e.g. 6 (for sample-size planning)
- **MDE lift (%)** (number): e.g. 5 (minimum detectable lift)
- **Daily visitors** (number): e.g. 4000 (for days estimate)
- **Traffic split (%)** (select)
- **Number of comparisons** (number): 1 unless running multiple variants
- **Decimal places** (number)

## When to use

- When evaluating conversion rate changes, sign-ups, or click-through rates between a control and a treatment group.
- When comparing continuous metrics like average order value, page load times, or session durations using mean and standard deviation.
- When planning an experiment and calculating the required sample size or recommended duration based on daily traffic and statistical power.

## How it works

- Select the test type: choose Proportion for conversion rates or Continuous for metrics involving means and standard deviations.
- Input the sample sizes, successes, or means and standard deviations for both the control and treatment groups.
- Configure your statistical parameters, including the significance level (alpha), desired power, daily traffic, and the number of comparisons for Bonferroni correction.
- Review the output containing the p-value, relative lift, confidence intervals, sample size recommendations, and statistical significance verdict.

## Use cases

- Verifying if a new checkout flow significantly increases the purchase conversion rate compared to the legacy design.
- Determining if a page speed optimization project successfully reduced average page load times across user sessions.
- Calculating the minimum sample size and runtime required to detect a 5% lift in sign-ups before launching a marketing campaign.

## Frequently asked questions

### What is the difference between the Proportion and Continuous test types?

Proportion tests analyze binary outcomes like conversion rates, while Continuous tests compare averages of numerical values like revenue or duration using a Welch t-test.

### Why does the tool warn against sequential testing or 'peeking'?

Repeatedly checking p-values during an active experiment increases the rate of false positives, which can lead to declaring a false winner.

### What does the Bonferroni correction do?

It adjusts the significance threshold (alpha) when running multiple comparisons to prevent false positives caused by testing multiple variants.

### How is the recommended experiment duration calculated?

It estimates the days needed to reach the required sample size based on your daily traffic volume and traffic split percentage.

### What does it mean if the confidence interval crosses zero?

If the confidence interval contains zero, the difference between the control and treatment groups is not statistically significant.

## Related tools

- [Monte Carlo Simulation Builder](https://elysiatools.com/en/tools/monte-carlo-simulation-builder): Define input distributions (normal/uniform/lognormal/triangular), write a formula, run thousands of trials, and get the output distribution histogram with confidence intervals.
- [RSS / Atom Feed Markdown Summarizer](https://elysiatools.com/en/tools/rss-feed-markdown-summarizer): Fetch an RSS or Atom feed by URL (or paste raw XML), parse it, sort items by publish date, deduplicate, apply a time window, and output a clean Markdown summary
- [AI Markdown Article Translator](https://elysiatools.com/en/tools/ai-markdown-article-translator): Translate full Markdown articles with AI while preserving headings, tables, links, images, and code blocks for direct publication
- [Filename Sanitizer](https://elysiatools.com/en/tools/filename-sanitizer): Clean and sanitize filenames by removing illegal characters for Windows, Linux, and Mac
- [Mean Platelet Volume (MPV) Analyzer](https://elysiatools.com/en/tools/mean-platelet-volume-analyzer): Analyze Mean Platelet Volume (MPV) in fL, an index of platelet production kinetics. Reference range ≈ 7.5–11.5 fL (lab- and method-dependent). Low MPV (< 7.5 fL, small old platelets) indicates suppressed marrow production (aplastic anemia, chemotherapy/radiation, recent platelet transfusion, iron-deficiency regeneration, Wiskott-Aldrich). High MPV (> 11.5 fL, large young platelets) indicates peripheral destruction with compensatory turnover or macrothrombocytopenia (immune thrombocytopenia ITP, post-splenectomy, myeloproliferative neoplasms, severe sepsis in recovery, Bernard-Soulier / MYH9, megaloblastic anemia; elevated MPV after acute MI marks platelet activation and adverse prognosis). Optionally enter the platelet count to refine the differential: low platelets + high MPV → peripheral destruction; low + low → underproduction; high + low → reactive thrombocytosis; high + high → myeloproliferative neoplasm. Pre-analytical caveat: platelets swell in EDTA over ~2 h raising MPV, and values are method-dependent (impedance > optical > immuno) — compare within the same lab and analyzer. Derived from Bath 1996, Leader 2012, Vagdatli 2010, and MDCalc. Not medical advice.
- [Normal Distribution Plotter](https://elysiatools.com/en/tools/normal-distribution-plotter): Plot a normal-distribution bell curve as SVG with z-score markers, interval or tail-probability shading, and 68/95/99.7 standard-deviation bands.
- [Array Grouper](https://elysiatools.com/en/tools/array-grouper): Group array elements based on various criteria such as length, alphabetical order, numeric ranges, custom conditions, and more
- [APACHE II Score Calculator](https://elysiatools.com/en/tools/apache-ii-score): Calculate the APACHE II (Acute Physiology and Chronic Health Evaluation II) ICU severity score. APACHE II = Acute Physiology Score (12 variables, each 0–4 using the worst value in the first 24 h) + Age points (0–6) + Chronic Health points (0/2/5). Range 0–71; higher scores indicate worse prognosis. Oxygenation uses the A-a gradient when FiO₂ ≥ 0.5, otherwise PaO₂. Creatinine points are doubled when acute renal failure is indicated. GCS contributes 15 − GCS. Also reports the base-model predicted mortality: R = 1/(1 + e^(−logit)), logit = −3.517 + 0.146 × score (diagnostic-category weight not included). Thresholds cross-verified against Knaus 1985 (Crit Care Med) and the SFAR scoring table. Not a substitute for clinical judgement. Not medical advice.

## Samples

- [Android Image Processing Java Samples](https://elysiatools.com/en/samples/android-image-processing-java): Android Java image processing examples including reading/saving images, scaling, and format conversion
- [Android Image Processing Kotlin Samples](https://elysiatools.com/en/samples/android-image-processing-kotlin): Android Kotlin image processing examples including reading/saving images, scaling, and format conversion
- [Web Image Processing Python Samples](https://elysiatools.com/en/samples/web-image-processing-python): Web Python image processing examples using PIL/Pillow including reading, saving, resizing, and format conversion
- [Web Image Processing Rust Samples](https://elysiatools.com/en/samples/web-image-processing-rust): Web Rust image processing examples including image read/save, scaling, and format conversion

## Related content

- [Markdown Export, OCR, and Document Conversion Tools](https://elysiatools.com/en/hubs/markdown-convert): Compare Markdown-to-PDF, PDF-to-Markdown, OCR, slide deck export, and structured Markdown conversion tools in one hub for documentation publishing workflows.
- [Markdown Writing and Publishing Tools](https://elysiatools.com/en/hubs/markdown-utility): Compare Markdown formatting, link review, merging, preview, translation, and export tools in one hub for docs, notes, and publishing workflows.
- [Barcode, QR Code, and Label Tools](https://elysiatools.com/en/hubs/barcode-utility): Compare barcode generation, QR code creation and decoding, UPC/EAN validation, WiFi QR sharing, and printable label tools in one hub.
- [Documentation Authoring, Extraction, and Publishing Tools](https://elysiatools.com/en/hubs/documentation-authoring-publishing): Write docs, extract docs from code or PDFs, review Markdown, and export polished documentation in one docs workflow hub.
