Bookmarks    
nPr and nCr    
Random Number    
SD Calculator    
Sample Size    
Percent Error    
Density    
Molarity    
Molar Mass    
Ohm's Law    
Watts to Amps    
Voltage Drop    
Long Division    
Mixed Numbers    
Rounding    
Nth Root    
Exponents    
Half-Life    
Polar Form    
De Moivre    
3D Distance    
Point to Line    
Cross Product    
Determinant    
Sin Cos Tan    
Triangle Area    
Scale Factor    
Sector Area    
Ellipse Area    
Cube Volume    
Box Volume    
Sphere Volume    
Cone Volume    
Pipe Volume    
Time Duration    
Time Card    
Present Value    
Future Value    
Churn Rate    
A/B Test Calc    
SEO Traffic    
Ideal Weight    
Fat Intake    
Child Height    
Golf Handicap    
Heat Index    
Wind Chill    
Dew Point    
Download Time    
kWh to Cost    
AC Size (BTU)    
Heating Costs    
LED Savings    
Trip Gas Cost    
Tire Size    
Solar Output    
Solar Payback    
Battery Size    
Wall Area    
Gravel Needed    
Mortar Mix    
Slope Grade    
Curtain Size    
Soil Needed    
Sod Needed    
Ramp Length    
Blind Size    
Drain Slope    
Board Feet    
Heat Loss    
Furniture Fit    
Moving Boxes    
Plywood Cuts    
Shelf Sag    
   Add
Probability and random number calculators
Independent Events
Independent Events
Two Events Solver
Two Events Solver
Repeated Trials
Repeated Trials
Bayes' Theorem
Bayes' Theorem
Expected Value
Expected Value
Binomial Distribution
Binomial Distribution
nPr and nCr
nPr and nCr
Circular Permutation
Circular Permutation
With Repetition
With Repetition
Random Number
Random Number
Averages and statistics calculators
Average Calculator
Average Calculator
Mean Median Mode
Mean Median Mode
SD Calculator
SD Calculator
Quartiles & IQR
Quartiles & IQR
Frequency Table
Frequency Table
Correlation (r)
Correlation (r)
Normal Probability
Normal Probability
Z-Score Calculator
Z-Score Calculator
Confidence Interval
Confidence Interval
Sample Size
Sample Size
Mark & Recapture
Mark & Recapture
P-Value Calculator
P-Value Calculator
Percentage and ratio calculators
Percentage Calc
Percentage Calc
Percent Change
Percent Change
Percent Difference
Percent Difference
Percent Error
Percent Error
Ratio Calculator
Ratio Calculator
Discount Calculator
Discount Calculator
Sales Tax Calculator
Sales Tax Calculator
Margin Calculator
Margin Calculator
Speed calculators
Speed Calculator
Speed Calculator
Density and concentration calculators
Density
Density
Molarity
Molarity
Molar Mass
Molar Mass
Physics and electricity calculators
Ohm's Law
Ohm's Law
Watts to Amps
Watts to Amps
Resistor Colors
Resistor Colors
Voltage Drop
Voltage Drop
Unit conversion calculators
Weight Converter
Weight Converter
Shoe Size Converter
Shoe Size Converter
Integer and signed number calculators
Long Division
Long Division
LCM Calculator
LCM Calculator
GCF Calculator
GCF Calculator
Integer Calculator
Integer Calculator
Prime Factorization
Prime Factorization
Diophantine Solver
Diophantine Solver
Modulo Calculator
Modulo Calculator
Factor Calculator
Factor Calculator
Roman Numerals
Roman Numerals
Fraction, decimal and rounding calculators
Fraction Calculator
Fraction Calculator
Mixed Numbers
Mixed Numbers
Simplify Fractions
Simplify Fractions
Fraction to Decimal
Fraction to Decimal
Decimal to Fraction
Decimal to Fraction
Rounding
Rounding
Equation and inequality calculators
Linear Equation
Linear Equation
Linear Systems
Linear Systems
Quadratic Formula
Quadratic Formula
Absolute Value
Absolute Value
Quadratic Inequality
Quadratic Inequality
Polynomial calculators
Binomial Theorem
Binomial Theorem
Square root and nth root calculators
Simplify Radicals
Simplify Radicals
Nth Root
Nth Root
Exponent and logarithm calculators
Exponents
Exponents
Log Calculator
Log Calculator
Number of Digits
Number of Digits
Scientific Notation
Scientific Notation
Sci. Notation Math
Sci. Notation Math
Half-Life
Half-Life
Complex number calculators
Complex Numbers
Complex Numbers
Polar Form
Polar Form
De Moivre
De Moivre
Function and graph calculators
Slope Calculator
Slope Calculator
Linear Function
Linear Function
Direct & Inverse Variation
Direct & Inverse Variation
y = ax² Calculator
y = ax² Calculator
Distance Formula
Distance Formula
3D Distance
3D Distance
Section Formula
Section Formula
Point to Line
Point to Line
Lat/Long Distance
Lat/Long Distance
Complete the Square
Complete the Square
Circle Equation
Circle Equation
Conic Sections
Conic Sections
Polar Coordinates
Polar Coordinates
Sequence calculators
Arithmetic Sequence
Arithmetic Sequence
Geometric Sequence
Geometric Sequence
Fibonacci Sequence
Fibonacci Sequence
Recurrence Relation
Recurrence Relation
Vector calculators
Vector Calculator
Vector Calculator
Cross Product
Cross Product
Matrix calculators
Matrix Calculator
Matrix Calculator
Determinant
Determinant
Inverse Matrix
Inverse Matrix
Plane geometry calculators
Sin Cos Tan
Sin Cos Tan
Degrees ⇔ Radians
Degrees ⇔ Radians
a sin θ + b cos θ
a sin θ + b cos θ
Triangle Solver
Triangle Solver
Triangle Area
Triangle Area
Right Triangle
Right Triangle
Pythagorean Theorem
Pythagorean Theorem
Polygon Angles
Polygon Angles
Scale Factor
Scale Factor
Parallel Lines
Parallel Lines
Rectangle Area
Rectangle Area
Parallelogram Area
Parallelogram Area
Trapezoid Area
Trapezoid Area
Circle Calculator
Circle Calculator
Sector Area
Sector Area
Inscribed Angle
Inscribed Angle
Ellipse Area
Ellipse Area
Solid geometry calculators
Cube Volume
Cube Volume
Cube Surface Area
Cube Surface Area
Box Volume
Box Volume
Box Surface Area
Box Surface Area
Cylinder Volume
Cylinder Volume
Cylinder Surface
Cylinder Surface
Sphere Volume
Sphere Volume
Sphere Surface
Sphere Surface
Spherical Cap Volume
Spherical Cap Volume
Cap Surface Area
Cap Surface Area
Ellipsoid Volume
Ellipsoid Volume
Ellipsoid Surface
Ellipsoid Surface
Pyramid Volume
Pyramid Volume
Pyramid Surface
Pyramid Surface
Cone Volume
Cone Volume
Cone Surface Area
Cone Surface Area
Frustum Volume
Frustum Volume
Frustum Surface Area
Frustum Surface Area
Pipe Volume
Pipe Volume
Capsule Volume
Capsule Volume
Capsule Surface Area
Capsule Surface Area
Date and time calculators
Age Calculator
Age Calculator
Days Between Dates
Days Between Dates
Date Calculator
Date Calculator
Hours From Now
Hours From Now
Day of the Week
Day of the Week
Time Calculator
Time Calculator
Time Zone Converter
Time Zone Converter
Hours Calculator
Hours Calculator
Time Duration
Time Duration
Time Card
Time Card
Finance and economics calculators
Compound Interest
Compound Interest
Simple Interest
Simple Interest
Interest Calculator
Interest Calculator
TVM Calculator
TVM Calculator
Present Value
Present Value
Future Value
Future Value
ROI Calculator
ROI Calculator
IRR Calculator
IRR Calculator
Payback Period
Payback Period
Average Return
Average Return
GDP Calculator
GDP Calculator
Web marketing and ad metric calculators
CTR Calculator
CTR Calculator
Conversion Rate
Conversion Rate
CPC, CPM & CPA
CPC, CPM & CPA
ROAS Calculator
ROAS Calculator
Break-Even CPA
Break-Even CPA
LTV Calculator
LTV Calculator
CAC Calculator
CAC Calculator
Churn Rate
Churn Rate
A/B Test Calc
A/B Test Calc
A/B Sample Size
A/B Sample Size
SEO Traffic
SEO Traffic
Break-Even Point
Break-Even Point
Markup vs. Margin
Markup vs. Margin
CAGR Calculator
CAGR Calculator
Health and fitness calculators
BMI Calculator
BMI Calculator
Sleep Calculator
Sleep Calculator
Calorie Calculator
Calorie Calculator
BMR Calculator
BMR Calculator
TDEE Calculator
TDEE Calculator
Ideal Weight
Ideal Weight
Body Fat Calculator
Body Fat Calculator
Lean Body Mass
Lean Body Mass
Calories Burned
Calories Burned
Protein Intake
Protein Intake
Macro Calculator
Macro Calculator
Carb Calculator
Carb Calculator
Fat Intake
Fat Intake
Child Height
Child Height
Sports calculators
Golf Handicap
Golf Handicap
Pace Calculator
Pace Calculator
1RM Calculator
1RM Calculator
Target Heart Rate
Target Heart Rate
Weather calculators
Heat Index
Heat Index
Wind Chill
Wind Chill
Dew Point
Dew Point
Computer calculators
Base Converter
Base Converter
Subnet Calculator
Subnet Calculator
Download Time
Download Time
Household energy and budget calculators
Electricity Cost
Electricity Cost
kWh to Cost
kWh to Cost
Yearly kWh to Cost
Yearly kWh to Cost
AC Size (BTU)
AC Size (BTU)
AC Running Cost
AC Running Cost
Heating Costs
Heating Costs
Gas vs Electric
Gas vs Electric
LED Savings
LED Savings
Salary Calculator
Salary Calculator
Budget Calculator
Budget Calculator
Car calculators
Trip Gas Cost
Trip Gas Cost
EV Charging Cost
EV Charging Cost
EV vs Gas Cost
EV vs Gas Cost
MPG Calculator
MPG Calculator
Tire Size
Tire Size
Solar power and battery calculators
Solar Output
Solar Output
Solar Panel Count
Solar Panel Count
Solar Payback
Solar Payback
Battery Size
Battery Size
Home and DIY calculators
Tile Calculator
Tile Calculator
Stair Calculator
Stair Calculator
Concrete Volume
Concrete Volume
Wall Area
Wall Area
Wallpaper Rolls
Wallpaper Rolls
Paint Calculator
Paint Calculator
Flooring Needed
Flooring Needed
Exterior Walls
Exterior Walls
Gravel Needed
Gravel Needed
Mortar Mix
Mortar Mix
Slope Grade
Slope Grade
Lumber Cut List
Lumber Cut List
Lot Coverage/FAR
Lot Coverage/FAR
Sheet Vinyl Roll
Sheet Vinyl Roll
Insulation Needed
Insulation Needed
Curtain Size
Curtain Size
TV Size & Distance
TV Size & Distance
Soil Needed
Soil Needed
Sod Needed
Sod Needed
Block Calculator
Block Calculator
Brick Calculator
Brick Calculator
Deck Materials
Deck Materials
Ramp Length
Ramp Length
Pilot Hole Size
Pilot Hole Size
Room Ventilation
Room Ventilation
Paint Thinning
Paint Thinning
Baseboard & Trim
Baseboard & Trim
Blind Size
Blind Size
Picture Hanging
Picture Hanging
Drain Slope
Drain Slope
Screw Calculator
Screw Calculator
Board Feet
Board Feet
Fence Calculator
Fence Calculator
Wood Shrinkage
Wood Shrinkage
Caulk Calculator
Caulk Calculator
Heat Loss
Heat Loss
Furniture Fit
Furniture Fit
Moving Boxes
Moving Boxes
Storage Capacity
Storage Capacity
Plywood Cuts
Plywood Cuts
Shelf Sag
Shelf Sag

A/B Test Significance Calculator (Conversion Rates, P-Value and Confidence Interval)

Enter the visitors and conversions for group A (the original version) and group B (the changed version). Along with each conversion rate, the difference and the relative lift, it runs a two-proportion z-test for the z-score and p-value, and gives a "significant / not significant" verdict at the significance level you choose.

Visitors can be clicks, sessions or users, but count them the same way in both groups. Enter whole numbers. For the significance level and one-tailed/two-tailed, choose what you decided before the test started (changing them after seeing the results is not a valid use).
Result and graph
Enter the visitors and conversions for groups A and B on the left and press "Calculate". The verdict and a graph will appear here.

What you can do on this page

  • Enter the visitors (clicks or sessions also work) and conversions for groups A and B, and get each conversion rate, the difference in conversion rates and the relative lift all at once
  • "B looks better, but could it just be random variation?" A two-proportion z-test answers this (z-score, two-tailed and one-tailed p-values, and a verdict at a 5%, 1% or 10% significance level)
  • It also gives the confidence interval for the difference (such as 95%) and the margin by which the z-score clears (or misses) the significance line
  • A standard normal distribution graph shows where your z-score falls and shades the rejection region (significant if the z-score lands there)
  • A plain-language explanation of the formulas and copy-and-paste formulas for Excel, Google Sheets and Python are all on this page
This page is only for judging whether the conversion rates (proportions) of groups A and B differ. If you just want to turn a z-score into a p-value, use the p-value calculator; for a confidence interval of an average, the confidence interval calculator; and for the number of people needed to study one proportion, the sample size calculator. "Statistically significant" does not guarantee that the effect is real or large (see "Key idea" below and the terms section).

What is this calculation used for?

Improving landing pages and ad copy (online marketing)

"Version B with a different button color raised the conversion rate from 2.0% to 2.5%." Checking whether this difference is not just chance is A/B test significance, done every day in online marketing. With 10,000 visitors each, the two-tailed p-value is about 0.017, "significant at 5%". But with similar conversion rates on 500 visitors each (10 conversions vs. 12 or 13), the p-value is about 0.5 to 0.7, and you can only say "it may be chance".
If you switch because "B looks better", it often happens that results drift back the next month. Deciding the number of visitors first and judging by a significance level prevents this (assuming visitors are split at random).

Testing email subject lines and send times (email marketing)

Send the same email with two different subject lines to half the list each and compare open rates of 48% and 52%. This is also a test of the difference between two proportions. With 5,000 emails each, the 4-point difference has a z-score of about 4, a difference that is hard to explain by chance.
Larger proportions such as open rates are easier to judge with the same number of emails than small ones such as conversion rates (the difference tends to be large compared with the standard error). This formula gives you a feel for how the metric you compare changes the number of sends you need.

Comparing response rates of treatments (medical research)

Comparing improvement rates between two groups, such as "60% improved with the standard treatment and 70% with the new one", is a basic test in clinical studies (real studies use strict designs such as randomization and blinding).
Papers behind drug approvals and treatment guidelines always report p-values and confidence intervals. Once you understand this formula, you are not swayed by the single word "significant"; you can also read "how large is the difference?" and "were there enough patients?".

Comparing defect rates on a production line (quality control)

When a new process B seems to lower the defect rate from 1.5% to 1.0%, the verdict changes with volume: with 2,000 units each it is "within chance", and with 20,000 units each it is "a significant improvement". This calculation keeps decisions on equipment and process changes from being swayed by chance in small samples.
For small proportions like defect rates, the normal approximation gets rough when the expected number of defects (units × defect rate) falls below 5, so an exact test that uses the binomial distribution directly (such as Fisher's exact test) is chosen instead.

Measuring the effect of teaching materials and methods (education)

The same test works for comparisons such as "classes using the new textbook had a 75% pass rate, versus 65% with the old one". But if you assign textbooks by class, students were not split at random, and differences in teachers and students get mixed into the result.
In education and social science research, along with checking significance with this formula, the design matters: were groups split at random, and did anyone change how the comparison was cut after seeing the results? Knowing the assumptions behind the formula is what keeps you from being fooled by numbers.

Formula

Conversion rate of each group
Standard notation (the usual math form)
\(\hat{p}_A\) \(=\) \(x_A\) \(\div\) \(n_A\)
In words (symbols replaced with words)
③ \(\hat{p}_A\): conversion rate of A \(=\) ① \(x_A\): conversions in A \(\div\) ② \(n_A\): visitors in A
The formula in words
① Take the \(x_A\): conversions in A
② divide them by the \(n_A\): visitors in A , and you get the
③ \(\hat{p}_A\): conversion rate of A (the same formula for B, \(\hat{p}_B = x_B \div n_B\))
Quick example
If A has 200 conversions from 10,000 visitors and B has 250 conversions from 10,000 visitors, the conversion rates are
conversion rate of A \(\hat{p}_A\) \(=\) conversions (200) \(\div\) visitors (10,000)
\(200 \div 10000 = 0.02\ \ (2\%)\)
\(250 \div 10000 = 0.025\ \ (2.5\%)\)
Key idea
The denominator of the conversion rate (visitors) can be clicks, sessions or users, but always count A and B the same way. In the formulas, proportions are decimals between 0 and 1, such as \(0.02\), and they are turned into percentages only for display. The symbol \(\hat{p}\) (p-hat) stands for "the true conversion rate \(p\), estimated from the data you have". Even if 10,000 people see the same page, the number of conversions goes up or down a little by chance. So "it was 2% this time" is not "the true conversion rate is exactly 2%"; it is only an estimate near it. The test below is how you deal with this random wobble properly.
Difference in conversion rates and relative lift
Standard notation (the usual math form)
\(d\) \(=\) \(\hat{p}_B\) \(-\) \(\hat{p}_A\)
\(L\) \(=\) \(d\) \(\div\) \(\hat{p}_A\) \(\times\) \(100\)
In words (symbols replaced with words)
③ \(d\): difference in conversion rates \(=\) ① \(\hat{p}_B\): conversion rate of B \(-\) ② \(\hat{p}_A\): conversion rate of A
⑦ \(L\): relative lift (%) \(=\) ④ \(d\): difference in conversion rates \(\div\) ⑤ \(\hat{p}_A\): conversion rate of A \(\times\) ⑥ \(100\): percent base
The formula in words
① From the \(\hat{p}_B\): conversion rate of B ,
② subtract the \(\hat{p}_A\): conversion rate of A , and you get the
③ \(d\): difference in conversion rates (shown times 100, in percentage points).
④ Divide the \(d\): difference in conversion rates
⑤ by the \(\hat{p}_A\): conversion rate of A ,
⑥ multiply by the \(100\): percent base to turn it into a percentage, and you get the
⑦ \(L\): relative lift (%)
Quick example
If A's conversion rate is 0.02 (2%) and B's is 0.025 (2.5%), the difference and the relative lift are
difference \(d\) \(=\) conversion rate of B (0.025) \(-\) conversion rate of A (0.02)
relative lift \(L\) (%) \(=\) difference (0.005) \(\div\) conversion rate of A (0.02) \(\times\) \(100\)
\(0.025 - 0.02 = 0.005\ \ (0.5\ \mathrm{pp})\)
\(0.005 \div 0.02 \times 100 = 25\ \ (25\%)\)
Key idea
"Up from 2% to 2.5%" is 0.5 percentage points as a difference, and a 25% increase as a ratio. To avoid confusion, the result of subtracting one percentage from another is called "percentage points" (pp), not "%" (2.5% − 2% = 0.5 percentage points). The relative lift shows "by what % B increases conversions compared with A", so it ties directly to estimates of ad spend and revenue and is widely used in practice. But these two numbers are the difference observed in this sample; the true difference may not be exactly the same. The test that follows checks that.
Pooled conversion rate (one rate for both groups combined)
Standard notation (the usual math form)
\(\hat{p}\) \(=\) \((x_A + x_B)\) \(\div\) \((n_A + n_B)\)
In words (symbols replaced with words)
③ \(\hat{p}\): pooled conversion rate \(=\) ① \((x_A + x_B)\): total conversions \(\div\) ② \((n_A + n_B)\): total visitors
The formula in words
① Take the \((x_A + x_B)\): total conversions
② divide them by the \((n_A + n_B)\): total visitors , and you get the
③ \(\hat{p}\): pooled conversion rate
Quick example
If A has 200 conversions from 10,000 visitors and B has 250 from 10,000, the pooled conversion rate is
pooled conversion rate \(\hat{p}\) \(=\) total conversions (200 + 250) \(\div\) total visitors (10,000 + 10,000)
\((200 + 250) \div (10000 + 10000) = 450 \div 20000 = 0.0225\ \ (2.25\%)\)
Key idea
The test starts from the hypothesis that "A and B really have the same conversion rate" (the null hypothesis). If they are the same, both groups' data come from one conversion rate, so you combine (pool) both groups to estimate it. This is the pooled conversion rate, used in the standard error calculation that comes next. "Assume they are the same, work it out, and a difference this large should almost never happen" - if that is what you find, you start to doubt the assumption (no difference). That is the logic of a hypothesis test.
Standard error and z-score (two-proportion z-test)
Standard notation (the usual math form)
\(\mathrm{SE}\) \(=\) \(\sqrt{\hat{p}(1-\hat{p})\left(\dfrac{1}{n_A}+\dfrac{1}{n_B}\right)}\)
\(z\) \(=\) \(d\) \(\div\) \(\mathrm{SE}\)
In words (symbols replaced with words)
② \(\mathrm{SE}\): standard error \(=\) ① square root of: spread of the pooled rate \(\hat{p}(1-\hat{p})\) × sum of the reciprocals of the visitors \(\left(\dfrac{1}{n_A}+\dfrac{1}{n_B}\right)\)
⑤ \(z\): z-score \(=\) ③ \(d\): difference in conversion rates \(\div\) ④ \(\mathrm{SE}\): standard error
The formula in words
① The square root of the spread of the pooled rate \(\hat{p}(1-\hat{p})\) times the sum of the reciprocals of the visitors \(\left(\dfrac{1}{n_A}+\dfrac{1}{n_B}\right)\) is the
② \(\mathrm{SE}\): standard error (a guide to how much the difference moves around by chance).
③ Divide the \(d\): difference in conversion rates
④ by the \(\mathrm{SE}\): standard error , and you get the
⑤ \(z\): z-score (how many times the chance wobble the difference is)
Quick example
With a pooled conversion rate of 0.0225, 10,000 visitors in each group and a difference of 0.005, the standard error and z-score are
standard error \(\mathrm{SE}\) \(=\) square root of "0.0225 × (1 − 0.0225) × (1/10000 + 1/10000)"
z-score \(z\) \(=\) difference (0.005) \(\div\) standard error (0.002097)
\(0.0225 \times (1 - 0.0225) \times \left(\dfrac{1}{10000} + \dfrac{1}{10000}\right) = 0.02199375 \times 0.0002 \approx 0.0000043988\)
\(\mathrm{SE} = \sqrt{0.0000043988} \approx 0.002097\)
\(z = 0.005 \div 0.002097 \approx 2.384\)
Key idea
The z-score tells you how many times the chance wobble (the standard error) the observed difference is. When there is no difference, the z-score follows the standard normal distribution with mean 0 and standard deviation 1, so a difference with \(|z|\) above 2 rarely (about 5% of the time or less) shows up by chance when there is no real difference. The reciprocal of visitors \(1/n\) is in the standard error formula, so the more visitors, the smaller the standard error, and the same 0.5-point difference gives a larger z-score (it becomes significant more easily). With few visitors, even a big difference can only be called "maybe chance". This formula approximates the spread of conversion counts (a binomial distribution) with a normal distribution. As a guide, the expected conversions \(n\hat{p}\) and expected non-conversions \(n(1-\hat{p})\) in each group should all be 5 or more (10 or more if possible); if not, the result shows a warning.
P-values (two-tailed and one-tailed)
Standard notation (the usual math form)
\(p_2\) \(=\) \(2\) \(\times\) \((\) \(1\) \(-\) \(\Phi(|z|)\) \()\)
\(p_1\) \(=\) \(1\) \(-\) \(\Phi(z)\)
In words (symbols replaced with words)
④ \(p_2\): two-tailed p-value \(=\) ③ \(2\): for both tails \(\times\) \((\) ① \(1\): total probability \(-\) ② \(\Phi(|z|)\): cumulative probability up to \(|z|\) \()\)
⑦ \(p_1\): one-tailed p-value \(=\) ⑤ \(1\): total probability \(-\) ⑥ \(\Phi(z)\): cumulative probability up to \(z\)
The formula in words
① From the \(1\): total probability ,
② subtract the \(\Phi(|z|)\): cumulative probability up to \(|z|\) (this leaves the area of one tail),
③ multiply by \(2\): for both tails , and you get the
④ \(p_2\): two-tailed p-value .
⑤ From the \(1\): total probability ,
⑥ subtract the \(\Phi(z)\): cumulative probability up to \(z\) , and you get the
⑦ \(p_1\): one-tailed p-value (for the hypothesis "B is higher than A")
Quick example
With a z-score of 2.384 (the cumulative probability of the standard normal distribution is \(\Phi(2.384) \approx 0.99144\)), the two-tailed and one-tailed p-values are
two-tailed p-value \(p_2\) \(=\) \(2\) \(\times\) \((\) \(1\) \(-\) \(\Phi(2.384)\) (0.99144) \()\)
\(p_2 = 2 \times (1 - 0.99144) = 2 \times 0.00856 \approx 0.0171\)
\(p_1 = 1 - 0.99144 \approx 0.0086\)
Key idea
The p-value is "the probability of a difference at least this large arising by chance when there is really no difference". A two-tailed p-value of 0.0171 says "if there were no difference, a difference this large (in either direction) would show up only about 1.7 times in 100". It is below the 5% significance level (0.05), so the verdict is "statistically significant"; it is above 1% (0.01), so it is "not significant at 1%". A two-tailed test asks "is there a difference between A and B (whichever is higher)?", and a one-tailed test asks only "is B higher than A?". The one-tailed p-value is half the two-tailed one, so it becomes significant more easily, but you may use it only if you decided **before the test started** that you would not consider B being lower. Switching to one-tailed after seeing the results is not a valid use. For the same reason, checking results again and again and stopping as soon as it turns significant (peeking), or testing many versions or metrics at once and adopting anything that turns significant (multiple comparisons), makes the chance of a "significant" result with no real difference far larger than 5%. The basic rule is to decide the visitors needed and the metrics to look at before the test, wait until then, and judge only once. \(\Phi\) (phi) is the cumulative distribution function of the standard normal distribution: the probability that the z-score is that value or less, the area to the left under the bell curve. You can get its values from a table or from Excel's NORM.S.DIST function (the calculator on this page computes them numerically).
Confidence interval for the difference in conversion rates
Standard notation (the usual math form)
\(d_L,\ d_U\) \(=\) \(d\) \(\pm\) \(z_{\alpha/2}\) \(\times\) \(\sqrt{\dfrac{\hat{p}_A(1-\hat{p}_A)}{n_A}+\dfrac{\hat{p}_B(1-\hat{p}_B)}{n_B}}\)
In words (symbols replaced with words)
④ \(d_L,\ d_U\): lower and upper limits of the confidence interval \(=\) ① \(d\): difference in conversion rates \(\pm\) ② \(z_{\alpha/2}\): set by the confidence level \(\times\) ③ \(\mathrm{SE}_d\): standard error from each group's spread
The formula in words
① Around the \(d\): difference in conversion rates , go up and down by
② the \(z_{\alpha/2}\): set by the confidence level (1.96 for 95%) times the
③ \(\mathrm{SE}_d\): standard error from each group's spread , and you get the
④ \(d_L,\ d_U\): lower and upper limits of the confidence interval
Quick example
If A's conversion rate is 0.02 (10,000 visitors), B's is 0.025 (10,000 visitors) and the difference is 0.005, the 95% confidence interval for the difference is
lower \(d_L\), upper \(d_U\) \(=\) difference (0.005) \(\pm\) \(z_{\alpha/2}\) (1.96) \(\times\) standard error (0.002097)
\(\mathrm{SE}_d = \sqrt{\dfrac{0.02 \times 0.98}{10000} + \dfrac{0.025 \times 0.975}{10000}} = \sqrt{0.00000196 + 0.0000024375} \approx 0.002097\)
\(0.005 \pm 1.96 \times 0.002097 = 0.005 \pm 0.00411\)
\(0.00089 \le d \le 0.00911\ \ (0.09\ \text{to}\ 0.91\ \mathrm{pp})\)
Key idea
A test only answers yes or no: "is there a difference?". A confidence interval also tells you the size: "the true difference is probably in about this range". In this example the 95% confidence interval is 0.09 to 0.91 percentage points. It does not include 0, so you can say "B is higher", but it also shows that the range is too wide to claim "it went up by exactly 0.5 points". For the standard error, use the spread of each group's own conversion rate, not the pooled value from the test (the interval does not assume "no difference"). \(z_{\alpha/2}\) is the standard normal value for the confidence level: 1.645 for 90%, 1.96 for 95% and 2.576 for 99%. If you set the confidence level to "100 − significance level" (95% for a 5% significance level), the two-tailed verdict and whether the interval includes 0 almost always agree (they do not correspond for a one-tailed test; also, the test uses the pooled standard error and the interval uses each group's, so they can rarely disagree near the borderline).
The z-score is the difference in conversion rates divided by the chance wobble (standard error) assuming no difference, and the p-value is the probability of a difference at least that large arising by chance. If the p-value is below the significance level (such as 5%), the verdict is "statistically significant", and the confidence interval for the difference shows the likely range of the true difference. Statistically significant says "hard to explain by chance"; it does not say "the effect is certain or large".

Symbols and terms

Symbols

\(n_A,\ n_B\) n sub A, n sub B Visitors (clicks or sessions) in groups A and B. \(n\) is the first letter of "number", the usual letter in statistics for sample size.
\(x_A,\ x_B\) x sub A, x sub B Conversions in groups A and B. Using \(x\) for the number of successes is the usual practice for the binomial distribution.
\(\hat{p}_A,\ \hat{p}_B\) p-hat sub A, p-hat sub B Conversion rates of groups A and B (proportions from the sample). \(p\) is the first letter of "proportion" or "probability", and the "\(\hat{}\)" (hat) on top marks a value estimated from the sample for the true value \(p\).
\(\hat{p}\) p-hat Pooled conversion rate. One estimated conversion rate from both groups combined, assuming no difference: \(\hat{p} = (x_A + x_B) \div (n_A + n_B)\).
\(d\) dee Difference in conversion rates (\(\hat{p}_B - \hat{p}_A\)), from the first letter of "difference". This page shows it times 100, in percentage points.
\(L\) el Relative lift (%), from the first letter of "lift". How many % B's conversion rate increased compared with A's.
\(\mathrm{SE}\) S-E Standard error. A guide to how much a value from the sample (here, the difference in conversion rates) moves around by chance; it is the standard deviation of that sample value. The test uses the pooled conversion rate, and the confidence interval uses \(\mathrm{SE}_d\) from each group's conversion rate.
\(z\) zee z-score (test statistic). The difference divided by the standard error: how many times the chance wobble the difference is. With no difference it follows the standard normal distribution (mean 0, standard deviation 1), so the p-value comes from how large it is.
\(|z|\) absolute value of z The size of the z-score without its sign. A two-tailed test does not care which direction the difference goes, so it uses the absolute value.
\(\Phi\) phi (capital) The cumulative distribution function of the standard normal distribution. \(\Phi(z)\) is the probability that a standard normal value is \(z\) or less, the area under the bell curve to the left of \(z\). The Greek capital letter phi is the usual symbol for it.
\(p_2,\ p_1\) p sub 2, p sub 1 The two-tailed and one-tailed p-values. The subscripts 2 and 1 stand for "two tails" and "one tail", a labeling used on this page only; statistics textbooks just write \(p\).
\(\alpha\) alpha The significance level: the line set in advance where "if the p-value is smaller, reject the null hypothesis". 5% (0.05) is common. Greek letter alpha; it is the probability of a Type I error (calling a difference real when there is none).
\(z_{\alpha/2}\) z sub alpha over 2 The boundary z-score (critical value) that leaves a total probability of \(\alpha\) in the two tails. 1.96 for a 95% confidence interval (\(\alpha = 0.05\)) and 2.576 for 99%. For one tail it is \(z_\alpha\) (1.645 for 5%).
\(d_L,\ d_U\) d sub L, d sub U The lower and upper limits of the confidence interval for the difference: the two ends of the range where the true difference is likely to be.

Terms

A/B test An experiment that shows visitors the original version (A) or a changed version (B) at random over the same period and compares the results. Widely used for websites, ads and email. Splitting at random makes factors other than the version (day of the week, traffic source and so on) the same in both groups.
conversion rate (CVR) The share of visits (clicks) that led to a result (conversion) such as a purchase or sign-up. To calculate or work back from a single conversion rate, see the conversion rate calculator page.
null hypothesis The hypothesis the test tries to reject: "A and B really have the same conversion rate" (\(p_A = p_B\)). The test checks how unlikely the observed difference would be if this were true.
alternative hypothesis The hypothesis you want to show, opposite to the null hypothesis. In a two-tailed test it is "A and B differ" (\(p_A \neq p_B\)); in a one-tailed test, "B is higher than A" (\(p_B > p_A\)).
significance level The line set before the test - if the p-value is smaller, reject the null hypothesis. 5% is common; use 1% for a stricter test. "Significant at 5%" says "the chance of a difference this large with no real difference is under 5%".
p-value The probability of a difference at least this large arising by chance, assuming the null hypothesis (no difference) is true. The smaller it is, the harder the result is to explain by chance. Note that it is not "the probability that B is better" (a common misunderstanding).
statistically significant The p-value is below the significance level, so the difference is judged hard to explain by random variation. It does not promise that the difference is large or the effect certain; with a huge number of visitors, even a tiny difference becomes significant.
test statistic A single number calculated from the data for the test. On this page it is the z-score, "difference ÷ standard error". How extreme it is decides the p-value.
two-proportion z-test A test that uses the normal distribution to check whether the proportions (such as conversion rates) of two groups differ. It is the most basic way to judge A/B test significance. You can do almost the same thing with a 2×2 chi-square test, and the two-tailed p-values match.
standard error The standard deviation of a statistic from a sample (here, the difference in conversion rates) - how much it moves around by chance depending on the sample. It shrinks in proportion to the square root of the visitors, so halving the standard error takes 4 times as many visitors.
pooling Combining both groups' data into one conversion rate estimate, on the assumption that there is no difference. The standard error for the test uses this pooled value.
confidence interval The range where the true difference is likely to be. A 95% confidence interval is a way of building intervals so that, if you repeated the experiment many times, 95% of them would contain the true difference. If the interval does not include 0, that almost always agrees with significance in a two-tailed test at 5%.
two-tailed test A test of "is there a difference (whichever is higher)?". Use it for an ordinary A/B test where you do not know which way it will go.
one-tailed test A test with a fixed direction, such as "is B higher than A?". Its p-value is half the two-tailed one, so it becomes significant more easily; use it only if you decided the direction before the test.
rejection region The tail area of the standard normal distribution where, if the z-score lands there, you reject the null hypothesis. For two-tailed 5% it is the two tails \(z < -1.96\) and \(z > 1.96\), with a total area of 0.05. It is the shaded part of the graph on this page.
critical value The z-score at the edge of the rejection region. 1.96 for two-tailed 5%, 2.576 for two-tailed 1% and 1.645 for one-tailed 5%. The "z-score margin" in the result shows how far this z-score is above (or below) the critical value.
normal approximation Conversion counts really follow a binomial distribution, but with many visitors a normal distribution (bell curve) approximates it well. The z-test uses this approximation, so as a guide the expected conversions \(n\hat{p}\) and expected non-conversions \(n(1-\hat{p})\) in each group should be 5 or more (10 or more if possible).
statistical power The probability that the test detects a real difference as significant. With few visitors, power is low, and even a real effect often comes out "not significant". "Not significant" does not prove "no difference".
Type I error Calling a result significant when there is really no difference (a false positive). The significance level \(\alpha\) is the chance of this error you accept; at 5%, "run 100 tests with no real difference and about 5 will be significant by chance".
multiple comparisons Testing many versions or metrics at once. Even at a 5% significance level, test 20 metrics and one will likely be "significant" by chance. You need to narrow the metrics in advance or use a stricter significance level (Bonferroni correction and others).
peeking (stopping early) Checking the results again and again during the test and stopping as soon as it looks significant. The p-value goes up and down during a test, so if you keep "stopping the first time it drops below 5%", the chance of significance with no real difference goes far above 5%. The basic rule is to decide the visitors needed first and wait until you reach them.
relative lift How many % B's conversion rate increased compared with A's: the difference \(d\) divided by A's conversion rate, times 100. "Up 25%" is easier to turn into estimates of extra conversions and revenue than "up 0.5 points".
percentage point The unit for the result of subtracting one percentage from another (abbreviated pp). Used as in "2.5% − 2% = 0.5 percentage points"; writing "0.5%" could be mistaken for "0.5% of 2%".

Good to know before you start

Here is what helps you use the calculation on this page with real understanding, not just by pressing the button.
If you get stuck, going back over these topics is the quickest way forward.

Percents (Grades 6–7)
  • Knowing that a percent is found as "part ÷ whole" (conversion rate = conversions ÷ visitors)
  • Being able to switch between percentages and decimals (\(2.5\% = 0.025\))
  • Knowing that a difference between percentages is stated in percentage points (2.5% − 2% = 0.5 percentage points)
Reciprocals and square roots (Grades 6–8)
  • Being able to add reciprocals, as in \(\dfrac{1}{5000} + \dfrac{1}{5000} = 0.0004\)
  • Understanding square roots and finding simple ones, as in \(\sqrt{0.0001} = 0.01\)
The normal distribution and standard deviation (high school statistics, AP Statistics)
  • Knowing that the standard normal distribution (mean 0, standard deviation 1) is bell-shaped, and that "the probability of \(z\) or less", \(\Phi(z)\), is the area to the left under the curve
  • Being able to read approximate values from a z-table, such as \(\Phi(1.96) \approx 0.975\)
  • A feel for "the chance of landing 2 or more standard deviations from the mean is about 5%"
The logic of hypothesis testing (high school statistics, AP Statistics)
  • Understanding the flow of setting a null hypothesis of "no difference" and judging by how unlikely the observed result would be under it
  • Understanding how significance level, p-value and rejection region relate (the p-value is below the significance level ⇔ the z-score falls in the rejection region)
The binomial distribution and its normal approximation (AP Statistics)
  • Knowing that "the number of successes in \(n\) tries of something with probability \(p\)" follows a binomial distribution, and that for large \(n\) a normal distribution approximates it

How to calculate it in Excel

Copy the whole table below and paste it into cell A1 in Excel. It works as is.
Table to find each conversion rate
Visitors in A n_A 10000
Conversions in A x_A 200
Visitors in B n_B 10000
Conversions in B x_B 250
Conversion rate of A p̂_A =B2/B1
Conversion rate of B p̂_B =B4/B3
Table to find the difference and relative lift
Conversion rate of A p̂_A 0.02
Conversion rate of B p̂_B 0.025
Difference d =B2-B1
Relative lift L (%) =B3/B1*100
Table to find the pooled conversion rate
Conversions in A x_A 200
Conversions in B x_B 250
Visitors in A n_A 10000
Visitors in B n_B 10000
Pooled conversion rate p̂ =(B1+B2)/(B3+B4)
Table to find the standard error and z-score
Pooled conversion rate p̂ 0.0225
Visitors in A n_A 10000
Visitors in B n_B 10000
Difference d 0.005
Standard error SE =SQRT(B1*(1-B1)*(1/B2+1/B3))
z-score z =B4/B5
Table to find the p-values (two-tailed and one-tailed)
z-score z 2.384
Two-tailed p-value p₂ =2*(1-NORM.S.DIST(ABS(B1),TRUE))
One-tailed p-value p₁ =1-NORM.S.DIST(B1,TRUE)
Table to find the confidence interval for the difference
Conversion rate of A p̂_A 0.02
Visitors in A n_A 10000
Conversion rate of B p̂_B 0.025
Visitors in B n_B 10000
Difference d 0.005
Confidence level (0.95 = 95%) 0.95
z_(α/2) =NORM.S.INV(1-(1-B6)/2)
Standard error SE_d =SQRT(B1*(1-B1)/B2+B3*(1-B3)/B4)
Lower limit d_L =B5-B7*B8
Upper limit d_U =B5+B7*B8
After pasting, the upper rows are your inputs and the rows with formulas are calculated automatically. Proportions are decimals between 0 and 1, such as 0.02; multiply by 100 for a percentage.
In the first table, B5 shows A's conversion rate 0.02 and B6 shows B's 0.025. The second table gives a difference of 0.005 (0.5 percentage points) and a lift of 25, the third a pooled conversion rate of 0.0225, and the fourth a standard error of about 0.002097 and a z-score of about 2.384.
In the fifth table, NORM.S.DIST(z, TRUE) returns the standard normal cumulative probability Φ(z); the two-tailed p-value is about 0.0171 and the one-tailed p-value about 0.0086. In the sixth table, NORM.S.INV finds the z-score from a cumulative probability (1.96 for a 95% confidence level), giving a lower limit of about 0.0009 and an upper limit of about 0.0091 (0.09 to 0.91 percentage points).

How to calculate it in Google Sheets

Copy the whole table below and paste it into cell A1 in Google Sheets. It works as is.
Table to find each conversion rate
Visitors in A n_A 10000
Conversions in A x_A 200
Visitors in B n_B 10000
Conversions in B x_B 250
Conversion rate of A p̂_A =B2/B1
Conversion rate of B p̂_B =B4/B3
Table to find the difference and relative lift
Conversion rate of A p̂_A 0.02
Conversion rate of B p̂_B 0.025
Difference d =B2-B1
Relative lift L (%) =B3/B1*100
Table to find the pooled conversion rate
Conversions in A x_A 200
Conversions in B x_B 250
Visitors in A n_A 10000
Visitors in B n_B 10000
Pooled conversion rate p̂ =(B1+B2)/(B3+B4)
Table to find the standard error and z-score
Pooled conversion rate p̂ 0.0225
Visitors in A n_A 10000
Visitors in B n_B 10000
Difference d 0.005
Standard error SE =SQRT(B1*(1-B1)*(1/B2+1/B3))
z-score z =B4/B5
Table to find the p-values (two-tailed and one-tailed)
z-score z 2.384
Two-tailed p-value p₂ =2*(1-NORMSDIST(ABS(B1)))
One-tailed p-value p₁ =1-NORMSDIST(B1)
Table to find the confidence interval for the difference
Conversion rate of A p̂_A 0.02
Visitors in A n_A 10000
Conversion rate of B p̂_B 0.025
Visitors in B n_B 10000
Difference d 0.005
Confidence level (0.95 = 95%) 0.95
z_(α/2) =NORMSINV(1-(1-B6)/2)
Standard error SE_d =SQRT(B1*(1-B1)/B2+B3*(1-B3)/B4)
Lower limit d_L =B5-B7*B8
Upper limit d_U =B5+B7*B8
Almost the same formulas as in Excel work. Write the standard normal cumulative probability as NORMSDIST(z) and its inverse as NORMSINV(p) (they do the same job as Excel's NORM.S.DIST and NORM.S.INV). Copy the whole table, paste it into cell A1, and replace the inputs with your own numbers.

How to calculate it in Python

import math

visits_a, conversions_a = 10000, 200   # visitors and conversions in A
visits_b, conversions_b = 10000, 250   # visitors and conversions in B
alpha = 0.05                           # significance level (5%)

def normal_cdf(z):
    # standard normal cumulative distribution function Φ(z)
    return 0.5 * math.erfc(-z / math.sqrt(2))

def normal_ppf(p):
    # inverse of Φ (bisection; used to find z_(α/2))
    low, high = -10.0, 10.0
    for _ in range(100):
        mid = (low + high) / 2
        if normal_cdf(mid) < p:
            low = mid
        else:
            high = mid
    return (low + high) / 2

cvr_a = conversions_a / visits_a
cvr_b = conversions_b / visits_b
diff = cvr_b - cvr_a
lift = diff / cvr_a * 100
pooled = (conversions_a + conversions_b) / (visits_a + visits_b)
se_pooled = math.sqrt(pooled * (1 - pooled) * (1 / visits_a + 1 / visits_b))
z = diff / se_pooled
p_two_sided = 2 * (1 - normal_cdf(abs(z)))
p_one_sided = 1 - normal_cdf(z)

se_diff = math.sqrt(cvr_a * (1 - cvr_a) / visits_a + cvr_b * (1 - cvr_b) / visits_b)
z_half = normal_ppf(1 - alpha / 2)
ci_low, ci_high = diff - z_half * se_diff, diff + z_half * se_diff

print(f"Conversion rate A: {cvr_a * 100:.2f}%  Conversion rate B: {cvr_b * 100:.2f}%")
print(f"Difference: {diff * 100:.2f} percentage points  Relative lift: {lift:.1f}%")
print(f"z-score: {z:.3f}  two-tailed p-value: {p_two_sided:.4f}  one-tailed p-value: {p_one_sided:.4f}")
print("Verdict (two-tailed):", "statistically significant" if p_two_sided < alpha else "not statistically significant")
print(f"{(1 - alpha) * 100:.0f}% confidence interval for the difference: {ci_low * 100:.2f} to {ci_high * 100:.2f} percentage points")
Runs with the standard library only (math.erfc gives the standard normal cumulative probability). In this example it prints a z-score of 2.384, a two-tailed p-value of 0.0171, a one-tailed p-value of 0.0086, the verdict "statistically significant", and a 95% confidence interval for the difference of 0.09 to 0.91 percentage points. Change the four numbers and the significance level at the top and run it. If scipy is available, you can use scipy.stats.norm.cdf and norm.ppf instead.

How to write it in LaTeX and other math languages (copy and paste)

Conversion rate of each group
p̂_A = x_A ÷ n_A
\hat{p}_A = \dfrac{x_A}{n_A}
<math xmlns="http://www.w3.org/1998/Math/MathML" display="block">
  <mrow>
    <msub><mover accent="true"><mi>p</mi><mo>^</mo></mover><mi>A</mi></msub>
    <mo>=</mo>
    <mfrac>
      <msub><mi>x</mi><mi>A</mi></msub>
      <msub><mi>n</mi><mi>A</mi></msub>
    </mfrac>
  </mrow>
</math>
hat(p)_A = x_A / n_A
xA/nA
pA := xA/nA;
pA = xA/nA;
p̂_A = x_A/n_A
Difference in conversion rates and relative lift
d = p̂_B − p̂_A,  L = d ÷ p̂_A × 100
d = \hat{p}_B - \hat{p}_A, \quad L = \dfrac{d}{\hat{p}_A} \times 100
<math xmlns="http://www.w3.org/1998/Math/MathML" display="block">
  <mrow>
    <mi>d</mi><mo>=</mo>
    <msub><mover accent="true"><mi>p</mi><mo>^</mo></mover><mi>B</mi></msub>
    <mo>&#x2212;</mo>
    <msub><mover accent="true"><mi>p</mi><mo>^</mo></mover><mi>A</mi></msub>
    <mo>,</mo><mspace width="1em"/>
    <mi>L</mi><mo>=</mo>
    <mfrac><mi>d</mi><msub><mover accent="true"><mi>p</mi><mo>^</mo></mover><mi>A</mi></msub></mfrac>
    <mo>&#xD7;</mo><mn>100</mn>
  </mrow>
</math>
d = hat(p)_B - hat(p)_A, L = d / hat(p)_A xx 100
d = pB - pA; d/pA*100
d := pB - pA; L := d/pA*100;
d = pB - pA; L = d/pA*100;
d = p̂_B − p̂_A, L = d/p̂_A × 100
Pooled conversion rate (one rate for both groups combined)
p̂ = (x_A + x_B) ÷ (n_A + n_B)
\hat{p} = \dfrac{x_A + x_B}{n_A + n_B}
<math xmlns="http://www.w3.org/1998/Math/MathML" display="block">
  <mrow>
    <mover accent="true"><mi>p</mi><mo>^</mo></mover>
    <mo>=</mo>
    <mfrac>
      <mrow><msub><mi>x</mi><mi>A</mi></msub><mo>+</mo><msub><mi>x</mi><mi>B</mi></msub></mrow>
      <mrow><msub><mi>n</mi><mi>A</mi></msub><mo>+</mo><msub><mi>n</mi><mi>B</mi></msub></mrow>
    </mfrac>
  </mrow>
</math>
hat(p) = (x_A + x_B) / (n_A + n_B)
(xA + xB)/(nA + nB)
pPool := (xA + xB)/(nA + nB);
pPool = (xA + xB)/(nA + nB);
p̂ = (x_A + x_B)/(n_A + n_B)
Standard error and z-score (two-proportion z-test)
SE = √(p̂(1 − p̂)(1/n_A + 1/n_B)),  z = d ÷ SE
\mathrm{SE} = \sqrt{\hat{p}(1-\hat{p})\left(\dfrac{1}{n_A}+\dfrac{1}{n_B}\right)}, \quad z = \dfrac{d}{\mathrm{SE}}
<math xmlns="http://www.w3.org/1998/Math/MathML" display="block">
  <mrow>
    <mi>z</mi><mo>=</mo>
    <mfrac>
      <mrow><msub><mover accent="true"><mi>p</mi><mo>^</mo></mover><mi>B</mi></msub><mo>&#x2212;</mo><msub><mover accent="true"><mi>p</mi><mo>^</mo></mover><mi>A</mi></msub></mrow>
      <msqrt>
        <mover accent="true"><mi>p</mi><mo>^</mo></mover>
        <mo>(</mo><mn>1</mn><mo>&#x2212;</mo><mover accent="true"><mi>p</mi><mo>^</mo></mover><mo>)</mo>
        <mo>(</mo>
        <mfrac><mn>1</mn><msub><mi>n</mi><mi>A</mi></msub></mfrac>
        <mo>+</mo>
        <mfrac><mn>1</mn><msub><mi>n</mi><mi>B</mi></msub></mfrac>
        <mo>)</mo>
      </msqrt>
    </mfrac>
  </mrow>
</math>
SE = sqrt(hat(p)(1 - hat(p))(1/n_A + 1/n_B)), z = d / SE
se = Sqrt[pPool*(1 - pPool)*(1/nA + 1/nB)]; z = d/se
SE := sqrt(pPool*(1 - pPool)*(1/nA + 1/nB)); z := d/SE;
SE = sqrt(pPool*(1 - pPool)*(1/nA + 1/nB)); z = d/SE;
SE = √(p̂(1 − p̂)(1/n_A + 1/n_B)), z = d/SE
P-values (two-tailed and one-tailed)
p₂ = 2(1 − Φ(|z|)),  p₁ = 1 − Φ(z)
p_2 = 2\left(1 - \Phi(|z|)\right), \quad p_1 = 1 - \Phi(z)
<math xmlns="http://www.w3.org/1998/Math/MathML" display="block">
  <mrow>
    <msub><mi>p</mi><mn>2</mn></msub><mo>=</mo>
    <mn>2</mn><mo>(</mo><mn>1</mn><mo>&#x2212;</mo>
    <mi>&#x3A6;</mi><mo>(</mo><mo>|</mo><mi>z</mi><mo>|</mo><mo>)</mo><mo>)</mo>
    <mo>,</mo><mspace width="1em"/>
    <msub><mi>p</mi><mn>1</mn></msub><mo>=</mo>
    <mn>1</mn><mo>&#x2212;</mo><mi>&#x3A6;</mi><mo>(</mo><mi>z</mi><mo>)</mo>
  </mrow>
</math>
p_2 = 2(1 - Phi(|z|)), p_1 = 1 - Phi(z)
p2 = 2*(1 - CDF[NormalDistribution[0, 1], Abs[z]]); p1 = 1 - CDF[NormalDistribution[0, 1], z]
p2 := 2*(1 - Statistics:-CDF(Normal(0, 1), abs(z))); p1 := 1 - Statistics:-CDF(Normal(0, 1), z);
p2 = 2*(1 - normcdf(abs(z))); p1 = 1 - normcdf(z);
p_2 = 2(1 − Φ(|z|)), p_1 = 1 − Φ(z)
Confidence interval for the difference in conversion rates
d ± z_(α/2) × √(p̂_A(1 − p̂_A)/n_A + p̂_B(1 − p̂_B)/n_B)
d \pm z_{\alpha/2}\sqrt{\dfrac{\hat{p}_A(1-\hat{p}_A)}{n_A}+\dfrac{\hat{p}_B(1-\hat{p}_B)}{n_B}}
<math xmlns="http://www.w3.org/1998/Math/MathML" display="block">
  <mrow>
    <mi>d</mi><mo>&#xB1;</mo>
    <msub><mi>z</mi><mrow><mi>&#x3B1;</mi><mo>/</mo><mn>2</mn></mrow></msub>
    <msqrt>
      <mfrac>
        <mrow><msub><mover accent="true"><mi>p</mi><mo>^</mo></mover><mi>A</mi></msub><mo>(</mo><mn>1</mn><mo>&#x2212;</mo><msub><mover accent="true"><mi>p</mi><mo>^</mo></mover><mi>A</mi></msub><mo>)</mo></mrow>
        <msub><mi>n</mi><mi>A</mi></msub>
      </mfrac>
      <mo>+</mo>
      <mfrac>
        <mrow><msub><mover accent="true"><mi>p</mi><mo>^</mo></mover><mi>B</mi></msub><mo>(</mo><mn>1</mn><mo>&#x2212;</mo><msub><mover accent="true"><mi>p</mi><mo>^</mo></mover><mi>B</mi></msub><mo>)</mo></mrow>
        <msub><mi>n</mi><mi>B</mi></msub>
      </mfrac>
    </msqrt>
  </mrow>
</math>
d +- z_(alpha/2) sqrt((hat(p)_A(1 - hat(p)_A))/n_A + (hat(p)_B(1 - hat(p)_B))/n_B)
seD = Sqrt[pA*(1 - pA)/nA + pB*(1 - pB)/nB]; {d - zHalf*seD, d + zHalf*seD}
seD := sqrt(pA*(1 - pA)/nA + pB*(1 - pB)/nB); [d - zHalf*seD, d + zHalf*seD];
seD = sqrt(pA*(1 - pA)/nA + pB*(1 - pB)/nB); ci = [d - zHalf*seD, d + zHalf*seD];
d ± z_(α/2) √(p̂_A(1 − p̂_A)/n_A + p̂_B(1 − p̂_B)/n_B)

How to have ChatGPT  do the calculation

You are a calculation assistant for statistics (hypothesis testing). Do the following calculation by actually running Python code, and base your answer only on the numbers from the execution result (do not answer by mental math or guessing).

The A/B test results are:
- Group A: 10000 visitors, 200 conversions
- Group B: 10000 visitors, 250 conversions
Find each of the following:
1. The conversion rate of each group (%), the difference in conversion rates (percentage points) and the relative lift (%)
2. The pooled conversion rate and the pooled standard error SE = √(p̂(1−p̂)(1/n_A + 1/n_B))
3. The z-score = difference ÷ SE, and the two-tailed and one-tailed p-values (standard normal distribution)
4. Whether the difference is statistically significant at a 5% significance level (two-tailed)
5. The 95% confidence interval for the difference (using the standard error from each group's conversion rate, √(p̂_A(1−p̂_A)/n_A + p̂_B(1−p̂_B)/n_B))

In Python, use math.erfc (or scipy.stats.norm) and show the formulas you used and the numbers from the execution result.

How to Use
  1. 1
    Enter your numbers
    Type the numbers you want to calculate with into the input fields
  2. 2
    Calculate
    Press the "Calculate" button
  3. 3
    Check the result
    The result appears on the spot. The same page also explains the idea behind the calculation and the formula
  DataChef Features
Easy and Free
Unlimited conversions for free.
No technical knowledge required.
Intuitive and user-friendly operation.
No Registration Required
Available immediately after access.
Can be used without registering personal information.
Safe and Secure
Fully SSL encrypted communication.
Automatic file deletion by clicking "download".
Fast
High-speed site access
and rapid file conversion.
No Watermark
No watermark.
No attribution required.
Commercial Use Available
Free for commercial use.
No need to contact us for commercial use permission.