Bookmarks    
nPr and nCr    
Random Number    
SD Calculator    
Sample Size    
Percent Error    
Density    
Molarity    
Molar Mass    
Ohm's Law    
Watts to Amps    
Voltage Drop    
Long Division    
Mixed Numbers    
Rounding    
Nth Root    
Exponents    
Half-Life    
Polar Form    
De Moivre    
3D Distance    
Point to Line    
Cross Product    
Determinant    
Sin Cos Tan    
Triangle Area    
Scale Factor    
Sector Area    
Ellipse Area    
Cube Volume    
Box Volume    
Sphere Volume    
Cone Volume    
Pipe Volume    
Time Duration    
Time Card    
Present Value    
Future Value    
Churn Rate    
A/B Test Calc    
SEO Traffic    
Ideal Weight    
Fat Intake    
Child Height    
Golf Handicap    
Heat Index    
Wind Chill    
Dew Point    
Download Time    
kWh to Cost    
AC Size (BTU)    
Heating Costs    
LED Savings    
Trip Gas Cost    
Tire Size    
Solar Output    
Solar Payback    
Battery Size    
Wall Area    
Gravel Needed    
Mortar Mix    
Slope Grade    
Curtain Size    
Soil Needed    
Sod Needed    
Ramp Length    
Blind Size    
Drain Slope    
Board Feet    
Heat Loss    
Furniture Fit    
Moving Boxes    
Plywood Cuts    
Shelf Sag    
   Add
Probability and random number calculators
Independent Events
Independent Events
Two Events Solver
Two Events Solver
Repeated Trials
Repeated Trials
Bayes' Theorem
Bayes' Theorem
Expected Value
Expected Value
Binomial Distribution
Binomial Distribution
nPr and nCr
nPr and nCr
Circular Permutation
Circular Permutation
With Repetition
With Repetition
Random Number
Random Number
Averages and statistics calculators
Average Calculator
Average Calculator
Mean Median Mode
Mean Median Mode
SD Calculator
SD Calculator
Quartiles & IQR
Quartiles & IQR
Frequency Table
Frequency Table
Correlation (r)
Correlation (r)
Normal Probability
Normal Probability
Z-Score Calculator
Z-Score Calculator
Confidence Interval
Confidence Interval
Sample Size
Sample Size
Mark & Recapture
Mark & Recapture
P-Value Calculator
P-Value Calculator
Percentage and ratio calculators
Percentage Calc
Percentage Calc
Percent Change
Percent Change
Percent Difference
Percent Difference
Percent Error
Percent Error
Ratio Calculator
Ratio Calculator
Discount Calculator
Discount Calculator
Sales Tax Calculator
Sales Tax Calculator
Margin Calculator
Margin Calculator
Speed calculators
Speed Calculator
Speed Calculator
Density and concentration calculators
Density
Density
Molarity
Molarity
Molar Mass
Molar Mass
Physics and electricity calculators
Ohm's Law
Ohm's Law
Watts to Amps
Watts to Amps
Resistor Colors
Resistor Colors
Voltage Drop
Voltage Drop
Unit conversion calculators
Weight Converter
Weight Converter
Shoe Size Converter
Shoe Size Converter
Integer and signed number calculators
Long Division
Long Division
LCM Calculator
LCM Calculator
GCF Calculator
GCF Calculator
Integer Calculator
Integer Calculator
Prime Factorization
Prime Factorization
Diophantine Solver
Diophantine Solver
Modulo Calculator
Modulo Calculator
Factor Calculator
Factor Calculator
Roman Numerals
Roman Numerals
Fraction, decimal and rounding calculators
Fraction Calculator
Fraction Calculator
Mixed Numbers
Mixed Numbers
Simplify Fractions
Simplify Fractions
Fraction to Decimal
Fraction to Decimal
Decimal to Fraction
Decimal to Fraction
Rounding
Rounding
Equation and inequality calculators
Linear Equation
Linear Equation
Linear Systems
Linear Systems
Quadratic Formula
Quadratic Formula
Absolute Value
Absolute Value
Quadratic Inequality
Quadratic Inequality
Polynomial calculators
Binomial Theorem
Binomial Theorem
Square root and nth root calculators
Simplify Radicals
Simplify Radicals
Nth Root
Nth Root
Exponent and logarithm calculators
Exponents
Exponents
Log Calculator
Log Calculator
Number of Digits
Number of Digits
Scientific Notation
Scientific Notation
Sci. Notation Math
Sci. Notation Math
Half-Life
Half-Life
Complex number calculators
Complex Numbers
Complex Numbers
Polar Form
Polar Form
De Moivre
De Moivre
Function and graph calculators
Slope Calculator
Slope Calculator
Linear Function
Linear Function
Direct & Inverse Variation
Direct & Inverse Variation
y = ax² Calculator
y = ax² Calculator
Distance Formula
Distance Formula
3D Distance
3D Distance
Section Formula
Section Formula
Point to Line
Point to Line
Lat/Long Distance
Lat/Long Distance
Complete the Square
Complete the Square
Circle Equation
Circle Equation
Conic Sections
Conic Sections
Polar Coordinates
Polar Coordinates
Sequence calculators
Arithmetic Sequence
Arithmetic Sequence
Geometric Sequence
Geometric Sequence
Fibonacci Sequence
Fibonacci Sequence
Recurrence Relation
Recurrence Relation
Vector calculators
Vector Calculator
Vector Calculator
Cross Product
Cross Product
Matrix calculators
Matrix Calculator
Matrix Calculator
Determinant
Determinant
Inverse Matrix
Inverse Matrix
Plane geometry calculators
Sin Cos Tan
Sin Cos Tan
Degrees ⇔ Radians
Degrees ⇔ Radians
a sin θ + b cos θ
a sin θ + b cos θ
Triangle Solver
Triangle Solver
Triangle Area
Triangle Area
Right Triangle
Right Triangle
Pythagorean Theorem
Pythagorean Theorem
Polygon Angles
Polygon Angles
Scale Factor
Scale Factor
Parallel Lines
Parallel Lines
Rectangle Area
Rectangle Area
Parallelogram Area
Parallelogram Area
Trapezoid Area
Trapezoid Area
Circle Calculator
Circle Calculator
Sector Area
Sector Area
Inscribed Angle
Inscribed Angle
Ellipse Area
Ellipse Area
Solid geometry calculators
Cube Volume
Cube Volume
Cube Surface Area
Cube Surface Area
Box Volume
Box Volume
Box Surface Area
Box Surface Area
Cylinder Volume
Cylinder Volume
Cylinder Surface
Cylinder Surface
Sphere Volume
Sphere Volume
Sphere Surface
Sphere Surface
Spherical Cap Volume
Spherical Cap Volume
Cap Surface Area
Cap Surface Area
Ellipsoid Volume
Ellipsoid Volume
Ellipsoid Surface
Ellipsoid Surface
Pyramid Volume
Pyramid Volume
Pyramid Surface
Pyramid Surface
Cone Volume
Cone Volume
Cone Surface Area
Cone Surface Area
Frustum Volume
Frustum Volume
Frustum Surface Area
Frustum Surface Area
Pipe Volume
Pipe Volume
Capsule Volume
Capsule Volume
Capsule Surface Area
Capsule Surface Area
Date and time calculators
Age Calculator
Age Calculator
Days Between Dates
Days Between Dates
Date Calculator
Date Calculator
Hours From Now
Hours From Now
Day of the Week
Day of the Week
Time Calculator
Time Calculator
Time Zone Converter
Time Zone Converter
Hours Calculator
Hours Calculator
Time Duration
Time Duration
Time Card
Time Card
Finance and economics calculators
Compound Interest
Compound Interest
Simple Interest
Simple Interest
Interest Calculator
Interest Calculator
TVM Calculator
TVM Calculator
Present Value
Present Value
Future Value
Future Value
ROI Calculator
ROI Calculator
IRR Calculator
IRR Calculator
Payback Period
Payback Period
Average Return
Average Return
GDP Calculator
GDP Calculator
Web marketing and ad metric calculators
CTR Calculator
CTR Calculator
Conversion Rate
Conversion Rate
CPC, CPM & CPA
CPC, CPM & CPA
ROAS Calculator
ROAS Calculator
Break-Even CPA
Break-Even CPA
LTV Calculator
LTV Calculator
CAC Calculator
CAC Calculator
Churn Rate
Churn Rate
A/B Test Calc
A/B Test Calc
A/B Sample Size
A/B Sample Size
SEO Traffic
SEO Traffic
Break-Even Point
Break-Even Point
Markup vs. Margin
Markup vs. Margin
CAGR Calculator
CAGR Calculator
Health and fitness calculators
BMI Calculator
BMI Calculator
Sleep Calculator
Sleep Calculator
Calorie Calculator
Calorie Calculator
BMR Calculator
BMR Calculator
TDEE Calculator
TDEE Calculator
Ideal Weight
Ideal Weight
Body Fat Calculator
Body Fat Calculator
Lean Body Mass
Lean Body Mass
Calories Burned
Calories Burned
Protein Intake
Protein Intake
Macro Calculator
Macro Calculator
Carb Calculator
Carb Calculator
Fat Intake
Fat Intake
Child Height
Child Height
Sports calculators
Golf Handicap
Golf Handicap
Pace Calculator
Pace Calculator
1RM Calculator
1RM Calculator
Target Heart Rate
Target Heart Rate
Weather calculators
Heat Index
Heat Index
Wind Chill
Wind Chill
Dew Point
Dew Point
Computer calculators
Base Converter
Base Converter
Subnet Calculator
Subnet Calculator
Download Time
Download Time
Household energy and budget calculators
Electricity Cost
Electricity Cost
kWh to Cost
kWh to Cost
Yearly kWh to Cost
Yearly kWh to Cost
AC Size (BTU)
AC Size (BTU)
AC Running Cost
AC Running Cost
Heating Costs
Heating Costs
Gas vs Electric
Gas vs Electric
LED Savings
LED Savings
Salary Calculator
Salary Calculator
Budget Calculator
Budget Calculator
Car calculators
Trip Gas Cost
Trip Gas Cost
EV Charging Cost
EV Charging Cost
EV vs Gas Cost
EV vs Gas Cost
MPG Calculator
MPG Calculator
Tire Size
Tire Size
Solar power and battery calculators
Solar Output
Solar Output
Solar Panel Count
Solar Panel Count
Solar Payback
Solar Payback
Battery Size
Battery Size
Home and DIY calculators
Tile Calculator
Tile Calculator
Stair Calculator
Stair Calculator
Concrete Volume
Concrete Volume
Wall Area
Wall Area
Wallpaper Rolls
Wallpaper Rolls
Paint Calculator
Paint Calculator
Flooring Needed
Flooring Needed
Exterior Walls
Exterior Walls
Gravel Needed
Gravel Needed
Mortar Mix
Mortar Mix
Slope Grade
Slope Grade
Lumber Cut List
Lumber Cut List
Lot Coverage/FAR
Lot Coverage/FAR
Sheet Vinyl Roll
Sheet Vinyl Roll
Insulation Needed
Insulation Needed
Curtain Size
Curtain Size
TV Size & Distance
TV Size & Distance
Soil Needed
Soil Needed
Sod Needed
Sod Needed
Block Calculator
Block Calculator
Brick Calculator
Brick Calculator
Deck Materials
Deck Materials
Ramp Length
Ramp Length
Pilot Hole Size
Pilot Hole Size
Room Ventilation
Room Ventilation
Paint Thinning
Paint Thinning
Baseboard & Trim
Baseboard & Trim
Blind Size
Blind Size
Picture Hanging
Picture Hanging
Drain Slope
Drain Slope
Screw Calculator
Screw Calculator
Board Feet
Board Feet
Fence Calculator
Fence Calculator
Wood Shrinkage
Wood Shrinkage
Caulk Calculator
Caulk Calculator
Heat Loss
Heat Loss
Furniture Fit
Furniture Fit
Moving Boxes
Moving Boxes
Storage Capacity
Storage Capacity
Plywood Cuts
Plywood Cuts
Shelf Sag
Shelf Sag

A/B Test Sample Size Calculator (MDE, Significance, Power and Test Duration)

To find the sample size, enter the baseline CVR, the MDE, the significance level and the power, then press "Calculate". You get the sample size per group for A and B (rounded up) and the total, plus the days needed if you enter your daily visitors. Switch "Solve for" to power to fix the sample size per group and find the power 1 − β instead.

Enter the CVR and the MDE as percentages (for 3%, enter "3"; for a 10% lift, enter "10"). The MDE can be a relative % (3% to 3.3% is 10%) or in percentage points, pp (3% to 3.3% is 0.3). For daily visitors, enter the total for A and B together (all visits to the page being tested).
Result and graph
Fill in the fields on the left and press "Calculate". The result and a graph of MDE vs. required sample size will appear here.

What you can do on this page

  • Enter the baseline CVR (the conversion rate of your current version A), the minimum detectable effect (MDE, as a relative % or in percentage points), the significance level \(\alpha\) (two-sided or one-sided) and the power \(1-\beta\). You get the sample size each group of your A/B test needs (rounded up) and the total
  • "My page converts at 3%. How many visitors do A and B each need to detect a 10% lift (3% to 3.3%)?" This page answers the most important question to ask before a test starts
  • Enter your daily visitors (optional) to also see how many days it takes to reach that sample size (rounded up)
  • If you can only get a fixed number of visitors per group, fix the sample size and solve for the power \(1-\beta\) (the chance of finding a difference when there really is one)
  • A graph of how the sample size changes with the MDE, a plain-language explanation of the formulas, and copy-and-paste formulas for Excel, Google Sheets and Python are all on this page
This page is only for A/B tests that check whether versions A and B differ in CVR (conversion rate). It uses the standard formula for a two-proportion z-test. To find how many people a survey needs to measure one percentage precisely, use the "Sample Size Calculator". To read the p-value from results you have already collected, use the "P-Value Calculator". Reaching the required sample size does not by itself prove that a significant result is a real effect. Please also read the cautions in the formula explanation.

What is this calculation used for?

Planning an A/B test of the "Buy" button in an online store

Suppose an online store's product page converts at 2%, and you want to detect a relative lift of 10% (2% to 2.2%) from new button text. With a 5% significance level (two-sided) and 80% power, you need about 80,682 visitors per group, about 161,364 in total.
Comparing a few dozen or a few hundred visitors ("B did better last week") cannot tell a chance difference from a real one. Working out the sample size before you start keeps you from spending weeks on a test that cannot give an answer.

Deciding how many days to test a B2B SaaS landing page

A landing page gets demo requests at a 5% CVR. To detect a 1-point lift (5% to 6%), you need 8,158 visitors per group, 16,316 in total. At 300 visits a day, that is \(16316 \div 300 \approx 54.4\), so 55 days, or 8 weeks if you want to even out the days of the week.
If someone asks for results in 2 weeks, this calculation lets you explain with numbers: "In 2 weeks (4,200 visitors, 2,100 per group), we would need a much larger MDE, around 3 points (5% to 8%, which needs 1,059 per group)."

Testing email newsletter subject lines (open rates)

If your newsletter has a 20% open rate and you want to detect a 2-point lift (20% to 22%) from a new subject line, you need to send to 6,510 people per group, 13,020 in total (5% significance level, 80% power).
A metric with a high rate, like an open rate, needs fewer people to show a difference than a CVR of a few percent. This is because the \(p(1-p)\) part of the formula is different: the same "2 points" needs a different sample size for each metric.

Checking the power before concluding "no difference"

Suppose a test aimed at a relative 10% lift on a page with a 3% CVR ran with 5,000 visitors per group and showed "no significant difference". Working back the power gives only about 13.5%. The test was set up to miss a real improvement more than 8 times out of 10.
The right conclusion here is not "it had no effect" but "this sample size could not tell". Working back the power keeps you from misreading a test and throwing away a good idea.

Comparing retention after a new app tutorial

An app has a 40% next-day retention rate. To detect a relative 10% lift (40% to 44%) from a new tutorial, about 2,389 new users per group, about 4,778 in total, are enough (5% significance level, 80% power).
A metric close to 50%, like retention, has a large spread, but the difference in rates (4 points) is also large, so results come faster than in a test on a CVR of a few percent. Estimating the sample size for each metric also helps you decide which improvement to work on first.

Using 90% power for tests you cannot afford to miss

For a test where missing a good idea would be costly, such as a change in how prices are shown, you can raise the power from 80% to 90%. For a 3% CVR and a relative 10% lift, the sample size per group rises from 53,211 to 71,233 (about 1.34 times).
This calculation shows exactly how many more visitors it takes to cut misses in half (from 20% to 10%), so you can match the sample size and the test length to how important the test is.

Formulas and figures

Formula for the target CVR (the CVR you hope version B reaches)
Standard notation (the usual math form)
\(p_2\) \(=\) \(p_1\) \(\times\) \(\left(1 + \dfrac{r}{100}\right)\)
\(p_2\) \(=\) \(p_1\) \(+\) \(d\)
In words (symbols replaced with words)
③ \(p_2\): target CVR \(=\) ① \(p_1\): baseline CVR \(\times\) ② \(1 +\) relative lift \(r \div 100\)
\(p_2\): target CVR \(=\) \(p_1\): baseline CVR \(+\) ④ \(d\): MDE in percentage points
The formula in words
① Take the \(p_1\): baseline CVR
② multiply it by \(1 +\) relative lift \(r \div 100\)
③ and you get the \(p_2\): target CVR
④ (If you give the MDE in percentage points, add the \(d\): MDE in percentage points directly to the baseline CVR \(p_1\) to get the target CVR.)
Quick example
With a baseline CVR of 3% and a relative lift of 10% to detect, the target CVR is
\(p_2\): target CVR \(=\) baseline CVR (3%) \(\times\) \(1 + 10 \div 100\)
\(3 \times (1 + 10 \div 100) = 3 \times 1.1 = 3.3\ \ (3.3\%)\)
Key idea
Be careful: "a 10% improvement" can be read in two ways. As a relative %, it is "3% plus a tenth of 3% = 3.3%". In percentage points, it is "3% + 10 points = 13%". The two give completely different sample sizes. In marketing, "a 10% lift in CVR" usually refers to the relative %, and that is also this page's default. The target CVR \(p_2\) is not a forecast of what version B will actually do. It is the smallest difference you do not want to miss (the MDE). If the real difference is smaller than this, the test will likely fail to show a significant difference, even at the full sample size.
Formula for the sample size per group (two-proportion z-test)
Figure
Standard notation (the usual math form)
\(n\) \(=\) \(\bigl(\) \(z_{\alpha/2}\sqrt{2\bar{p}(1-\bar{p})}\) \(+\) \(z_{\beta}\sqrt{p_1(1-p_1)+p_2(1-p_2)}\) \(\bigr)^2\) \(\div\) \((p_2 - p_1)^2\)
In words (symbols replaced with words)
④ \(n\): sample size per group \(=\) \(\bigl(\) ① \(z_{\alpha/2}\sqrt{2\bar{p}(1-\bar{p})}\): significance term \(+\) ② \(z_{\beta}\sqrt{p_1(1-p_1)+p_2(1-p_2)}\): power term \(\bigr)^2\) \(\div\) ③ \((p_2 - p_1)^2\): the difference to detect, squared
The formula in words
① Add the significance term (the \(z_{\alpha/2}\) from the significance level × the spread when there is no difference, \(\sqrt{2\bar{p}(1-\bar{p})}\))
② and the power term (the \(z_{\beta}\) from the power × the spread when there is a difference, \(\sqrt{p_1(1-p_1)+p_2(1-p_2)}\)) and square the sum,
③ divide it by the square of the difference to detect, \((p_2 - p_1)^2\)
④ and you get the \(n\): sample size per group (round any fraction up)
Quick example
With a baseline CVR of 3% (0.03), a target of 3.3% (0.033), a two-sided test at the 5% significance level (\(z_{\alpha/2} = 1.9600\)) and 80% power (\(z_{\beta} = 0.8416\)), the sample size per group is
\(n\): sample size per group \(=\) \(\bigl(\) \(1.9600 \times \sqrt{2 \times 0.0315 \times 0.9685}\) \(+\) \(0.8416 \times \sqrt{0.03 \times 0.97 + 0.033 \times 0.967}\) \(\bigr)^2\) \(\div\) \((0.033 - 0.03)^2\)
\(\bar{p} = (0.03 + 0.033) \div 2 = 0.0315\)
\(1.9600 \times \sqrt{2 \times 0.0315 \times 0.9685} = 1.9600 \times 0.24701 \approx 0.48414\)
\(0.8416 \times \sqrt{0.03 \times 0.97 + 0.033 \times 0.967} = 0.8416 \times 0.24700 \approx 0.20788\)
\((0.48414 + 0.20788)^2 \div (0.033 - 0.03)^2 = 0.47889 \div 0.000009 \approx 53210.3 \rightarrow 53211\)
Key idea
\(\bar{p}\) is the average of the two CVRs, \((p_1 + p_2) \div 2\): the shared CVR you assume when there is no difference. \(\sqrt{2\bar{p}(1-\bar{p})}\) and \(\sqrt{p_1(1-p_1)+p_2(1-p_2)}\) both measure how much the difference in CVR bounces around by chance. The first is for "no difference" and the second is for "a real difference". Look at the formula: the sample size is divided by the square of the difference you want to detect. So halving the MDE makes the sample size about 4 times larger, and cutting it to a third makes it about 9 times larger. This is why the number of visitors climbs so fast when you want to catch even small lifts. The z-value from the significance level is \(z_{\alpha/2}\) for a two-sided test (1.9600 at 5%, 2.5758 at 1%, 1.6449 at 10%) and \(z_{\alpha}\) for a one-sided test (1.6449 at 5%, 2.3263 at 1%, 1.2816 at 10%). The z-value from the power, \(z_{\beta}\), is 0.8416 at 80% and 1.2816 at 90%. Textbooks and other tools sometimes use a simplified formula that merges the two spreads into the single \(\sqrt{2\bar{p}(1-\bar{p})}\): \(n \approx 2\bar{p}(1-\bar{p})(z_{\alpha/2}+z_{\beta})^2 \div (p_2-p_1)^2\). This calculator uses the formula on the card above (with two separate spreads) and also shows the simplified value in the result for reference. When the CVR is small, the two barely differ (53,211 vs. 53,212 in the example above). Reaching the required sample size does not make a significant result a sure effect. The formula only holds if visitors are assigned at random, you do not keep peeking at the results and stop early, and you use a stricter cutoff when you compare several metrics or several B versions at once (multiple comparisons).
Formula for the days needed (from daily visitors)
Standard notation (the usual math form)
\(D\) \(=\) \(\lceil\) \(2n\) \(\div\) \(V\) \(\rceil\)
In words (symbols replaced with words)
③ \(D\): days needed \(=\) \(\lceil\) ① \(2n\): total sample size for A and B \(\div\) ② \(V\): daily visitors \(\rceil\)
The formula in words
① Take the \(2n\): total sample size for A and B (twice the sample size per group \(n\))
② divide it by the \(V\): daily visitors and round up any fraction (\(\lceil\ \rceil\) is the symbol for rounding up),
③ and you get the \(D\): days needed
Quick example
If the sample size per group is 53,211 and the page being tested gets 2,000 visits a day, the days needed are
\(D\): days needed \(=\) \(\lceil\) \(2 \times 53211\) \(\div\) 2,000 visits \(\rceil\)
\(2 \times 53211 \div 2000 = 106422 \div 2000 = 53.211 \rightarrow 54\)
Key idea
\(\lceil\ \rceil\) is the symbol for rounding up (the ceiling function), so 53.211 days becomes 54 days. Visitors are split half and half between A and B, so enter the daily visitors for A and B together (all visits to the page being tested). Even if the math says 54 days, real tests are usually run in whole weeks. CVR changes by day of the week (weekday and weekend visitors differ), so running for 7, 14, ... days spreads the day-of-week effect evenly over both groups. In this example, 8 weeks (56 days) is a good plan. If the days needed run to several months, it is more realistic to choose a larger MDE or to test on a page with more traffic.
Formula for the power (with a fixed sample size per group)
Figure
Standard notation (the usual math form)
\(z_{1-\beta}\) \(=\) \(\bigl(\) \(|p_2 - p_1|\sqrt{n}\) \(-\) \(z_{\alpha/2}\sqrt{2\bar{p}(1-\bar{p})}\) \(\bigr)\) \(\div\) \(\sqrt{p_1(1-p_1)+p_2(1-p_2)}\)
\(1-\beta\) \(=\) \(\Phi\bigl(\) \(z_{1-\beta}\) \(\bigr)\)
In words (symbols replaced with words)
④ \(z_{1-\beta}\): z-value for the power \(=\) \(\bigl(\) ① \(|p_2 - p_1|\sqrt{n}\): difference to detect × \(\sqrt{\text{sample size per group}}\) \(-\) ② \(z_{\alpha/2}\sqrt{2\bar{p}(1-\bar{p})}\): significance term \(\bigr)\) \(\div\) ③ \(\sqrt{p_1(1-p_1)+p_2(1-p_2)}\): spread when there is a difference
⑥ \(1-\beta\): power \(=\) \(\Phi\bigl(\) ⑤ \(z_{1-\beta}\): z-value for the power \(\bigr)\)
The formula in words
① Take the difference to detect \(|p_2 - p_1|\) times the square root of the sample size per group, \(\sqrt{n}\)
② subtract the \(z_{\alpha/2}\sqrt{2\bar{p}(1-\bar{p})}\): significance term
③ divide by the \(\sqrt{p_1(1-p_1)+p_2(1-p_2)}\): spread when there is a difference
④ and you get the \(z_{1-\beta}\): z-value for the power
⑤ Use the cumulative distribution function \(\Phi\) of the standard normal distribution to find the probability of a value at or below that \(z_{1-\beta}\): z-value for the power
⑥ and you get the \(1-\beta\): power
Quick example
With a baseline CVR of 3%, a target of 3.3% and a two-sided test at the 5% significance level (\(z_{\alpha/2} = 1.9600\)), if you can only get 30,000 visitors per group, the power is
\(z_{1-\beta}\): z-value for the power \(=\) \(\bigl(\) \(0.003 \times \sqrt{30000}\) \(-\) \(1.9600 \times 0.24701\) \(\bigr)\) \(\div\) \(0.24700\)
\(0.003 \times \sqrt{30000} = 0.003 \times 173.205 \approx 0.51962\)
\((0.51962 - 0.48414) \div 0.24700 \approx 0.1436\)
\(1 - \beta = \Phi(0.1436) \approx 0.5571\ \ (55.71\%)\)
Key idea
This is the "sample size per group" formula solved again for \(z_{\beta}\). \(\Phi\) (phi) is the cumulative distribution function of the standard normal distribution. It returns the probability that a z-value is at or below a given value (\(\Phi(0.8416) \approx 0.80\), \(\Phi(1.2816) \approx 0.90\)). In the example above, the power is only about 56%. So even if version B really does reach 3.3%, a test with 30,000 visitors per group will call it "no difference" almost half the time. A test with power below 80% often misses real effects. Concluding "there was no difference" from a low-power test is jumping to conclusions. Strictly speaking, a two-sided test should also add the other tail (the chance that B comes out significantly worse). That value is tiny, so this calculator follows common practice and leaves it out.
The sample size per group for an A/B test is "(significance term + power term) squared ÷ (difference to detect) squared", rounded up. Halving the MDE makes the sample size about 4 times larger, so the trick is to decide first how large a lift you want to catch (the MDE). Divide by daily visitors to get the days needed, or fix the sample size and solve the formula again to get the power.

Symbols and terms

Symbols

\(p_1\) p one The baseline CVR (the conversion rate of the current version A). \(p\) stands for proportion or probability. In the formulas, 3% is written as the decimal 0.03.
\(p_2\) p two The target CVR (the conversion rate you hope version B reaches). It is \(p_1 (1 + r/100)\) for a relative % and \(p_1 + d\) for percentage points.
\(\bar{p}\) p-bar The average of the two CVRs, \((p_1 + p_2) \div 2\). It is the shared CVR you assume when there is no difference. In statistics, the bar on top marks an average.
\(r\) r The MDE given as a relative %. It stands for ratio and tells you by what percent the baseline CVR improves (3% to 3.3% gives \(r = 10\)).
\(d\) d The MDE given in percentage points. It stands for difference and is the plain difference in CVR (3% to 3.3% gives \(d = 0.3\) points).
\(p_2 - p_1\) p two minus p one The difference to detect, in percentage points written as a decimal (0.3 points is 0.003). The sample size is divided by its square, so the smaller the difference, the faster the sample size grows.
\(\alpha\) alpha The significance level: the probability of calling a difference real when there is none (the probability of a Type I error). 5% (0.05) is the usual choice. Alpha is the first Greek letter, and in statistics it is used for the probability of the "first" kind of error.
\(\beta\) beta The probability of missing a difference that is really there (the probability of a Type II error). Beta is the second Greek letter, used for the probability of the "second" kind of error.
\(1-\beta\) one minus beta The power: the probability of finding a difference when there really is one. 80% or 90% is the usual choice.
\(z_{\alpha/2}\) z sub alpha over two The z-value for the significance level \(\alpha\) in a two-sided test (the point with \(\alpha/2\) of the standard normal distribution above it). It is 1.9600 at 5%, 2.5758 at 1% and 1.6449 at 10%. A one-sided test uses \(z_{\alpha}\) instead (1.6449 at 5%).
\(z_{\beta}\) z sub beta The z-value for the power \(1-\beta\) (the point with probability \(1-\beta\) of the standard normal distribution below it). It is 0.8416 at 80% and 1.2816 at 90%.
\(z_{1-\beta}\) z sub one minus beta In the power formula, the z-value worked back from the sample size. Put it into \(\Phi\) and you get the power. (It is the same thing as \(z_{\beta}\), written from the side of the value you are solving for.)
\(\Phi\) capital phi The cumulative distribution function of the standard normal distribution. It returns the probability that a z-value is at or below a given value: \(\Phi(1.96) \approx 0.975\) and \(\Phi(0.8416) \approx 0.80\). The capital Greek letter phi is the usual symbol for this function.
\(n\) n The sample size per group (for each of version A and version B). It stands for number. The total for A and B is \(2n\).
\(\sqrt{\ \ }\) square root The square root: the positive number that gives the number inside when squared, for example \(\sqrt{30000} \approx 173.205\). It appears when finding the spread (standard deviation) of a CVR.
\(V\) V Daily visitors (all visits to the page being tested, A and B together). It stands for visits.
\(D\) D The days needed. It stands for days.
\(\lceil\ \rceil\) ceiling The symbol for rounding up (the ceiling function). It gives the smallest whole number that is greater than or equal to the number inside (\(\lceil 53.211 \rceil = 54\)).

Terms

A/B test An experiment that splits visitors at random between the current version (A) and a new version (B) over the same period, to see which performs better (for example, has a higher CVR). It is used to check with data, not hunches, whether a change such as a button color, a headline, the way a price is shown or the layout of a landing page really helps.
CVR (conversion rate) The share of visitors or clicks that end in a goal action (a conversion), such as a purchase or a sign-up. In the formulas on this page, version A's CVR is \(p_1\) and version B's is \(p_2\).
MDE (minimum detectable effect) The smallest lift you want to be able to detect. It is the first value to decide when you plan an A/B test. The smaller it is, the faster the sample size grows (it is inversely proportional to the square of the difference), so choose the smallest lift that would matter to your business.
significance level The probability \(\alpha\) of the mistake you are willing to accept - seeing a chance difference and calling it real when there is no difference. 5% is standard, which allows a false alarm about 1 time in 20. A stricter level (1%) needs a larger sample size.
statistical power The probability \(1-\beta\) of finding a difference when there really is one. 80% is standard, which accepts missing a real lift about 1 time in 5. If a low-power test says "no difference", that is not evidence that there is no difference.
Type I error The mistake of calling a difference real when there is none (a false positive). The probability you allow for it is the significance level \(\alpha\). It is the flip side of a Type II error (a miss): pushing one down tends to push the other up. The most basic way to make both small is a larger sample size.
Type II error The mistake of missing a difference that is really there (a false negative). Its probability is \(\beta\), and \(1-\beta\) is the power. It is the flip side of a Type I error (a false alarm): pushing one down tends to push the other up. The most basic way to make both small is a larger sample size.
null hypothesis The hypothesis "versions A and B have the same CVR" that the test tries to reject. The test asks how unusual the data would be if the null hypothesis were true. If it is unusual enough, the null hypothesis is rejected and the result is "a difference".
alternative hypothesis The hypothesis "versions A and B have different CVRs" that the test tries to show. It is accepted when the null hypothesis (no difference) is rejected.
two-sided test A test that counts both "B is higher" and "B is lower" as a difference. It is used in a normal A/B test, where you do not know which way it will go. It needs a larger sample size than a one-sided test, but it also catches B doing worse.
one-sided test A test that only looks for "B is higher". It needs a smaller sample size than a two-sided test, but it cannot catch B doing worse, so you must choose it before the test starts. (Switching to one-sided after seeing the results is not valid.)
two-proportion z-test A test that checks whether two groups differ in a proportion (such as CVR), using a z-value (the standardized difference) and the standard normal distribution. The formulas on this page find the sample size for this test. The normal approximation works when each group has enough conversions (as a rule of thumb, at least 5 to 10).
standard normal distribution The normal distribution with mean 0 and standard deviation 1 (a symmetric bell shape). A z-value is a point on its horizontal axis, and each z-value matches one probability, as in "the area to the right of 1.96 is 2.5%".
cumulative distribution function A function that returns the probability of getting a given value or less. The one for the standard normal distribution is written \(\Phi\) and appears at the end of the power formula. In Excel it is NORM.S.DIST (with TRUE as the second argument), and in Python it is NormalDist().cdf.
standard deviation (spread) How spread out data is. The spread of a proportion \(p\) is proportional to \(\sqrt{p(1-p)}\). The \(\sqrt{2\bar{p}(1-\bar{p})}\) and \(\sqrt{p_1(1-p_1)+p_2(1-p_2)}\) in the formulas are the spread of the difference between the two groups' CVRs.
multiple comparisons Comparing several B versions or several metrics at the same time. The more comparisons you make, the more likely it is that at least one looks significant by chance, so an adjustment is needed, such as dividing the significance level by the number of comparisons (the Bonferroni correction).
peeking (early stopping) Checking the results again and again before the required sample size is reached and stopping the moment the result looks significant. It makes chance differences easy to catch, so a test meant to have a 5% significance level actually makes wrong calls far more often. The formulas on this page assume that you fix the sample size first and wait until you reach it.

Good to know before you start

Here is what helps you use the calculation on this page with real understanding, not just by pressing the button.
If you get stuck, going back over these topics is the quickest way forward.

Percentages (Grade 6)
  • Being able to switch between percentages and decimals (\(3\% = 0.03\), \(0.033 = 3.3\%\))
  • Knowing the difference between a relative % increase (3% up by a tenth is 3.3%) and a percentage-point increase (3% + 0.3 points = 3.3%)
Square roots (Grade 8)
  • Knowing that \(\sqrt{a}\) is the positive number that gives \(a\) when squared, and being able to find values such as \(\sqrt{30000} \approx 173.2\) on a calculator
  • Knowing that when you divide by a square, halving the number you divide by makes the answer 4 times larger (\((1/2)^2 = 1/4\))
Basic probability (Grade 7 to high school)
  • Being able to treat a rate such as "30 purchases out of 1,000 visits" as a probability (a number from 0 to 1)
  • Having a feel for how a difference can appear by chance when there is none (getting 7 heads in 10 coin flips is not unusual)
The normal distribution and z-scores (high school statistics)
  • Knowing that the normal distribution is a symmetric bell shape and that about 95% of it lies within 1.96 standard deviations of the mean
  • Knowing that a z-score tells how many standard deviations a value is from the mean
The idea of hypothesis testing (AP Statistics or intro college statistics)
  • Knowing how a test works: you set up the null hypothesis "no difference", and if the data would be unusual under it, you conclude "there is a difference"
  • Knowing that the significance level \(\alpha\) (the chance of calling a difference real when there is none) and the power \(1-\beta\) (the chance of finding a real difference) are about two different mistakes
Rounding up (Grade 4)
  • Knowing why numbers of people and days are rounded up to whole numbers, as in 53,210.3 visitors becoming 53,211 and 53.211 days becoming 54

How to calculate it in Excel

Copy the whole table below and paste it into cell A1 in Excel. It works as is.
Table to find the target CVR (relative %)
Baseline CVR p1 (%) 3
Relative lift r (%) 10
Target CVR p2 (%) =B1*(1+B2/100)
Table to find the sample size per group
Baseline CVR p1 (%) 3
Target CVR p2 (%) 3.3
Significance level α (%, two-sided) 5
Power 1−β (%) 80
z_α/2 =NORM.S.INV(1-B3/100/2)
z_β =NORM.S.INV(B4/100)
Average CVR p̄ =(B1+B2)/2/100
Sample size per group n (rounded up) =ROUNDUP((B5*SQRT(2*B7*(1-B7))+B6*SQRT(B1/100*(1-B1/100)+B2/100*(1-B2/100)))^2/((B2-B1)/100)^2,0)
Table to find the days needed
Sample size per group n 53211
Daily visitors V (A + B total) 2000
Days needed D (rounded up) =ROUNDUP(2*B1/B2,0)
Table to find the power (fixed sample size per group)
Baseline CVR p1 (%) 3
Target CVR p2 (%) 3.3
Significance level α (%, two-sided) 5
Sample size per group n 30000
z_α/2 =NORM.S.INV(1-B3/100/2)
Average CVR p̄ =(B1+B2)/2/100
z_1−β =(ABS(B2-B1)/100*SQRT(B4)-B5*SQRT(2*B6*(1-B6)))/SQRT(B1/100*(1-B1/100)+B2/100*(1-B2/100))
Power 1−β (%) =NORM.S.DIST(B7,TRUE)*100
After pasting, the upper rows are your inputs and the last row is calculated automatically.
The first table shows 3.3 in B3 (a target CVR of 3.3%), the second shows 53211 in B8 (the sample size per group), the third shows 54 in B3 (the days needed) and the fourth shows about 55.71 in B8 (a power of 55.71%).
NORM.S.INV turns a probability into a z-value, and NORM.S.DIST (with TRUE as the second argument) turns a z-value into a probability. For a one-sided test, delete the "/2" from the z_α/2 formula so that it reads =NORM.S.INV(1-B3/100). ROUNDUP(…,0) rounds up.

How to calculate it in Google Sheets

Copy the whole table below and paste it into cell A1 in Google Sheets. It works as is.
Table to find the target CVR (relative %)
Baseline CVR p1 (%) 3
Relative lift r (%) 10
Target CVR p2 (%) =B1*(1+B2/100)
Table to find the sample size per group
Baseline CVR p1 (%) 3
Target CVR p2 (%) 3.3
Significance level α (%, two-sided) 5
Power 1−β (%) 80
z_α/2 =NORM.S.INV(1-B3/100/2)
z_β =NORM.S.INV(B4/100)
Average CVR p̄ =(B1+B2)/2/100
Sample size per group n (rounded up) =ROUNDUP((B5*SQRT(2*B7*(1-B7))+B6*SQRT(B1/100*(1-B1/100)+B2/100*(1-B2/100)))^2/((B2-B1)/100)^2,0)
Table to find the days needed
Sample size per group n 53211
Daily visitors V (A + B total) 2000
Days needed D (rounded up) =ROUNDUP(2*B1/B2,0)
Table to find the power (fixed sample size per group)
Baseline CVR p1 (%) 3
Target CVR p2 (%) 3.3
Significance level α (%, two-sided) 5
Sample size per group n 30000
z_α/2 =NORM.S.INV(1-B3/100/2)
Average CVR p̄ =(B1+B2)/2/100
z_1−β =(ABS(B2-B1)/100*SQRT(B4)-B5*SQRT(2*B6*(1-B6)))/SQRT(B1/100*(1-B1/100)+B2/100*(1-B2/100))
Power 1−β (%) =NORM.S.DIST(B7,TRUE)*100
The same formulas as in Excel work as is (Google Sheets also has NORM.S.INV, NORM.S.DIST and ROUNDUP). Copy the whole table, paste it into cell A1, and replace the inputs in column B with your own numbers.

How to calculate it in Python

from math import sqrt, ceil
from statistics import NormalDist

base_cvr = 3.0          # baseline CVR (%)
relative_mde = 10.0     # MDE to detect (relative %)
alpha = 0.05            # significance level (two-sided)
power = 0.80            # power
daily_visits = 2000     # daily visitors (A + B total)

p1 = base_cvr / 100
p2 = p1 * (1 + relative_mde / 100)          # target CVR
z_alpha = NormalDist().inv_cdf(1 - alpha / 2)  # for a one-sided test, use 1 - alpha
z_beta = NormalDist().inv_cdf(power)
p_bar = (p1 + p2) / 2

n_exact = (z_alpha * sqrt(2 * p_bar * (1 - p_bar))
           + z_beta * sqrt(p1 * (1 - p1) + p2 * (1 - p2))) ** 2 / (p2 - p1) ** 2
n_per_group = ceil(n_exact)                  # sample size per group (rounded up)
days = ceil(2 * n_per_group / daily_visits)  # days needed

print(f"Target CVR: {p2 * 100:.4g}%")
print(f"Sample size per group: {n_per_group} (total {2 * n_per_group})")
print(f"Days needed: {days}")

# Solving the other way: power with 30,000 visitors per group
n_fixed = 30000
z_power = (abs(p2 - p1) * sqrt(n_fixed) - z_alpha * sqrt(2 * p_bar * (1 - p_bar))) \
          / sqrt(p1 * (1 - p1) + p2 * (1 - p2))
print(f"Power with {n_fixed} per group: {NormalDist().cdf(z_power) * 100:.2f}%")
Runs with the standard library only (statistics.NormalDist needs Python 3.8 or later). In this example, the target CVR is 3.3%, the sample size per group is 53211 (106422 in total), the days needed are 54, and the power with 30,000 visitors per group is about 55.71%. Change the values at the top and run it.

How to write it in LaTeX and other math languages (copy and paste)

Formula for the target CVR (the CVR you hope version B reaches)
p₂ = p₁ × (1 + r/100)
p_2 = p_1 \left(1 + \dfrac{r}{100}\right)
<math xmlns="http://www.w3.org/1998/Math/MathML" display="block">
  <mrow>
    <msub><mi>p</mi><mn>2</mn></msub>
    <mo>=</mo>
    <msub><mi>p</mi><mn>1</mn></msub>
    <mo>(</mo><mn>1</mn><mo>+</mo><mfrac><mi>r</mi><mn>100</mn></mfrac><mo>)</mo>
  </mrow>
</math>
p_2 = p_1 (1 + r/100)
p1*(1 + r/100)
p2 := p1*(1 + r/100);
p2 = p1*(1 + r/100);
p_2 = p_1 (1 + r/100)
Formula for the sample size per group (two-proportion z-test)
n = (z_α/2 √(2p̄(1 − p̄)) + z_β √(p₁(1 − p₁) + p₂(1 − p₂)))² ÷ (p₂ − p₁)²
n = \dfrac{\left( z_{\alpha/2}\sqrt{2\bar{p}(1-\bar{p})} + z_{\beta}\sqrt{p_1(1-p_1)+p_2(1-p_2)} \right)^2}{(p_2 - p_1)^2}
<math xmlns="http://www.w3.org/1998/Math/MathML" display="block">
  <mrow>
    <mi>n</mi>
    <mo>=</mo>
    <mfrac>
      <msup>
        <mrow>
          <mo>(</mo>
          <msub><mi>z</mi><mrow><mi>&#x3B1;</mi><mo>/</mo><mn>2</mn></mrow></msub>
          <msqrt><mn>2</mn><mover><mi>p</mi><mo>&#xAF;</mo></mover><mo>(</mo><mn>1</mn><mo>&#x2212;</mo><mover><mi>p</mi><mo>&#xAF;</mo></mover><mo>)</mo></msqrt>
          <mo>+</mo>
          <msub><mi>z</mi><mi>&#x3B2;</mi></msub>
          <msqrt>
            <msub><mi>p</mi><mn>1</mn></msub><mo>(</mo><mn>1</mn><mo>&#x2212;</mo><msub><mi>p</mi><mn>1</mn></msub><mo>)</mo>
            <mo>+</mo>
            <msub><mi>p</mi><mn>2</mn></msub><mo>(</mo><mn>1</mn><mo>&#x2212;</mo><msub><mi>p</mi><mn>2</mn></msub><mo>)</mo>
          </msqrt>
          <mo>)</mo>
        </mrow>
        <mn>2</mn>
      </msup>
      <msup>
        <mrow><mo>(</mo><msub><mi>p</mi><mn>2</mn></msub><mo>&#x2212;</mo><msub><mi>p</mi><mn>1</mn></msub><mo>)</mo></mrow>
        <mn>2</mn>
      </msup>
    </mfrac>
  </mrow>
</math>
n = (z_(alpha/2) sqrt(2 bar p (1 - bar p)) + z_beta sqrt(p_1(1 - p_1) + p_2(1 - p_2)))^2 / (p_2 - p_1)^2
Ceiling[(za*Sqrt[2*pbar*(1 - pbar)] + zb*Sqrt[p1*(1 - p1) + p2*(1 - p2)])^2/(p2 - p1)^2]
n := ceil((za*sqrt(2*pbar*(1 - pbar)) + zb*sqrt(p1*(1 - p1) + p2*(1 - p2)))^2/(p2 - p1)^2);
n = ceil((za*sqrt(2*pbar*(1 - pbar)) + zb*sqrt(p1*(1 - p1) + p2*(1 - p2)))^2/(p2 - p1)^2);
n = (z_(α/2) √(2p̄(1 − p̄)) + z_β √(p_1(1 − p_1) + p_2(1 − p_2)))^2/(p_2 − p_1)^2
Formula for the days needed (from daily visitors)
D = ⌈2n ÷ V⌉
D = \left\lceil \dfrac{2n}{V} \right\rceil
<math xmlns="http://www.w3.org/1998/Math/MathML" display="block">
  <mrow>
    <mi>D</mi>
    <mo>=</mo>
    <mo>&#x2308;</mo>
    <mfrac><mrow><mn>2</mn><mi>n</mi></mrow><mi>V</mi></mfrac>
    <mo>&#x2309;</mo>
  </mrow>
</math>
D = |~ (2n)/V ~|
Ceiling[2*n/v]
days := ceil(2*n/V);
D = ceil(2*n/V);
D = ⌈2n/V⌉
Formula for the power (with a fixed sample size per group)
1 − β = Φ((|p₂ − p₁|√n − z_α/2 √(2p̄(1 − p̄))) ÷ √(p₁(1 − p₁) + p₂(1 − p₂)))
1 - \beta = \Phi\left( \dfrac{|p_2 - p_1|\sqrt{n} - z_{\alpha/2}\sqrt{2\bar{p}(1-\bar{p})}}{\sqrt{p_1(1-p_1)+p_2(1-p_2)}} \right)
<math xmlns="http://www.w3.org/1998/Math/MathML" display="block">
  <mrow>
    <mn>1</mn><mo>&#x2212;</mo><mi>&#x3B2;</mi>
    <mo>=</mo>
    <mi>&#x3A6;</mi>
    <mo>(</mo>
    <mfrac>
      <mrow>
        <mo>|</mo><msub><mi>p</mi><mn>2</mn></msub><mo>&#x2212;</mo><msub><mi>p</mi><mn>1</mn></msub><mo>|</mo>
        <msqrt><mi>n</mi></msqrt>
        <mo>&#x2212;</mo>
        <msub><mi>z</mi><mrow><mi>&#x3B1;</mi><mo>/</mo><mn>2</mn></mrow></msub>
        <msqrt><mn>2</mn><mover><mi>p</mi><mo>&#xAF;</mo></mover><mo>(</mo><mn>1</mn><mo>&#x2212;</mo><mover><mi>p</mi><mo>&#xAF;</mo></mover><mo>)</mo></msqrt>
      </mrow>
      <msqrt>
        <msub><mi>p</mi><mn>1</mn></msub><mo>(</mo><mn>1</mn><mo>&#x2212;</mo><msub><mi>p</mi><mn>1</mn></msub><mo>)</mo>
        <mo>+</mo>
        <msub><mi>p</mi><mn>2</mn></msub><mo>(</mo><mn>1</mn><mo>&#x2212;</mo><msub><mi>p</mi><mn>2</mn></msub><mo>)</mo>
      </msqrt>
    </mfrac>
    <mo>)</mo>
  </mrow>
</math>
1 - beta = Phi((|p_2 - p_1| sqrt(n) - z_(alpha/2) sqrt(2 bar p (1 - bar p))) / sqrt(p_1(1 - p_1) + p_2(1 - p_2)))
CDF[NormalDistribution[0, 1], (Abs[p2 - p1]*Sqrt[n] - za*Sqrt[2*pbar*(1 - pbar)])/Sqrt[p1*(1 - p1) + p2*(1 - p2)]]
power := Statistics[CDF](Normal(0, 1), (abs(p2 - p1)*sqrt(n) - za*sqrt(2*pbar*(1 - pbar)))/sqrt(p1*(1 - p1) + p2*(1 - p2)));
power = normcdf((abs(p2 - p1)*sqrt(n) - za*sqrt(2*pbar*(1 - pbar)))/sqrt(p1*(1 - p1) + p2*(1 - p2)));
1 − β = Φ((|p_2 − p_1| √n − z_(α/2) √(2p̄(1 − p̄)))/√(p_1(1 − p_1) + p_2(1 − p_2)))

How to have ChatGPT  do the calculation

You are an assistant for planning A/B tests. Do the following calculation by actually running Python code, and base your answer only on the numbers from the execution result (do not answer by mental math or guessing).

The current page's CVR (the baseline CVR) is 3%. I want to be able to detect a relative lift of 10% with version B (a CVR of 3.3%).
With a 5% significance level (two-sided test) and 80% power, find each of the following:
1. The target CVR p2 (p2 = p1 × (1 + 10/100))
2. The sample size per group n (the two-proportion z-test formula: n = (z_{α/2}·√(2p̄(1−p̄)) + z_β·√(p1(1−p1)+p2(1−p2)))² ÷ (p2−p1)², with p̄ = (p1+p2)/2, rounded up) and the total sample size for A and B
3. The days needed (rounded up) with 2,000 daily visitors (A and B together)
4. The power if only 30,000 visitors per group are available (1−β = Φ((|p2−p1|·√n − z_{α/2}·√(2p̄(1−p̄))) ÷ √(p1(1−p1)+p2(1−p2))))

Use statistics.NormalDist for the z-values and the cumulative distribution function, and show the formulas you used and the numbers from the execution result.

How to Use
  1. 1
    Enter your numbers
    Type the numbers you want to calculate with into the input fields
  2. 2
    Calculate
    Press the "Calculate" button
  3. 3
    Check the result
    The result appears on the spot. The same page also explains the idea behind the calculation and the formula
  DataChef Features
Easy and Free
Unlimited conversions for free.
No technical knowledge required.
Intuitive and user-friendly operation.
No Registration Required
Available immediately after access.
Can be used without registering personal information.
Safe and Secure
Fully SSL encrypted communication.
Automatic file deletion by clicking "download".
Fast
High-speed site access
and rapid file conversion.
No Watermark
No watermark.
No attribution required.
Commercial Use Available
Free for commercial use.
No need to contact us for commercial use permission.