…
Skip to content
Topics
On this page

A/B Testing in Marketing

A/B testing is a method of showing two versions of a page, email or ad to two random groups at the same time and measuring which one gets more of a chosen action, such as orders or sign-ups. Version A is the current one, called the control, and version B contains a single change, which is why marketers also call the method split testing.

  • Hypothesis: A written guess about what change will help, for which visitors, and why.
  • Random split: Visitors are assigned to A or B by chance, so the two groups are alike in every other way.
  • Sample size: The number of visitors each version needs before the result can be trusted, decided before the test starts.
  • Statistical significance: A check that the difference is larger than the normal random ups and downs.
  • One change at a time: If A and B differ in five different ways, you will never know which particular change mattered.
A/B test: 26,000 visitors split at random, version B converts at 3.6 percent against 3.0 percentOver 14 days, 26,000 product page visitors are split at random, 50/50. Version A, the current page, gets 13,000 visitors and 390 orders, a 3.0 percent conversion rate. Version B, which shows the delivery date, gets 13,000 visitors and 468 orders, a 3.6 percent rate. That is a 20 percent relative lift with a z score of about 2.7, above 1.96, so it is significant at the 95 percent level.26,000 visitorsProduct pages14 days50%50%A: current page13,000 visitors390 ordersB: delivery date shown13,000 visitors468 ordersConversion rate3.0%3.6%20% relative liftz = 2.7, above 1.96Significant at 95%Same days, random visitors, one change: the delivery date near Add to Cart
A/B test: 26,000 visitors split at random, version B converts at 3.6 percent against 3.0 percent

This lesson follows one example: a D2C millet snacks brand from Chennai that sells ragi chips and jowar puffs on its own website. About 2,000 people reach its product pages every day, and 3 percent of them place an order. Many shoppers ask on WhatsApp when their parcel will arrive, so the team wants to test showing the delivery date near the Add to Cart button.

Why A/B Testing Matters

Opinions about design are cheap, and even experienced marketers regularly guess wrong about customer behaviour, so A/B testing replaces the loudest opinion in the meeting with evidence from real shoppers. Because both versions run during the same days, a festival, a payday or a rainy weekend affects both groups equally, which a simple before and after comparison cannot promise. Small gains also add up: a page that converts slightly better earns more from every rupee already spent on ads, which is the core idea behind conversion rate optimization.

Step-by-Step Framework for A/B Testing

Step 1: Find the Problem with Data

Start from evidence rather than personal taste, and look at where visitors drop off in GA4, watch session recordings and heatmaps in a tool such as Microsoft Clarity, and read customer questions. The snacks brand saw that many shoppers opened the shipping policy page from the product page and then left without buying.

Step 2: Write a Hypothesis

A good hypothesis has three parts: the change, the expected effect, and the reason. The team wrote: "If we show 'Delivered by Thursday' above the Add to Cart button, more visitors will order, because many visitors leave the page to check delivery times." Writing the reason down helps the team learn something useful even when the variation loses.

Step 3: Choose One Primary Metric

Pick one number that decides the winner before the test begins, which for the snacks brand is the order conversion rate, meaning orders divided by visitors who saw the product page. Watch secondary metrics, such as average order value, only to make sure the winner does not cause harm elsewhere. The purchase event must already be tracked correctly, as covered in GA4 events and conversions.

Step 4: Work Out the Sample Size

Sample size depends on your current conversion rate and the smallest improvement worth finding. A common rule of thumb, for 95 percent significance and an 80 percent chance of catching a real effect, is shown below, where p is the current conversion rate and d is the change you want to detect, both written as decimals.

Example
visitors per version = 16 x p x (1 - p) / d^2

p = 0.03 (3 percent now)      d = 0.006 (3.0 to 3.6 percent)
16 x 0.03 x 0.97 = 0.4656
0.006 x 0.006    = 0.000036
0.4656 / 0.000036 = 12,933, so about 13,000 per version

The test needs about 26,000 visitors in total. At 2,000 visitors a day that takes 13 days, so the team plans a full two weeks to include both weekends. A smaller improvement needs far more traffic: detecting 3.0 to 3.3 percent would need about four times as many visitors, because halving d multiplies the sample by four. Online sample size calculators give more exact figures using the same ideas.

Step 5: Build the Variant and Split Traffic

Build version B with only the delivery date added, and split visitors 50/50 at random. Use a testing tool or your store platform's built-in experiments, and make sure a returning visitor always sees the same version. For ads, platforms such as Meta and Google Ads offer their own experiment features, covered in creative testing on Meta.

Step 6: Run It Without Peeking

Let the experiment reach its planned sample size before deciding anything, because checking every day and stopping the first time B looks ahead makes a false winner much more likely, because random swings are largest early on. Do not change prices, stock or ads for only one version while the test runs.

Step 7: Read the Result and Check Significance

Compare conversion rates, then check whether the gap is bigger than random chance would normally produce. The example below walks through the numbers, although most testing tools do this calculation automatically and show a confidence or probability figure.

Step 8: Decide, Roll Out and Record

If B wins clearly, make it the default for every visitor, and if the result is unclear, keep A, since the change did not prove its value. Either way, record the hypothesis, dates, sample, result and lesson in a shared test log, so the team does not repeat old tests.

A/B Testing Template

Example
Test name:        Delivery date above Add to Cart
Page or asset:    Product pages, mobile and desktop
Problem seen:     Many visitors open the shipping page, then leave
Hypothesis:       If we show the delivery date above the button,
                  more visitors will order, because delivery time
                  is a common worry
Primary metric:   Orders / product page visitors
Guard metrics:    Average order value, refund requests
Current rate (p): 3.0%     Smallest lift worth finding: 3.6%
Sample needed:    13,000 per version (26,000 total)
Planned length:   14 days, starting on a Monday
Result:           A ___%   B ___%   z = ___   Decision: ___
Lesson learned:   ___

Example: The Millet Snacks Test

After 14 days, each version had reached 13,000 visitors.

VersionVisitorsOrdersConversion rate
A: current page13,0003903.0%
B: delivery date shown13,0004683.6%
  • Lift: B converts at 3.6 percent and A at 3.0 percent, a gain of 0.6 points. Divided by the starting 3.0 percent, that is a 20 percent relative lift.
  • Pooled rate: Together the two versions had 858 orders from 26,000 visitors, a rate of 3.3 percent.
  • Normal wobble: The standard error, the size of the random swing you would expect, is the square root of 0.033 x 0.967 x (2 / 13,000), which is about 0.0022, or 0.22 points.
  • Z score: The gap of 0.006 divided by 0.0022 gives a z score of about 2.7. A z score above 1.96 means the result is significant at the 95 percent level, so this gap is very unlikely to be pure chance.
  • Decision: The team made the delivery date standard on all product pages, and logged the lesson that delivery certainty matters to its buyers.

A significant result on the website still only tells you about that page. Whether the brand's ads create extra sales at all is a different question, answered by incrementality testing.

Mistakes to Avoid in A/B Testing

  • Stopping early: Ending the experiment the moment B pulls ahead, before reaching the planned sample.
  • Too many changes: Testing a new headline, photo, price and button colour together, then not knowing which one worked.
  • Too little traffic: Hoping to detect tiny gains on a page with a few hundred visitors a week.
  • Uneven timing: Running A in one week and B in the next, so a sale or holiday decides the result.
  • Ignoring segments you planned: A change can help mobile buyers and hurt desktop ones, so plan key segments in advance instead of hunting for them afterwards.
  • Chasing the wrong metric: A bright promotional banner can increase clicks while actually lowering orders.

For page-level ideas worth testing, see landing page optimization.

How AI Changes A/B Testing

What AI Automates Now

AI tools can suggest hypotheses from heatmaps and reviews, write many headline and image variants in minutes, and explain test results in plain words. Some testing tools and ad platforms also shift traffic automatically toward the version that is doing better, a method often called a multi-armed bandit.

What Still Needs a Human

People still choose which problem is worth testing, make sure each variant is accurate and on brand, and decide whether a 20 percent lift is worth the cost of the change. A person also has to spot when a test was broken, for example when one version failed to load on some phones.

Risk to Watch

Generating fifty variants is easy, but traffic is not, and splitting visitors fifty ways means no version gets enough visitors for a clear answer. AI summaries can also declare a winner before the result reaches significance, so always check the sample and the z score yourself.

Do It with AI

Use this prompt to turn a problem into a test plan. It works in ChatGPT, Claude or Gemini.

Prompt for ChatGPT, Claude or Gemini

You are a conversion specialist for an online store in India. Page to test: [page and what it sells] Current conversion rate: [for example 3%] Daily visitors to this page: [number] Problem I have seen: [what analytics, recordings or customer messages show] 1. Write three hypotheses in the form "If we [change], then [metric] will [rise or fall], because [reason]". 2. For the strongest one, describe version B in detail, changing only one element. 3. Using visitors per version = 16 x p x (1 - p) / d^2, calculate the sample size for detecting a 20 percent relative lift, showing each step. 4. Say how many days the test needs at my traffic, rounded up to full weeks. Do not invent results, benchmarks or customer quotes.

  1. Collect evidence of the problem from analytics, recordings and customer questions.
  2. Run the prompt and pick the hypothesis the team believes in most.
  3. Recalculate the sample size by hand or with a calculator.
  4. Build version B, run the test for the planned days, and record the result in the test log.

Check Before You Use It

  • Facts: Recheck every calculation, and make sure delivery dates, prices and offers in version B are true.
  • Brand fit: Variant copy must sound like the brand, not like a generic template.
  • Compliance: Both versions must be honest; never test fake countdown timers, false stock warnings or hidden charges, which can count as dark patterns under Indian consumer rules.

Quick Quiz

Pick an answer to check yourself. Nothing is saved.

Question 1 / 3

  1. 1. The millet snacks store has a 3 percent conversion rate and wants to detect a rise to 3.6 percent. Using the rule of thumb 16 x p x (1 - p) / d squared, about how many visitors does each version need?

Frequently Asked Questions

What is A/B testing in simple words?

A/B testing shows two versions of the same page, email or ad to two random groups of people at the same time, then compares which version gets more of the action you want. Because the groups are random and run together, the difference is likely caused by the change itself.

How long should an A/B test run?

Run it until each version has reached the sample size you calculated before starting, and for at least one or two full weeks so that weekdays and weekends are both included. Stopping early because one version looks ahead is the most common way to get a false winner.

What does 95 percent statistical significance mean?

It means that if the two versions truly performed the same, a difference as large as the one you saw would appear by chance less than 5 percent of the time. It does not mean there is a 95 percent chance the winner is better, and it says nothing about whether the gain is large enough to matter.

What is the difference between A/B testing and multivariate testing?

An A/B test changes one thing and compares two versions. A multivariate test changes several elements at once and compares many combinations, which needs far more traffic. Most small businesses should run simple A/B tests.

Can I A/B test with low website traffic?

Yes, but only big changes will show a clear result. With little traffic, test bold ideas such as a new offer or a much shorter form, use a conversion that happens often, such as add to cart, or test on ad platforms where reach is larger.