A/B Testing in Marketing
A/B testing is a method of showing two versions of a page, email or ad to two random groups at the same time and measuring which one gets more of a chosen action, such as orders or sign-ups. Version A is the current one, called the control, and version B contains a single change, which is why marketers also call the method split testing.
- Hypothesis: A written guess about what change will help, for which visitors, and why.
- Random split: Visitors are assigned to A or B by chance, so the two groups are alike in every other way.
- Sample size: The number of visitors each version needs before the result can be trusted, decided before the test starts.
- Statistical significance: A check that the difference is larger than the normal random ups and downs.
- One change at a time: If A and B differ in five different ways, you will never know which particular change mattered.
This lesson follows one example: a D2C millet snacks brand from Chennai that sells ragi chips and jowar puffs on its own website. About 2,000 people reach its product pages every day, and 3 percent of them place an order. Many shoppers ask on WhatsApp when their parcel will arrive, so the team wants to test showing the delivery date near the Add to Cart button.
Why A/B Testing Matters
Opinions about design are cheap, and even experienced marketers regularly guess wrong about customer behaviour, so A/B testing replaces the loudest opinion in the meeting with evidence from real shoppers. Because both versions run during the same days, a festival, a payday or a rainy weekend affects both groups equally, which a simple before and after comparison cannot promise. Small gains also add up: a page that converts slightly better earns more from every rupee already spent on ads, which is the core idea behind conversion rate optimization.
Step-by-Step Framework for A/B Testing
Step 1: Find the Problem with Data
Start from evidence rather than personal taste, and look at where visitors drop off in GA4, watch session recordings and heatmaps in a tool such as Microsoft Clarity, and read customer questions. The snacks brand saw that many shoppers opened the shipping policy page from the product page and then left without buying.
Step 2: Write a Hypothesis
A good hypothesis has three parts: the change, the expected effect, and the reason. The team wrote: "If we show 'Delivered by Thursday' above the Add to Cart button, more visitors will order, because many visitors leave the page to check delivery times." Writing the reason down helps the team learn something useful even when the variation loses.
Step 3: Choose One Primary Metric
Pick one number that decides the winner before the test begins, which for the snacks brand is the order conversion rate, meaning orders divided by visitors who saw the product page. Watch secondary metrics, such as average order value, only to make sure the winner does not cause harm elsewhere. The purchase event must already be tracked correctly, as covered in GA4 events and conversions.
Step 4: Work Out the Sample Size
Sample size depends on your current conversion rate and the smallest improvement worth finding. A common rule of thumb, for 95 percent significance and an 80 percent chance of catching a real effect, is shown below, where p is the current conversion rate and d is the change you want to detect, both written as decimals.
visitors per version = 16 x p x (1 - p) / d^2
p = 0.03 (3 percent now) d = 0.006 (3.0 to 3.6 percent)
16 x 0.03 x 0.97 = 0.4656
0.006 x 0.006 = 0.000036
0.4656 / 0.000036 = 12,933, so about 13,000 per versionThe test needs about 26,000 visitors in total. At 2,000 visitors a day that takes 13 days, so the team plans a full two weeks to include both weekends. A smaller improvement needs far more traffic: detecting 3.0 to 3.3 percent would need about four times as many visitors, because halving d multiplies the sample by four. Online sample size calculators give more exact figures using the same ideas.
Step 5: Build the Variant and Split Traffic
Build version B with only the delivery date added, and split visitors 50/50 at random. Use a testing tool or your store platform's built-in experiments, and make sure a returning visitor always sees the same version. For ads, platforms such as Meta and Google Ads offer their own experiment features, covered in creative testing on Meta.
Step 6: Run It Without Peeking
Let the experiment reach its planned sample size before deciding anything, because checking every day and stopping the first time B looks ahead makes a false winner much more likely, because random swings are largest early on. Do not change prices, stock or ads for only one version while the test runs.
Step 7: Read the Result and Check Significance
Compare conversion rates, then check whether the gap is bigger than random chance would normally produce. The example below walks through the numbers, although most testing tools do this calculation automatically and show a confidence or probability figure.
Step 8: Decide, Roll Out and Record
If B wins clearly, make it the default for every visitor, and if the result is unclear, keep A, since the change did not prove its value. Either way, record the hypothesis, dates, sample, result and lesson in a shared test log, so the team does not repeat old tests.
A/B Testing Template
Test name: Delivery date above Add to Cart
Page or asset: Product pages, mobile and desktop
Problem seen: Many visitors open the shipping page, then leave
Hypothesis: If we show the delivery date above the button,
more visitors will order, because delivery time
is a common worry
Primary metric: Orders / product page visitors
Guard metrics: Average order value, refund requests
Current rate (p): 3.0% Smallest lift worth finding: 3.6%
Sample needed: 13,000 per version (26,000 total)
Planned length: 14 days, starting on a Monday
Result: A ___% B ___% z = ___ Decision: ___
Lesson learned: ___Example: The Millet Snacks Test
After 14 days, each version had reached 13,000 visitors.
| Version | Visitors | Orders | Conversion rate |
|---|---|---|---|
| A: current page | 13,000 | 390 | 3.0% |
| B: delivery date shown | 13,000 | 468 | 3.6% |
- Lift: B converts at 3.6 percent and A at 3.0 percent, a gain of 0.6 points. Divided by the starting 3.0 percent, that is a 20 percent relative lift.
- Pooled rate: Together the two versions had 858 orders from 26,000 visitors, a rate of 3.3 percent.
- Normal wobble: The standard error, the size of the random swing you would expect, is the square root of 0.033 x 0.967 x (2 / 13,000), which is about 0.0022, or 0.22 points.
- Z score: The gap of 0.006 divided by 0.0022 gives a z score of about 2.7. A z score above 1.96 means the result is significant at the 95 percent level, so this gap is very unlikely to be pure chance.
- Decision: The team made the delivery date standard on all product pages, and logged the lesson that delivery certainty matters to its buyers.
A significant result on the website still only tells you about that page. Whether the brand's ads create extra sales at all is a different question, answered by incrementality testing.
Mistakes to Avoid in A/B Testing
- Stopping early: Ending the experiment the moment B pulls ahead, before reaching the planned sample.
- Too many changes: Testing a new headline, photo, price and button colour together, then not knowing which one worked.
- Too little traffic: Hoping to detect tiny gains on a page with a few hundred visitors a week.
- Uneven timing: Running A in one week and B in the next, so a sale or holiday decides the result.
- Ignoring segments you planned: A change can help mobile buyers and hurt desktop ones, so plan key segments in advance instead of hunting for them afterwards.
- Chasing the wrong metric: A bright promotional banner can increase clicks while actually lowering orders.
For page-level ideas worth testing, see landing page optimization.
How AI Changes A/B Testing
What AI Automates Now
AI tools can suggest hypotheses from heatmaps and reviews, write many headline and image variants in minutes, and explain test results in plain words. Some testing tools and ad platforms also shift traffic automatically toward the version that is doing better, a method often called a multi-armed bandit.
What Still Needs a Human
People still choose which problem is worth testing, make sure each variant is accurate and on brand, and decide whether a 20 percent lift is worth the cost of the change. A person also has to spot when a test was broken, for example when one version failed to load on some phones.
Risk to Watch
Generating fifty variants is easy, but traffic is not, and splitting visitors fifty ways means no version gets enough visitors for a clear answer. AI summaries can also declare a winner before the result reaches significance, so always check the sample and the z score yourself.
Do It with AI
Use this prompt to turn a problem into a test plan. It works in ChatGPT, Claude or Gemini.
You are a conversion specialist for an online store in India. Page to test: [page and what it sells] Current conversion rate: [for example 3%] Daily visitors to this page: [number] Problem I have seen: [what analytics, recordings or customer messages show] 1. Write three hypotheses in the form "If we [change], then [metric] will [rise or fall], because [reason]". 2. For the strongest one, describe version B in detail, changing only one element. 3. Using visitors per version = 16 x p x (1 - p) / d^2, calculate the sample size for detecting a 20 percent relative lift, showing each step. 4. Say how many days the test needs at my traffic, rounded up to full weeks. Do not invent results, benchmarks or customer quotes.
- Collect evidence of the problem from analytics, recordings and customer questions.
- Run the prompt and pick the hypothesis the team believes in most.
- Recalculate the sample size by hand or with a calculator.
- Build version B, run the test for the planned days, and record the result in the test log.
Check Before You Use It
- Facts: Recheck every calculation, and make sure delivery dates, prices and offers in version B are true.
- Brand fit: Variant copy must sound like the brand, not like a generic template.
- Compliance: Both versions must be honest; never test fake countdown timers, false stock warnings or hidden charges, which can count as dark patterns under Indian consumer rules.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. The millet snacks store has a 3 percent conversion rate and wants to detect a rise to 3.6 percent. Using the rule of thumb 16 x p x (1 - p) / d squared, about how many visitors does each version need?
Frequently Asked Questions
What is A/B testing in simple words?
A/B testing shows two versions of the same page, email or ad to two random groups of people at the same time, then compares which version gets more of the action you want. Because the groups are random and run together, the difference is likely caused by the change itself.
How long should an A/B test run?
Run it until each version has reached the sample size you calculated before starting, and for at least one or two full weeks so that weekdays and weekends are both included. Stopping early because one version looks ahead is the most common way to get a false winner.
What does 95 percent statistical significance mean?
It means that if the two versions truly performed the same, a difference as large as the one you saw would appear by chance less than 5 percent of the time. It does not mean there is a 95 percent chance the winner is better, and it says nothing about whether the gain is large enough to matter.
What is the difference between A/B testing and multivariate testing?
An A/B test changes one thing and compares two versions. A multivariate test changes several elements at once and compares many combinations, which needs far more traffic. Most small businesses should run simple A/B tests.
Can I A/B test with low website traffic?
Yes, but only big changes will show a clear result. With little traffic, test bold ideas such as a new offer or a much shorter form, use a conversion that happens often, such as add to cart, or test on ad platforms where reach is larger.
Related Articles
- Conversion Rate Optimization (CRO)Conversion rate optimization step by step: find funnel leaks, rank ideas, test fixes and use a CRO audit checklist, shown with a Bengaluru SaaS startup.
- Landing Page Optimization with AILanding page optimization in eight steps: match the ad, keep one goal, build trust, speed up mobile and test AI variants, with a Jaipur restaurant example.
- Creative Testing on MetaCreative testing on Meta made fair: one change at a time, enough results before deciding and a simple test log, shown with a Kota NEET coaching institute.
- Incrementality Testing and Lift TestsIncrementality testing shows how many sales your ads truly caused. Learn holdouts, geo tests and conversion lift through a worked Diwali sale example.
- Microsoft Clarity (Heatmaps)Microsoft Clarity tutorial: install it, read heatmaps and session recordings, mask sensitive data and handle consent, with a Hyderabad travel agency case.
- GA4 Events and ConversionsGA4 events explained: automatic, enhanced, recommended and custom events, and how to mark key events, with a coaching institute example and an AI prompt.