Ecommerce A/B Testing: The Store Owner's Guide for 2026
Most stores that say they A/B test are running a lottery with their own traffic. Here's the whole system: what to test first, how much traffic you need, how bandits roll out winners, and how to keep a test journal your team learns from.

Most Shopify stores that say they "do A/B testing" are running a lottery with their own traffic.
They test a green button against an orange one. They check the dashboard on day three, see a 14% lift, crown a winner, and move on. Six weeks later the store earns exactly what it earned before, and nobody can explain why.
So here's the short answer. Ecommerce AB testing means showing two versions of a page to a split of your visitors and keeping the version that earns more money per visitor. Done right, it works like this:
- Test the offer, the price framing or the headline first.
- Pick your sample size before the test starts.
- Judge winners on revenue per visitor, not clicks.
- Run one test per page at a time.
- Write every result in a test journal.
The rest of this guide is how to do it without fooling yourself, with Black Friday eight weeks away.
A test that can't change what the buyer is deciding about can't change what the buyer pays you.
What should you test first: the offer, the price or the headline?
Here's my hot take. The order you test in matters more than the tool you test with.
Every store has three levels of change. The top level changes what the buyer is deciding about: the offer and the price. The middle level changes how fast the buyer understands it: the headline, the first screen, the order of the proof. The bottom level changes how the page looks: button colors, fonts, icons, the badge row.
Almost every beginner starts at the bottom, because it's easy. That's also why it rarely matters. A buyer who wasn't sure your $64 serum was worth it doesn't suddenly become sure because the button went from black to orange.
So start at the top.
Test the offer first
The offer is the thing the buyer is actually saying yes to. One bottle, or a 60-day supply with a free travel size. The single mattress topper, or the topper plus two pillows as a sleep kit. A 30-day guarantee, or "sleep on it for 100 nights."
When you change the offer, you change the math of the decision. That's why offer tests swing hardest, in both directions.
A good offer test changes one thing about the deal and keeps the page the same. Bundle versus single. Guarantee A versus guarantee B. Free gift versus no gift. If you've never tested a bundle, start there, because it moves average order value, and average order value is the half of revenue per visitor most stores never touch.
Test the price framing second
I said price framing, not price. Most founders are scared to test price because they picture raising it and watching sales collapse. But the words around the price move more money than the number itself.
Cost per use. The comparison to what the buyer pays now. The anchor you show first. We break down three framing moves in our guide to how price framing changes Shopify conversion, and the anchoring side of it lives in the price anchoring strategy that makes a price feel fair.
If you do want to test the number, go in with eyes open. Our breakdown of whether charm pricing actually lifts conversion covers why ending in 9 is the wrong digit to obsess over. And know your tool: some testing apps can split prices on Shopify and some can't (more on that in the tools section).
Test the headline third
The headline is the first sentence of the sales conversation. If it describes the product ("Organic Cotton Sheet Set") instead of the outcome the buyer wants ("Sleep cool without waking up at 3 a.m. kicking off the covers"), you're asking the rest of the page to do all the work.
Headline tests are cheap to build and fast to read, because every visitor sees the headline. The first 800 pixels decide whether anyone reads the rest, which is the whole argument in what to put above the fold on a Shopify product page. For the copy itself, how to write Shopify product page copy that sells walks through it line by line.
What to leave for later
Button colors. Font sizes. Icon styles. Trust badge rows. Sticky add to cart bars. These aren't useless, they're just small. When we pulled the published tests on sticky bars for our sticky add to cart button research, the real wins were smaller than the app listings promise. Small wins need huge samples to detect, and huge samples are the one thing a $30,000 a month store doesn't have.
Which tools do you need: testing apps, heatmaps and session recordings?
You need two kinds of tools. One kind tells you what to test. The other kind runs the test.
Here's the thing most tool roundups skip: the first kind is free now, and it's where the ideas come from. The second kind is where the money gets spent, and it's useless without the first.
| Tool | What it does | What it can test or show | Watch out for |
|---|---|---|---|
| Shopify Rollouts | Native experiments inside the Shopify admin | Two themes or two checkout setups against each other | Experiments need the Grow plan or higher; no price or discount tests |
| Intelligems | Shopify testing app focused on price and offer tests | Product prices, subscription prices, discounts, shipping rates, free shipping thresholds, content | Built around profit per visitor, so set your costs up correctly first |
| Microsoft Clarity | Free heatmaps and session recordings | Clicks, scroll depth, rage clicks, dead clicks, recordings, AI summaries | Shows what happened, never why; you still need a hypothesis |
| Hotjar (part of Contentsquare) | Heatmaps, recordings and surveys | Click maps, scroll maps, move maps | Move maps are desktop only; mobile is where most stores lose buyers |
| Optimizely and similar platforms | Full experimentation suites | A/B tests, multivariate tests, bandits | Built for enterprise traffic and budgets |
The native option on Shopify
On June 5, 2026, Shopify added A/B testing to Rollouts. You can now run an experiment between two entirely different themes or two checkout setups without installing anything. By default an experiment splits traffic 50/50 and ends 90 days after it starts, and your store has to be on the Grow plan or higher to run one.
It's a real step forward. But read the fine print on the report. According to Shopify's rollout analytics documentation, a theme experiment tracks conversion rate over time, bounce rate, reached checkout rate and add to cart rate. Average order value and gross sales show up in theme launches and checkout rollouts, but not in the theme experiment list. That means a theme test can "win" on conversion rate while quietly losing money per visitor, and the report won't flag it.
The fix is easy. Pull orders for each arm yourself and do the revenue per visitor math by hand before you call a winner.
Why a whole-theme test is a blunt tool
Rollouts compares whole experiences, not single elements. That's great for a redesign. It's terrible for learning. If theme B wins, you don't know if it was the headline, the image order or the faster load time. And as we found in whether a new Shopify theme lifts conversions, the theme itself is rarely what's holding a page back.
Where AI actually helps today
AI has changed two jobs in testing, and I'll be specific about both.
First, reading behavior. Clarity now ships AI summaries and an AI chat on top of its heatmaps and recordings, so you can ask what's happening on a page instead of watching 200 recordings. Treat those summaries as a starting point for a hypothesis, never as the hypothesis itself.
Second, writing variants. A language model can draft 10 headline angles in a minute. That speed is dangerous if you test all 10, because 10 small variants split your traffic into 10 tiny samples. Use AI to generate many options, then pick the 2 that make the biggest bet and test those.
How do you read heatmaps and session recordings without guessing?
A heatmap is a picture of where people clicked, tapped and scrolled on a page, added up across many visits. A session recording is a replay of one visit. You need both, because the heatmap shows the pattern and the recording shows the person.
The trap is staring at them until you see what you already believed. So give yourself three questions and answer them in order.
Question 1: Where does the scroll die?
Open the scroll map on mobile first. Find the line where the color drops off hard. Now look at what sits just below that line. If your reviews, your guarantee or your "why this price" section lives under the drop, most buyers never saw it. You don't need a test to know that. You need to move it up, then test the new order.
Contentsquare's guide explains that scroll maps show the percentage of people who reach each point on the page. That's the single most useful number on any heatmap, and most founders never look at it.
Question 2: Where do people click on things that aren't buttons?
Microsoft defines dead clicks as clicks on an element that give no feedback in a reasonable amount of time, and rage clicks as multiple rapid clicks in the same small area. Both mean the same thing on a product page: the buyer expected something to happen and nothing did.
On stores I've audited, the usual suspects are the same every time. Product images that look zoomable and aren't. A shipping line that looks like a link. The star rating, which buyers tap expecting to jump to reviews. Every one of those is a free test idea, straight from the buyer's thumb.
Question 3: What do the buyers who leave do right before they leave?
Filter your recordings to mobile sessions on your best-selling product page that didn't add to cart. Watch 20. Not 200. Twenty.
Take notes in plain words: "scrolled to the price, back up to the images, opened the size chart, left." If six people open the size chart and leave, your test is the size guidance, not the button.
For the full list of places a page bleeds, see how to find conversion rate leaks on your Shopify store. If most people leave before they scroll at all, start with how to cut product page bounce rate. And if they add to cart but don't buy, the leak is further down, which is the story in what your add to cart rate is really telling you.
A heatmap tells you where. A recording tells you who. Neither tells you why. The why is your hypothesis, and your hypothesis is what you test.
Turn what you saw into one sentence
Every test should start with a sentence in this shape: "Because we saw [behavior], we believe [change] will lift [metric] for [page]." For example: "Because 6 of 20 recordings opened the size chart and left, we believe putting a fit finder above the add to cart button will lift revenue per visitor on the linen pants page."
If you can't fill in that sentence, you're not ready to test. You're ready to watch more recordings.
How much traffic does a test need? Sample size, duration and significance in plain words
This is the part everybody skips, and it's the part that decides whether your test means anything.
Three words, in plain English:
- Sample size is how many visitors each version needs to see before the result is trustworthy.
- Duration is how long it takes your store to send that many visitors.
- Significance is how sure you are that the difference is real and not luck.
The rule that saves you from yourself
Pick your sample size before you launch. Then don't call the test until you reach it.
Here's why. Evan Miller, whose sample size calculator half the testing world uses, showed in How Not to Run an A/B Test that checking results over and over and stopping the moment they look significant wrecks your math. In his worst case, testing after every visitor at a 5% significance threshold produces a false positive rate of 26.1%, more than five times what you think you're getting.
Read that again. One in four "winners" would be noise.
That's the day-three "14% lift" from the top of this post. It was never real. It was a coin landing heads three times in a row.
How big a sample do you need?
Miller gives a rule of thumb for a test with a 5% significance level and 80% power: visitors per version equals 16 times the variance, divided by the square of the difference you want to detect. For a conversion rate, the variance is the rate times one minus the rate.
Here's the math done for you on a page converting at 2%:
| Lift you want to detect | Conversion rate goes from | Visitors needed per version | Total visitors for an A/B test |
|---|---|---|---|
| 10% | 2.0% to 2.2% | 78,400 | 156,800 |
| 20% | 2.0% to 2.4% | 19,600 | 39,200 |
| 30% | 2.0% to 2.6% | about 8,711 | about 17,422 |
| 50% | 2.0% to 3.0% | 3,136 | 6,272 |
The formula: 16 x (0.02 x 0.98) / (difference squared). For a 20% lift the difference is 0.004, so 16 x 0.0196 / 0.000016 = 19,600 visitors per version.
Evan Miller's rule of thumb: 16 x p(1-p) / difference squared, at 5% significance and 80% power.
Look at the shape of that chart. Halving the lift you want to detect roughly quadruples the traffic you need. That's the whole reason small tests die on small stores. A button color change might really be worth 3%. You'll never have the traffic to prove it.
Plug your own numbers into Evan Miller's sample size calculator before every test. It takes a minute.
How long should a test run?
Divide the total visitors you need by the daily visitors to the page being tested. Then round up to full weeks.
Run the math on a store like this (a hypothetical): the hero product page gets 12,000 visitors a month, about 400 a day, converting at 2%. To detect a 20% lift, you need 39,200 visitors total. That's 98 days, roughly 14 weeks. To detect a 50% lift, you need 6,272 visitors. That's 16 days, which you round up to 3 full weeks.
Two rules on duration:
- Never run a test less than two full weeks, even if you hit your sample on day nine. Weekday buyers and weekend buyers behave differently, and paydays move behavior too.
- Never let a test run past the point where the page, the ads or the season change under it. A test that straddles a big promo is two tests glued together.
Judge the winner on revenue per visitor
A test can lift conversion rate and lose money. It happens every time a variant adds a discount, a cheaper bundle or a free shipping threshold that sits too low.
So score every arm on revenue per visitor: conversion rate times average order value. If version A converts at 2.0% with a $70 average order value, revenue per visitor is $1.40. If version B converts at 2.3% with a $58 average order value, revenue per visitor is $1.33. B "won" on conversion rate and lost 7 cents on every visitor. On 10,000 visitors, that's $14,000 versus $13,340, a $660 loss the conversion rate report calls a win.
If that trade-off is new to you, start with revenue per visitor vs conversion rate and how to calculate Shopify revenue per visitor step by step. The deeper case is in why your conversion rate is a vanity metric.
Bandits and Thompson sampling: how do automatic winner rollouts work?
A classic A/B test is patient. It holds a 50/50 split until the sample size is reached, even if one version is clearly losing. Every visitor sent to the loser is a sale you paid to learn.
A multi-armed bandit is impatient on purpose. The name comes from slot machines ("one-armed bandits"). Picture a row of machines, each with a different payout you can't see. You want to find the best one while losing as little money as possible looking for it.
How Thompson sampling decides where traffic goes
Thompson sampling is the most common way bandits make that choice. Here's the plain version.
For each version of your page, the algorithm keeps a range of believable conversion rates based on what it's seen so far. Early on the ranges are wide, because it hasn't seen much. For every new visitor it draws one random guess from each version's range and sends the visitor to whichever version drew the highest guess.
Watch what happens. A version that's truly better draws high guesses more often, so it gets more visitors. A version that's losing still gets some traffic while its range is wide, so it gets a fair chance to prove itself. As data piles up, the ranges narrow and traffic flows toward the winner on its own.
Optimizely's documentation describes the same mechanic: for each variation it uses the observed conversions and visitors to build a distribution, samples those distributions many times, and allocates traffic according to how often each one wins.
The honest trade-off
Bandits aren't free money. Optimizely is blunt about the costs in that same page: bandits ignore statistical significance, so their results page doesn't report it. They generally take longer than an A/B test to separate winners from losers. And if you pause a variation or change the goal mid-test, the bandit has to start from scratch.
So here's how I think about the choice:
| Question | Classic A/B test | Bandit with Thompson sampling |
|---|---|---|
| Main goal | Learn exactly how much better one version is | Earn the most money while the test runs |
| Traffic split | Fixed until the sample size is reached | Shifts toward the leader as data comes in |
| Cost of a losing version | Half your traffic sees it the whole time | Losing versions get less traffic over time |
| Significance report | Yes | Usually not |
| Best for | Big decisions you'll build on: offer, price, page structure | Headlines, images, promos and short sale windows |
| Worst for | Short campaigns that end before the sample is reached | Proving a result to a skeptical partner or investor |
How we run it
This is how the testing side of RevenueFlows AI works. Split tests use Thompson sampling, so traffic moves toward the version that's earning more while the test is still live. When one version clearly wins, it rolls out to every visitor automatically. Nobody has to remember to log in and flip a switch, which (I'll admit it) is the step that used to get skipped most in my own stores.
And the winner doesn't just disappear into the page. It gets written down. More on that in the test journal section, because that's where the compounding happens.
How do you A/B test on low traffic?
Most stores doing $10,000 to $50,000 a month don't have 39,200 visitors to spare on one page. That's fine. Low traffic doesn't mean no testing. It means different testing.
Rule 1: Only test big swings
On a page with 5,000 visitors a month, you can't detect a 10% lift in any reasonable time. You can detect a 50% lift in about six weeks. So only test changes that could plausibly move the needle by half: a new offer, a new price frame, a rewritten first screen. If a change can't possibly be worth 30% or more, don't test it. Ship it or skip it.
Rule 2: Test where the traffic already is
Don't split 20,000 monthly visitors across 40 product pages and test each one. Put the test on the one page that gets most of the traffic, usually the hero product. One test with enough traffic beats five tests with none.
Rule 3: Pick a metric closer to the click
If purchases are too rare to count, test on add to cart rate first. It happens several times more often than a purchase, so you reach a readable sample sooner. Then confirm the winner on revenue per visitor over the following month. A proxy metric is a shortcut, not a verdict.
Rule 4: Use a bandit when you'd rather earn than prove
On a small store, a bandit is often the better tool. You may never get a clean significance number, but you'll lose fewer sales to the weaker version while you wait. For a headline or a hero image, that's the right trade.
Rule 5: Sometimes skip the test
If a recording shows buyers tapping a broken image zoom 30 times a day, fix it. You don't need a four-week experiment to prove that a broken thing should work. Save testing for real questions where you honestly don't know the answer.
Low traffic doesn't mean you can't test. It means you can't afford to test small ideas.
What are the most common testing mistakes?
I've made most of these. Every store I've audited has made at least three.
Calling it early. Day three, 14% lift, winner declared. Evan Miller's 26.1% false positive number is the cost of this habit. Pick the sample size first. Wait for it.
Testing tiny things on tiny traffic. A 3% effect needs about 871,000 visitors per version at a 2% conversion rate. No store doing $40,000 a month will ever finish that test.
Judging on the wrong number. Conversion rate goes up, average order value goes down, revenue per visitor drops, and the team celebrates. Always score on revenue per visitor.
Running overlapping tests on the same page. If the headline test and the bundle test both run on the hero page at once, you can't tell which one moved the number. Shopify's own experiments give each experiment dedicated traffic, so each visitor only sees one. Hold yourself to the same rule.
Changing things mid-test. New ad creative, a price change, a new email flow pointing at the page. Any of these resets the audience your test was measuring. Note every outside change in the journal, and if it's big, restart.
Ignoring mobile. Check your device split. On most DTC stores I look at, phones win by a wide margin. A variant that wins on desktop can lose on mobile, and a blended result hides it. Check the split by device before rolling out.
Believing the novelty. Returning visitors click on what's new because it's new. That bump fades. It's one more reason to run at least two full weeks.
Testing without a hypothesis. "Let's see what happens" isn't a test. It's a guess with a traffic bill. If you can't write the "because we saw" sentence, go back to the recordings.
Faking urgency to juice a test. A countdown timer that resets on refresh can win a two-week test and cost you trust for a year. We made that case in how to create urgency that's actually true.
Never writing the result down. The most expensive mistake on this list. Which brings me to the journal.
How do you build a test journal your whole team learns from?
Here's what I see at almost every store that says it tests. Someone ran a test last spring. Nobody remembers if it won. The person who ran it left. Now the new hire wants to test the same headline.
A test that isn't written down didn't happen. Worse, it'll happen again, and you'll pay for it twice.
What goes in every entry
Keep one row per test. Here are the fields that matter:
| Field | What to write | Example |
|---|---|---|
| Test name | Page plus the change | Hero serum page: 60-day bundle vs single bottle |
| Hypothesis | The "because we saw" sentence | Because 9 of 20 recordings scrolled to the price and left, we believe a bundle framed as cost per day will lift revenue per visitor |
| Versions | What each arm showed | A: single bottle $64. B: 60-day bundle $108 with cost per day shown |
| Dates and traffic | Start, end, visitors per arm | 8 weeks, 14,100 and 14,000 visitors |
| Results | Conversion rate, average order value and revenue per visitor for each arm | A: 1.8%, $64, $1.15. B: 1.7%, $86, $1.46 |
| Verdict | Winner, loser or no clear difference | B wins on revenue per visitor |
| What we learned | The reusable lesson, in one sentence | Our buyers accept a higher price when it's framed per day |
| Outside changes | Anything that happened during the test | New Meta creative launched Sept 15 |
| Next test | What this result points to | Test cost per day framing on the single bottle too |
(That's a hypothetical example to show the format. Check the arithmetic: 1.8% x $64 = $1.152, and 1.7% x $86 = $1.462.)
The column that matters most
"What we learned" is the only column anyone will read in a year. Write it as a rule about your buyers, not a fact about a page. "The bundle won" is a fact. "Our buyers accept a higher price when it's framed per day" is a rule you can use on the next 12 product pages.
Log losers with the same care. A clean loss tells you what your buyers don't care about, so you never test it again.
Keep a learnings page on top of the journal
After 10 or 15 tests, patterns show up. Put them on one page the whole team reads before they propose anything new: what has worked for this store, what hasn't, and what's still unknown.
That's exactly how we set it up inside RevenueFlows AI. Every store keeps its own test journal, and a "what has worked for this store" view pulls the lessons together so the next test starts from what your buyers have already told you. And the ideas for those next tests come from the session recordings and heatmaps, which closes the loop: watch, guess, test, write it down, repeat.
Watch what compounding does to the money
Here's the math on why the journal is worth it. Run the numbers on a store like this (a hypothetical skincare brand, not a client): the hero serum page converts at 1.8% with a $64 average order value. Revenue per visitor is $1.15. On 10,000 visitors, that's $11,520.
Over one quarter, two tests win and get written down. A bundle test lifts average order value to $78. A headline test lifts conversion rate to 2.2%. Now revenue per visitor is 2.2% x $78 = $1.72. On the same 10,000 visitors, that's $17,160.
That's $5,640 more a month from the same traffic. Neither test was magic. They stacked, because the second test started where the first one left off.
Before: 1.8% x $64. After a bundle test and a headline test: 2.2% x $78.
For real numbers, here's the flagship. A bedding brand we rebuilt went from a conversion rate of 1.0% and an average order value of $125 (revenue per visitor $1.25) to 3.5% and $231 (revenue per visitor $8.10). On 10,000 visitors, that's $12,500 before and $81,000 after. See the full case study numbers. Real client numbers, not typical results, and not a promise of what your store will do.
How should you test around Black Friday?
Black Friday is the worst week of the year to learn something and the best week of the year to earn something. Treat it that way.
The traffic is huge. Adobe reported that U.S. shoppers spent $11.8 billion online on Black Friday 2025 and $14.25 billion on Cyber Monday, with $44.2 billion across the five days from Thanksgiving through Cyber Monday. Adobe also reported peak Cyber Monday discounts of 31% on electronics, 28% on toys and 25% on apparel.
But the people are different. Holiday buyers are deal hunters and gift buyers. They arrive with a discount already in mind, often buying for someone else. A headline that wins with them may lose with your normal buyer in February.
The Black Friday testing calendar
Now through mid-October: finish your page tests. Any test that needs 14 weeks is already too late for this year. Pick tests that can reach their sample size by the third week of October, and roll the winners out.
The two weeks before Black Friday: freeze the page. Don't start new page tests. Your traffic mix is shifting daily as early-deal shoppers show up, and a test that straddles that shift is measuring two audiences at once.
Sale week: test the deal, not the page. This is where a bandit earns its keep. You don't have time to wait for significance on a five-day sale, but you can let Thompson sampling push traffic toward the stronger offer while the sale is live. Percent off versus dollars off. Tiered discount versus a free gift. Bundle deal versus sitewide code. Our full breakdown of A/B testing during Black Friday walks through the offer math and the freeze date.
Before you test any discount, know what it costs. A 25% discount on a product with a 60% margin eats a big share of your profit per order, and you need a lot more orders to break even. Our research on whether discounts really lift Shopify conversion shows why the deepest discount rarely wins on profit.
December and January: don't trust the sale-week winner. Log every Black Friday result with a big note in the "outside changes" column: holiday traffic. Then retest the winning idea on normal traffic in January before you make it permanent.
Black Friday tells you what deal hunters want. January tells you what your customers want. Don't confuse the two.
Start here: the full reading list
This guide is the hub. Every section above has a deeper post behind it. Here they are, grouped by what you're trying to fix.
Choosing what to test (offer, price and headline)
- How price framing changes Shopify conversion
- The price anchoring strategy that makes a price feel fair
- Whether charm pricing actually lifts conversion
- What to put above the fold on a Shopify product page
- How to write Shopify product page copy that sells
Finding test ideas in heatmaps and recordings
- How to find conversion rate leaks on your Shopify store
- How to cut product page bounce rate
- What your add to cart rate is really telling you
Measuring winners the right way
- Revenue per visitor vs conversion rate
- How to calculate Shopify revenue per visitor step by step
- Why your conversion rate is a vanity metric
Tests that usually disappoint (read before you run them)
Holiday and promo testing
- Whether discounts really lift Shopify conversion
- How to create urgency that's actually true
- A/B testing during Black Friday: what to test and what to freeze
FAQ
What should I A/B test first on an ecommerce store? The offer, then the price framing, then the headline on your best-selling product page. Those change what the buyer is deciding about. Button colors can wait.
How long should an ecommerce A/B test run? Until it reaches the sample size you picked before launch, and never less than two full weeks. At 400 visitors a day and a 2% conversion rate, spotting a 20% lift takes about 14 weeks.
How much traffic do I need to A/B test? About 19,600 visitors per version to spot a 20% lift at a 2% conversion rate, and about 3,100 per version for a 50% lift. Under 10,000 monthly visitors on a page, test only bold changes.
What's the difference between an A/B test and a multi-armed bandit? An A/B test keeps a fixed split so you learn exactly how much better a version is. A bandit using Thompson sampling moves traffic toward the leader as it goes, so you lose fewer sales but get a less precise answer.
Can I A/B test on Shopify without an app? Yes. Shopify Rollouts runs experiments between two themes or checkout setups on the Grow plan and up. It can't test prices, and its theme experiment report doesn't show revenue per visitor, so do that math yourself.
Should I run A/B tests during Black Friday? Freeze new page tests two weeks before. During the sale, test the deal itself with a bandit. Then retest any winner on normal traffic in January before you keep it.
What to do next
Open Microsoft Clarity (or whatever recording tool you already have), filter to mobile visitors on your best-selling product page who didn't add to cart, and watch 20 sessions this week. Write down the one thing at least five of them did before leaving. That's your first hypothesis, and your first real test.
Book Your Profit Audit
Before you spend 14 weeks of traffic on a test, find out which part of your page is actually leaking revenue per visitor. Get your free profit audit and we'll show you where the biggest gap is, then show you how to rebuild a high-converting product sales page in less than 15 minutes.
Or go here to check it out → revenueflows.ai
P.S. A test you called on day three is a coin flip with a spreadsheet attached. Pick the sample size, wait for it, and write down what your buyers told you.
Frequently asked questions
What should I A/B test first on an ecommerce store?
Test the offer first, then the price framing, then the headline on your best-selling product page. Those three change what the buyer is deciding about, so they move revenue per visitor far more than button colors, fonts or icon swaps ever will.
How long should an ecommerce A/B test run?
Long enough to hit the sample size you picked before launch, and never less than two full weeks so every weekday and weekend is counted. A product page with 400 visitors a day needs about 14 weeks to reliably spot a 20% lift at a 2% conversion rate.
How much traffic do I need to A/B test?
At a 2% conversion rate you need roughly 19,600 visitors per version to spot a 20% lift, and about 3,100 per version to spot a 50% lift. If one page gets under 10,000 visitors a month, test bold changes only or use a bandit.
What's the difference between an A/B test and a multi-armed bandit?
An A/B test holds a fixed split until it reaches its sample size, so you learn exactly how much better one version is. A bandit, often run with Thompson sampling, shifts traffic toward the version that's winning while the test runs, so you lose fewer sales but get a less precise answer.
Can I A/B test on Shopify without an app?
Yes. Since June 2026, Shopify Rollouts can run experiments between two themes or two checkout setups on the Grow plan or higher. It can't test prices or discount logic, and its theme experiment report shows conversion rate and add to cart rate, not revenue per visitor.
Should I run A/B tests during Black Friday?
Don't start new page tests in the two weeks before Black Friday, and never trust a winner from the sale week as a year-round winner. Holiday shoppers behave like a different audience. Use the week to test discount structure with a bandit, and log everything for next year.

