4 A/B Tests Worth Running Now So Black Friday Isn't a Guess
By Muhammed Tüfekyapan
The Black Friday offer in most stores was decided in about ninety seconds. Someone said "same as last year." Someone else said "maybe a little deeper this time." And the single biggest margin decision of the quarter was settled by a memory nobody in the room could source. There is no document where 20 percent won. There is just a number that has been winning by default since 2023.
The inheritance feels reasonable because the offer worked. The store survived the weekend, revenue spiked, and depth debates feel unanswerable without data. So last year's settings get new dates, and the planning meeting moves on to creative. But here is the break: the offer was never one decision. It is four. How deep the discount goes. How long the window stays open. How many visitors ever see it. And where the spend thresholds sit. Your Black Friday offer has four dials, and right now all four are set to a memory.
September is the last month the traffic is both large enough and ordinary enough to turn each dial and hear the click. By late October the visitor mix turns promotional. In November, every session you spend learning is a session you could have spent earning. By the end of this you will have four fully specified tests, one per dial, each with a hypothesis, variants, one primary KPI, a traffic split, a stop date, and the decision rule written before the test starts. Start with the rule that separates a test from a slower way of guessing.
A Test Without a Decision Rule Is Just a Slower Guess
Most pre-season "testing" is one variant left running until somebody forms a feeling about it. That is not measurement. It is delay with a dashboard. A real test is four specs written before launch: a falsifiable hypothesis, two or three variants, one primary KPI, and a stop rule with an if/then decision attached. And the KPI matters more than people think, because the number most dashboards default to, conversion rate, is structurally biased toward the deepest discount. A test without a chosen KPI has already chosen the wrong one.
September traffic is the last large sample of the year that still behaves like ordinary customers. It is the last month the traffic still looks like ordinary customers, which is what makes these four tests readable now and unreadable in five weeks.
The Price of a Guessed Dial
Run the arithmetic on the dial most stores guess: depth. Take a store expecting 40,000 sessions across the four Black Friday days, an $85 average order, and a 45 percent gross margin before the discount. The scoreboard that matters is gross profit per exposed session: conversion rate multiplied by average order value multiplied by your margin after the discount. Arm A at 15 percent converts 6.0 percent of sessions: 0.06 x $85 x 0.30 is $1.53 of gross profit per session, or $61,200 across the weekend. Arm B at 20 percent converts better, 6.4 percent: 0.064 x $85 x 0.25 is $1.36 per session, or $54,400. The variant that wins the conversion rate contest loses $6,800 of gross profit. Nobody ran this test, so the store shipped 20 percent, celebrated the higher conversion rate, and never saw the bill.
Four Specs Before Anything Goes Live
Every test below carries the same four specs. The hypothesis is one falsifiable sentence, like "15 percent converts visitors likely to leave without purchasing nearly as well as 20 percent, and keeps the margin." The variants are two or three, never five, because each added arm stretches the calendar. The primary KPI is chosen before launch, and for anything touching discount depth it is gross profit per exposed session, not conversion rate. The stop rule pairs a volume floor with a calendar date, roughly 100 orders per arm or two full weekends, whichever comes first. And the decision rule is written as an if/then: "if the 15 percent arm holds gross profit per session within 5 percent of the 20 percent arm, we ship 15 percent."
| Testing the way most stores do it | Testing with a decision rule | |
|---|---|---|
| What gets tested | Whatever is easiest to change | The four settings that price the whole weekend |
| KPI | Decided afterward, by whatever moved | Chosen before launch, profit per exposed session |
| Stop rule | When someone remembers to look | Pre-set: orders per arm plus a calendar date |
| Decision | A meeting in November | An if/then written in September |
| What reaches Black Friday | Last year's settings with new dates | Four tested dials |
Four dials, four tests. Any setting that reaches November without a test behind it is a guess wearing last year's confidence.
Test the Offer Itself: Depth and Duration
Depth and duration are the two dials that set what the offer costs and what it says. Both are inherited in almost every store, usually a guess that worked well enough to become a tradition. And both tests must run on the traffic the store would actually discount in November, which means walk-away customers, not the whole store.
Test 1: The Depth Dial
Hypothesis: "15 percent converts visitors likely to leave without purchasing nearly as well as 20 percent, and keeps the margin." Variants: 10, 15, and 20 percent off, run on offer-eligible traffic only, split in even thirds. Stores under roughly 500 orders a month should collapse this to two arms, 15 against 20, so it reads inside three weekends. Primary KPI: gross profit per exposed session, using the Section 1 formula. Readability window: about 100 orders per arm, which for a store doing 1,000 orders a month is two to three weekends, so depth starts first because it reads slowest. Decision rule: ship the shallowest depth whose profit per session is within 5 percent of the best arm, because ties always go to margin.
Test 2: The Duration Dial
Hypothesis: "a 15-minute window that genuinely closes converts within 10 percent of a 45-minute window." Variants: 15 minutes against 45 minutes, 50/50 split, same depth across both arms. Primary KPI: conversion rate, because the question here is whether urgency compresses the decision or kills it. Readability window: one to two weekends, since window length moves behavior more visibly than depth. Decision rule: take the shortest window whose conversion rate is no more than 10 percent below the longest window's, because every extra minute of window trains "I will decide later" and collides with the queue pressure of the real event.
One condition sits under the whole test: the timer has to genuinely close, or you are measuring theater, not behavior. Some Shopify countdown timer apps restart the clock when the page reloads, so check yours before you trust the result.
These two specs are configuration, not custom development. Growth Suite's A/B Testing Module splits offer traffic across variants by discount depth, duration, and allocation percentage, and lets you pick the scoreboard before launch: conversion rate, average order value, or total revenue. Results come in as the test runs, so a two-weekend read does not require an export and a spreadsheet, and the tests run on live offer traffic, the same walk-away customers the Black Friday campaign will actually address.
Conversion rate is the loudest number in a depth test and the least honest: it will crown the deepest discount every single time.
Test the Reach and the Lever: Allocation and Threshold
Depth and duration decide what the offer costs and says. These two dials decide who ever sees the offer and what it does to the basket. The allocation dial sets how many visitors ever see the offer, and the only way to read it is to withhold the offer from a holdout group and compare. The threshold dial sets what the offer does to basket size: a flat percentage discounts the baskets you already had, while a tiered structure pays only for baskets that grew.
Test 3: The Allocation Dial
Hypothesis: "a meaningful slice of the visitors who take the offer would have bought without it." Setup: 80 percent of eligible visitors see the winning offer from Section 2, and a 20 percent holdout sees nothing. Primary KPI: total revenue per eligible session, computed across both groups together.
Read it like this. Over two weekends, say 2,000 eligible sessions per group, the exposed group buys at 7.2 percent and the holdout at 5.9 percent. A dashboard would book 144 offer wins. The holdout says 118 of those orders were coming anyway, and the offer truly moved 26. Decision rule: if the holdout buys at 80 percent or more of the exposed rate, the Black Friday plan should narrow who qualifies for an offer, not deepen the discount, because dedicated buyers were never the problem and paying them changes nothing except your margin.
Test 4: The Threshold Dial
Hypothesis: "a tiered structure lifts average order value more than a flat 15 percent while spending less on small baskets." Variants: a flat 15 percent against tiers set from your own numbers. For a store at an $85 average order, something like 10 percent over $100, 15 percent over $150, and 20 percent over $200, with the bottom rung just above the median basket and the top rung 40 to 50 percent above average. Primary KPI: gross profit per session, with average order value as the supporting read, and average order value is also the other AOV lever worth building before November. Because tiers run storewide, alternate the two structures in 3 to 4 day blocks across two weeks, Thursday to Sunday against Monday to Wednesday, so weekday and weekend mix stay balanced. Decision rule: ship the structure with the higher profit per session, then check the take-rate. If fewer than 15 percent of discounted orders reach the middle rung, the rungs are set for baskets you do not have.
The winning threshold structure has to become a real campaign on a real date. Growth Suite's scheduled storewide campaigns run spend-based tiers with fixed start and end dates, so the structure that won the October blocks goes live for Black Friday week exactly as it was measured and ends exactly when the event ends, with no one rebuilding it from memory at midnight.
The holdout group is the only line in your dashboard that tells you what the offer actually bought.
Sequence the Four Tests Before the Window Closes
The order is fixed by read speed and money at stake: depth first, duration second, allocation third, threshold across the final blocks. One dial at a time on the same audience, because overlapping tests mix their effects and leave you unable to say which change moved the number. And each test gets read once, on its stop date, with the decision recorded in the same document the Black Friday plan lives in, feeding the four winning settings into the Q3 numbers. All four dials lock by mid-November, because the week of the event is for load, creative, and shipping, not for learning.
The Order Is Depth, Duration, Allocation, Threshold
Working from this week's calendar: depth starts Friday, September 25 and runs through Sunday, October 11, three weekends, because it reads slowest and prices everything else. Duration runs Monday, October 12 through Sunday, October 18 on the winning depth. Allocation runs Monday, October 19 through Sunday, November 1, two weekends, because the holdout read only needs a purchase-rate gap, not a new offer build. Threshold alternates its blocks from November 2 through November 13. All four dials lock by November 16, which leaves the final stretch for load testing, creative, and shipping cutoffs instead of reopened questions. Smaller stores compress the depth test to two arms so the whole sequence still fits.
Read Once, on the Day You Said You Would
Do not peek daily. A Tuesday lead flips by Saturday in offer tests, and every early look is a chance to stop the test the moment your favorite variant is ahead, which is how 20 percent won in the first place. Read each test once, on the stop date, apply the if/then you wrote in September, and record the result and the decision in the plan document, so the version of you standing in mid-November cannot reopen settled questions under pressure.
| Dial | Hypothesis | Variants | Primary KPI | Traffic split | Readability window |
|---|---|---|---|---|---|
| Depth | 15% converts walk-away customers nearly as well as 20% | 10% / 15% / 20% | Gross profit per exposed session | Even thirds, offer-eligible only | ~100 orders per arm, 2-3 weekends |
| Duration | A 15-minute genuine window holds within 10% of 45 | 15 min / 45 min | Conversion rate | 50/50 | 1-2 weekends |
| Allocation | A slice of offer-takers were buying anyway | 80% exposed / 20% holdout | Total revenue per eligible session | 80/20 | 2 weekends |
| Threshold | Tiers lift AOV more than a flat 15% | Flat 15% / 10-15-20 tiers | Gross profit per session + AOV | Alternating 3-4 day blocks | 2 weeks of blocks |
Run the test, read it on the day you said you would, and move the dial. A result you never act on is an opinion with a chart attached.
Four Dials, Four Tests, One Document
A Black Friday offer is four dials, not one decision, and any dial that reaches November untested is a guess billed at full-event volume. Depth and duration set what the offer costs and says. Allocation and threshold set who ever sees it and what it does to the basket. Conversion rate crowns the deepest discount every time, so the only honest scoreboard is gross profit per exposed session, with the holdout group as the auditor. And a test without a pre-written decision rule is a slower guess: hypothesis, variants, one KPI, stop rule, if/then, or do not bother running it.
Tonight, open whatever document holds your Black Friday plan and write down the four settings: depth, duration, exposure, and threshold. Next to each one, name the test that chose it. Every blank space is a dial still set to somebody's memory, and September is the last month cheap enough to change that.
Growth Suite exists for the store that would rather know than remember: the A/B Testing Module runs the depth, duration, and allocation tests on live offer traffic with the scoreboard chosen up front, and scheduled tiered campaigns carry the winning structure into Black Friday week. It installs free from the Shopify App Store, and the 14-day trial is long enough to read your first test. If you are still comparing tools, our breakdown of the top Shopify discount apps covers the main options.
Frequently Asked Questions
How long should I run an A/B test before Black Friday?
Run each test until every arm has roughly 100 orders, or two full weekends, whichever comes first, then stop on the date you set before launch. For a store doing about 1,000 orders a month, most offer tests read inside two to three weeks. What ruins tests is not small samples; it is peeking daily and stopping the moment your favorite variant leads.
Should I test discount depth or duration first?
Depth first. It carries the most margin, so a wrong setting costs the most, and it reads the slowest because conversion differences between depths are small. Duration reads fast because window length changes behavior more visibly. Start depth in the last week of September and it still finishes in time for the duration test to run on the winning depth in mid-October.
What KPI should a discount A/B test use?
Gross profit per exposed session: conversion rate multiplied by average order value multiplied by your margin after the discount. Conversion rate alone almost always crowns the deepest discount because it ignores what the discount cost. A 20 percent offer that converts better can still return less profit per session than 15 percent. Choose the KPI before launch, or the dashboard chooses conversion rate for you.
Can I run more than one A/B test at a time?
Only on audiences that do not overlap. Running a depth test and a threshold test on the same visitors mixes the effects, and you cannot say which change moved the number. With limited traffic, run tests sequentially: one dial at a time, winner locked, next test started. Four sequential short tests still finish before mid-November if the first one starts now.
How much traffic do I need for a Black Friday A/B test?
Less than testing folklore claims. You are choosing between two or three settings, not publishing research. As a working rule, about 100 orders per arm separates a real difference from noise at the effect sizes discounts produce. If your store cannot reach that in two weekends per arm, drop to two arms instead of three and keep the test honest rather than fast.
Ready to Implement These Strategies?
Start applying these insights to your Shopify store with Growth Suite. It takes less than 60 seconds to launch your first campaign.
Muhammed Tüfekyapan
Founder of Growth Suite
Muhammed Tüfekyapan is a growth marketing expert and the founder of Growth Suite, an AI-powered Shopify app trusted by over 300 stores across 40+ countries. With a career in data-driven e-commerce optimization that began in 2012, he has established himself as a leading authority in the field.
In 2015, Muhammed authored the influential book, "Introduction to Growth Hacking," distilling his early insights into actionable strategies for business growth. His hands-on experience includes consulting for over 100 companies across more than 10 sectors, where he consistently helped brands achieve significant improvements in conversion rates and revenue. This deep understanding of the challenges facing Shopify merchants inspired him to found Growth Suite, a solution dedicated to converting hesitant browsers into buyers through personalized, smart offers. Muhammed's work is driven by a passion for empowering entrepreneurs with the data and tools needed to thrive in the competitive world of e-commerce.
More Insights from Our Blog
Continue reading for more expert tips and strategies to grow your Shopify store
60 Days Until Black Friday: What to Lock In This Week and What Can Still Wait
Black Friday is 60 days out. Some decisions get more expensive every week they stay open, and some get smarter. Here is how to tell which is which.
5 Numbers From Your Q3 Data That Should Decide Your Black Friday Offers
Your Black Friday offer is five settings, and Q3 already contains all of them. Read the five numbers that set depth, scope, mechanic, window, and audience.
Your Black Friday AOV Is Decided in September: Build the Upsell Funnel Now
Your Black Friday AOV is set by upsell take rate, and take rate is calibrated in September. The math: trigger rules, product picks, priority.
Explore more
Resources
Shopify Upselling & Cross-Selling
Most Shopify stores leave 10-30% of revenue on the table because they do not upsell or cross-sell effectively. This...
Shopify Conversion Rate: The Complete Resource Hub
Your conversion rate is the most important number in your Shopify store. Learn how to measure it, diagnose problems,...
Shopify Holiday & Seasonal Campaign Strategies
Every holiday has its own shopping psychology. Plan year-round campaigns with the right discount, right timing, and...
Shopify Countdown Timer
Adding a timer takes 5 minutes. Making it convert without destroying trust? That's strategy. The complete guide to...
Shopify Cart Abandonment
70% of Shopify carts are abandoned. Learn why customers leave, how to prevent abandonment in real-time, and recover...
Shopify Discount
Master Shopify discounts without sacrificing your profit margins. Learn strategic discount techniques, timing...