NavonaAINavonaAI
A/B Testing

How Long a Test Takes

Traffic decides how long an A/B test runs — and whether it can finish at all. What to expect at your store's volume.

A/B Testing is coming soon. This page describes how experiments will work once the feature launches. Want early access when it's ready? Let us know at support@navona.ai.

The honest answer is that it depends almost entirely on how many shoppers are shown a prevention popup. At high traffic a test can settle inside a week or two. At low traffic the same test can take a year, and for some stores it will never settle at all.

This page gives you the real shape of that so you can decide whether a test is worth starting.

What actually drives the wait

A test finishes when the gap between the two versions is big enough that we can tell it apart from random luck — which shoppers happened to land in which group. Two things decide when that happens:

  • How many interventions you show. More shoppers, faster answer. This is the dominant factor by a wide margin.
  • How big the real difference is. A change that improves things by 20% is spotted far sooner than one that improves things by 2%.

Neither is something NavonaAI can speed up. A test cannot be made to finish faster by wanting it to.

What counts is interventions shown, not visitors and not carts. An intervention is one prevention popup shown to one shopper — the same thing your dashboard's Interventions Shown card counts. A shopper who never triggers a popup contributes nothing to the test, because both versions treated them identically. See Why we measure per intervention.

How much data each metric needs

These are measured figures from four real NavonaAI stores, for detecting a 10% relative improvement — for example a prevention rate of 10.0% rising to 11.0%, not rising by 10 percentage points.

MetricInterventions shown needed, per group
Accept rate~900 – 1,200
Prevention rate (decides the winner)~1,400 – 3,300
Net revenue per intervention (guardrail)~4,400 – 10,400

Each figure is a range because the requirement depends on your own baseline rates. Your store sits somewhere inside it.

Revenue needs by far the most data because order values vary so much. At every one of those four stores, the spread in order value was between 3 and 17 times the average order value. A metric that jumps around that much takes a lot of orders to pin down — which is exactly why net revenue acts as a guardrail rather than as the verdict.

What that means in calendar time

Divide the requirement above by how many interventions you show per day, per group. That is the whole calculation, and you can check it yourself.

Worked through at four real store volumes:

Interventions shown per day, per groupAccept ratePrevention rate (the verdict)Net revenue per intervention
~550~2 daysunder a week~1 – 3 weeks
~28~4 – 6 weeks~2 – 4 months~5 – 12 months
~20~6 – 8 weeks~2.5 – 5.5 months~7 months – 1.5 years
~5~6 – 8 months~10 months – 2 years~2.5 – 6 years

The four daily volumes are real measured store traffic. The durations are arithmetic on the table above, so they inherit its ranges — treat them as the shape of the answer, not a promise.

Read the bottom row honestly. At around five interventions a day per group, A/B testing is not a tool your store can use — a verdict on the primary metric would take up to two years, by which time the season, the catalogue and the traffic mix have all changed. If that is your store, you are better served by Best Practices and the Strategies section than by a test that will never resolve.

There is no shame in this. Statistical testing has a minimum traffic requirement the same way a scale has a minimum weight, and most Shopify stores sit below it for most changes worth making.

Your own store's number

Before you start a test, NavonaAI shows the estimate for your store, based on your recent intervention volume — how many days until each metric could produce a verdict. Check it before you commit. If the estimate for the primary metric is longer than you are willing to wait, do not start the test: an abandoned half-finished test tells you nothing, and a half-finished test you act on is worse than no test at all.

Why the finish date moves

The estimate is recalculated as data arrives, and it will move in both directions.

If the early data shows a large gap, the estimate shortens. If that gap narrows as more shoppers come in — which is common, because early gaps are the noisiest — the estimate gets longer. A shrinking difference genuinely does take more data to confirm.

This is normal. An estimate that only ever counted down would be lying to you.

Run for at least one full week, whatever the estimate says

Even a high-traffic store should let a test run through at least one complete week, and preferably two.

Shopper behaviour differs by day of the week. A test that starts Thursday and ends Monday has weighted your weekend disproportionately. Two full weeks captures two complete weekly cycles and stops the calendar from picking your winner for you.

On this page