NavonaAINavonaAI
A/B Testing

Best Practices

Tips for running effective A/B testing experiments

A/B Testing is coming soon. This page describes how experiments will work once the feature launches. Want early access when it's ready? Let us know at support@navona.ai.

Run for at Least 7–14 Days

Shopper behavior varies significantly by day of week. Running for less than a week can skew results based on which days happen to fall in your test window. A two-week minimum captures at least two full weekly cycles.

Check the Traffic Estimate Before You Start

This is the step that decides whether the whole exercise is worth it.

What counts is interventions shown per group, not carts and not visitors. On the stores we measured, deciding a winner on prevention rate took somewhere between 1,400 and 3,300 interventions shown per group. At a few hundred a day that is a fortnight; at five a day it is years.

Look at your store's own estimate before starting a test, and do not start one whose estimate is longer than you are actually willing to wait. A half-finished test tells you nothing — and a half-finished test you act on is worse than never testing at all. How Long a Test Takes has the real numbers.

Low-traffic stores are better served by the volume-based termination rule — set a target sample size instead of planning to stop on a date, so the experiment runs until the data is there rather than ending on an arbitrary day.

The stop rule counts something different from what the statistics need. The volume target counts carts assigned, across both groups combined. The figures above are interventions shown, per group — a much smaller number, because most assigned carts never trigger a popup and because it is per group rather than pooled. Set the volume target well above your intervention requirement, or the experiment will stop before it can decide anything.

Test One Variable at a Time

Changing multiple settings between control and treatment makes it impossible to know which change caused the outcome. Keep all other settings identical between your two variants.

What varies between variants is the offer and its rules — see What You Can Test. The popup's wording and design are not variant settings, so an experiment cannot compare two different headlines.

Good: Control at 10% discount, treatment at 15% discount (same everything else)

Risky: Control at 10% percentage, treatment at €5 fixed amount with a minimum purchase (three changes at once)

Don't End Experiments Early

It's tempting to stop when one variant looks like it's winning after a few days, but early results can be misleading. A small number of high-value orders in one variant can make it look dominant when the real long-term performance is different. Let the experiment run to its intended duration or volume target.

Start with a Holdout Test

The most valuable first experiment: control with AI prevention off vs. treatment with AI prevention on. This gives you a clean baseline measurement of how much lift NavonaAI generates for your store — and the data to prove it.

Keep a Record

After promoting a winner, note what you tested and the results. This helps you build institutional knowledge about what works for your customers and informs future experiments.

On this page