Metrics & Results
The metric that decides an A/B test, the formula behind it, and how to read the verdict.
An A/B test compares Group A and Group B on one primary metric — the goal you pick when you set up the experiment. Everything else on the results page exists to explain why that metric moved.
The primary metric
You choose the primary goal from exactly two options:
| KPI | What it measures |
|---|---|
| Cart to Order Rate | Orders ÷ carts that reached the funnel — the default |
| Total Abandonment Rate | The complement of Cart to Order Rate |
There is one dropdown for this, populated from these two values — no separate "simple" or "advanced" mode, and no user-adjustable minimum detectable effect. Both KPIs are scored the same way: a pooled two-proportion z-test on orders over assigned carts.
Before you start a test, NavonaAI shows an estimate of how long it will take to reach a verdict, based on your store's recent traffic. That estimate uses a fixed internal assumption (a 10% relative effect size) — it is not something you configure per test. See How Long a Test Takes.
The formula and its denominator
Cart to Order Rate = Orders / Sample Size (Carts)Sample Size (Carts) — also called assigned valid carts — is every cart that was assigned to that variant and produced at least one valid tracked event. A cart is assigned the first time NavonaAI loads for it, which happens before any popup trigger condition (idle timeout, exit intent, and so on) is evaluated. Assignment does not require that a popup ever actually showed.
Orders only count if they belong to an assigned valid cart and fall inside the experiment's window (from starts_at, with a 72-hour grace period after ends_at for a cart that was already abandoning when the test ended).
Why some assigned carts never see a popup
Most carts that load NavonaAI are assigned to a group immediately — but only some of them go on to trigger a popup at all. A shopper whose cart is assigned and who then completes checkout without ever seeing a trigger condition had an identical experience in both groups: nothing about the test could have changed their outcome.
Those shoppers still count in the denominator, because assignment — not the popup firing — is what puts a cart in Group A or Group B. Including them is not a bug: excluding them would require tracking "was this shopper ever eligible to see the offer," which the product does not do. But it does mean the measured difference between groups is diluted by however many assigned carts never had a chance to be affected, which is exactly why a real effect can take longer to show up than a naive "orders per cart" intuition would suggest.
This is a real property of the measurement, not a numbers claim about any specific store — how much dilution you see depends on your own popup trigger rate. See How Long a Test Takes for what that means for how long a test needs to run.
Reading the verdict
The results page shows one of four states, computed entirely on the backend from the z-test and two sample-size floors — 200 assigned valid carts per arm and 10 orders (conversions) per arm. Nothing is recomputed in the browser.
| State | Meaning |
|---|---|
| Significant | The gap cleared the two-proportion z-test at 95% confidence, and both sample-size floors are met |
| Collecting data | Not enough assigned carts yet (badge shows roughly how many more per group are needed), or carts are sufficient but conversions aren't yet — no count is shown for that case, since the product doesn't compute one |
| Not significant | Both floors are met, but the gap is smaller than this traffic can reliably tell apart from noise — this does not mean the two versions perform identically, only that this test can't distinguish them yet |
| Can't be scored | The experiment's primary goal isn't one the verdict computation measures — terminal, more traffic will never fix it |
"Not significant" is not "no difference." A result below the detectable threshold could be a real small effect, a real effect this traffic is too small to see, or genuinely nothing — the badge cannot tell you which. See Reading Your Results for how to act on it responsibly.
Alongside the badge, the results page also reports:
- p-value — how often a gap this size would appear by chance alone if the two groups truly performed the same
- 95% confidence interval on the gap between B and A, in percentage points
- Detectable effect — the smallest true difference this experiment could reliably catch right now, given its current sample size
There is exactly one winner determination per experiment, driven by the primary goal you picked at setup. Nothing else on the page can override or veto it — a store admin can separately and manually declare a winner regardless of the statistical verdict, which is an explicit override action, not an automatic second opinion.
Supporting metrics
Each group also reports the figures that explain why it's ahead or behind on the primary metric:
| Metric | What it is |
|---|---|
| Sample Size (Carts) | Assigned valid carts for this group — the denominator above |
| Orders | Assigned valid carts that went on to place an order |
| Average Order Value | Revenue ÷ Orders |
| Revenue | Total order value from this group's converting carts |
| Discount Cost | Value of NavonaAI discount codes redeemed by this group — percentage discounts are approximated against the order's post-discount total, and free-shipping discounts count as $0, since neither the pre-discount subtotal nor shipping cost is tracked |
| Net Revenue | Revenue minus Discount Cost |
| Cart Abandonment Rate | 100 − Cart to Order Rate, shown lower-is-better — the same figure as the Total Abandonment Rate KPI above, labeled differently in this table |
Net Revenue is revenue, not profit — it subtracts the discount you gave away but not your cost of goods, so it does not by itself decide a winner. It is reported next to the primary metric so you can sanity-check that a group converting more shoppers isn't doing so by handing out proportionally larger discounts.
Daily performance
Below the summary, two views break results down over time:
- Daily Performance — Orders and conversion rate for both groups, bucketed by the day each cart was first assigned (not the day it later ordered), paginated
- Gap over time — the cumulative Group B minus Group A gap plotted against a shrinking "if the groups were identical" band, so you can see a real effect separate from early noise settling down
Next steps
- Reading Your Results — what the verdict does and does not tell you
- How Long a Test Takes — traffic, time, and feasibility
- Glossary — plain-language definitions