Pick a Winner, or Never Stop Optimizing: Two Ways an A/B Test Can Decide

Every A/B testing tool asks the same question: which version won? But there are two different things you might mean by that, and they want two different kinds of test. Sometimes you need an answer — a redesign either beats the old page or it does not, and you will ship one of them and move on. Sometimes you need a manager — the hero banner will run all year, the best version in January is not the best version in July, and nobody is going to come back and re-test it every quarter. Personyze runs both kinds from the same campaign, and the setting that chooses between them is called How this test decides. This post is about what each mode actually does, what it costs you, and how to choose.
The same test, decided two ways. On the left the split holds until one version proves itself, then the winner takes the tested traffic. On the right the split moves a little every night, and never stops.
Two questions, two kinds of test
A classic A/B test is an experiment. You fix the split, you wait, you get a verdict with a confidence attached, and the test is over. That fairness is the whole point: because the split never moved, the comparison is clean, and you can show anyone why you made the call. The price is that the losing version keeps receiving its full share of traffic until the end, and that the verdict is only true of the period it ran in.
The other kind of test is what statisticians call a bandit: instead of waiting to decide, it keeps re-deciding. Each night it looks at what has happened recently and gives a little more traffic to whichever version is performing now, while always holding some back to keep learning. It never produces a verdict, because it never needs one. The price is that the split is always trailing the truth by a few days, and the comparison is not clean enough to publish as a finding. That is fine when the goal was never a finding.
Pick a winner at significance: the experiment
In this mode the traffic stays as you set it while the versions compete on your goals, and the test concludes when one of them proves itself. What “proves itself” means is four thresholds, and Personyze ships them stricter than most tools do.
- Goals, one or more. The first goal you pick is the primary; readiness and the winner call track it. A test that optimizes for clicks and a test that optimizes for revenue do not always agree, so you say which one you meant.
- Confidence, sessions, runtime. Three presets set all three at once: Cautious at 99% confidence, 5,000 sessions and 21 days; Balanced at 95%, 1,000 sessions and 14 days; Fast at 90%, 500 sessions and 7 days. Or type your own. The runtime floor is what stops a good Tuesday from winning on Tuesday.
- A minimum improvement. Cautious asks for at least 1% lift, Balanced for 0.5%. A version that is ahead by a statistically real but commercially pointless margin does not get promoted.
- A holdout, if you want one. Give the untouched site its own share — two versions at 45% and 10% held out — and the lift you report is the lift you caused, not the lift against a guess.
When the bar is cleared, the winner is promoted: all of the tested traffic, or the share you nominate, or none at all if you would rather be emailed and decide yourself. A control group you kept stays held back, so the campaign can still be measured after the call. Then it is over, which is exactly what you wanted from it.
AI: optimize continuously: the manager
In this mode nothing concludes. Every night a pass scores every version on what has happened recently and re-balances the split toward the one performing now. Three properties make it safe to leave running.
- Recent traffic counts most. Views and clicks are decayed with a 28-day half-life, and anything older than 84 days is dropped. A version that was best two months ago can still lose today, which is the point: it notices when a season turns.
- It moves only with statistical confidence. Each version gets a chance it is best, from thousands of simulated re-runs of the data. Near-even numbers mean the data cannot tell the versions apart yet, and then the split holds still. It also refuses to compute anything at all under a few hundred decayed views, because a bandit on a handful of sessions is a random-number generator with a straight face.
- 10% always exploring. A tenth of the traffic is split evenly among the versions every night, no matter what. A losing version keeps earning the data that could redeem it, no version is ever driven to zero, and a data blackout cannot freeze one version at 100%.
What the nightly pass does, in order. The last step is the only one that touches serving, and it only writes when the recommended split differs from the current one by three points or more.
Two more rules are about respecting what you set. A version you put at 0% is your decision, not a data point, and it stays at 0%; giving it a share by hand is what re-enters it. And the optimizer only re-allocates the traffic the split already assigns — a campaign holding 10% back for the untouched site keeps holding it back.
One honest limit: the continuous mode scores interaction — clicks-or-better over confirmed views, where a click that went on to convert still counts — because that is the signal always-on content produces in volume, every day. If the thing you care about is revenue per session on a checkout change, that is a question for the experiment, pointed at that goal.
Which one should you choose?
| Pick a winner at significance | AI: optimize continuously | |
|---|---|---|
| The question | Which version is better, and by how much? | Which version should get the traffic tonight? |
| How it ends | With a verdict. The winner is promoted and the test concludes. | It does not. The split keeps following the recent data. |
| Traffic while it runs | Stays exactly as you set it, so the comparison is fair. | Shifts nightly toward the leader; 10% always divided evenly. |
| What it optimizes | Any goal you pick: purchases, revenue per session, sign-ups, clicks, bounce rate, time on site. | Interaction: clicks-or-better over confirmed views, so a click that led to a purchase counts too. |
| Good for | A decision you will ship once: redesign, pricing layout, checkout step, page against page. | Always-on content whose best version drifts: hero banners, offers, seasonal blocks, promoted content. |
| Cost of being wrong | The loser keeps its full share until the call, then you can misread a stale result later. | A short lag: the split trails the truth by days, never weeks. |
| What you can defend | Confidence, sessions, runtime, minimum improvement: a call you can show anyone. | A chance-of-best per version, refreshed nightly, and the split that follows from it. |
A rule of thumb that has held up: if you would be embarrassed to still be testing it next quarter, pick a winner. If you would be embarrassed to have stopped, optimize continuously. A pricing page belongs in the first group. A homepage hero belongs in the second. And you can change your mind: the mode is a setting on the campaign, applied when you save, and the same campaign can run as an experiment first and be handed to the optimizer once it has become furniture.
There is a hybrid worth knowing about. The experiment mode has an option to gradually shift the rotation toward the likely winner while the test runs. It gets to a decision faster, at the cost of some of the rigour that made the experiment worth running. Use it for time-sensitive tests, not for the ones you will be asked to defend.
The third reading: per audience
Both modes share a reading that most tools do not have at all. Personyze’s Audience Discovery mines your visitor data for the groups that convert far above or below average, and the test is scored inside each of those audiences separately: the rate per version, which version leads there, and how sure that is.
The per-audience verdicts under a running test. A version can win the site overall and lose the visitors who arrive direct, and the rows say so, with how sure each call is.
This changes what a “winner” is. B won overall, but A won for phone visitors at 86% sure. In a tool that only reports the total, you ship B and silently make the phone experience worse. Here the report names the disagreement, and the campaign editor shows it on every test, whichever mode it decides in.
What to do about it is deliberately simple today: the campaign still serves one split to everyone, so to act on a per-audience winner you target a copy of the campaign at that audience with its winning version. Discovered audiences are a category in the campaign rule picker, so that is a few clicks. Serving a different split per audience automatically is on the roadmap, and the per-audience scoring is what will drive it.
Setting it up
- On the campaign’s Content step, add your versions as rotation groups and set the split. Leave Site original a share if you want a holdout.
- Choose how visitors are assigned. Rotate users keeps a visitor on the same version every visit and is the right default; Visits re-assigns each visit, for banners and headlines where the impression is what is measured.
- Under Goals, pick one or more; the first is the primary.
- Under How this test decides, choose Pick a winner at significance and a preset, or AI: optimize continuously. Save.
- Read the result on the Performance step. In the continuous mode the table shows views, clicked, chance it is best, the current split and the recommended one; below it, which version wins per audience.
- A/B and multivariate testing — the feature page, with the editor shown beat by beat.
- Audience Discovery — where the per-audience reading comes from.
- How a personalization engine decides — where testing sits among the other layers.
Book a demo with a personalization expert
30 minutes with a personalization expert. Bring your stack, your goals, your skepticism. We'll show you what changes when every visit feels like the only one.
