Multi-armed bandit A/B test

Multi-armed bandit vs A/B testing

Every vendor fact below was read from that vendor's own documentation or pricing page on 26 July 2026; the load-bearing ones are quoted inline.

The verdict

Where the traffic goes, over the same fourteen days.

Schematic — endpoints only. No vendor publishes its reallocation curve.

A/B test fixed split

Half your shoppers see the weaker arm on day 14, the same as on day 1.

Multi-armed bandit reallocated on a schedule

The trailing arm thins but never closes — a bandit keeps exploring it.

Same test, same fourteen days. One is buying an answer. The other is buying conversions.

At a glance

A fixed-split A/B test and a multi-armed bandit, at a glance
A/B test (fixed split) Multi-armed bandit
Traffic split Fixed for the whole run Reallocated toward the leader on a schedule
Significance and control Both Optimizely’s implementation: neither. MABs “do not generate statistical significance and do not use a control or baseline experience.” Vendors differ — ask yours what its bandit reports
What you learn about the losers An estimate for every arm, and segments you can slice afterwards Little. VWO on post-test segmentation: “This analysis is possible in A/B Tests but might not be possible in MAB as sufficient data might not be available for underperforming variations”
Metrics it optimises As many as you can read One. VWO: “…they don’t work well for multiple goals as they only factor in the Primary Goal…” AB Tasty: “The Dynamic allocation algorithm is exclusively based on this primary goal”
Best for A decision you’ll defend in six months A short window, many variants, one fast metric

All quotes read 26 Jul 2026 from the vendor’s own documentation.

“Bandits find the winner faster” is backwards

The category’s most repeated claim, contradicted by the best-known vendor selling bandits, in its own support docs (read 26 Jul 2026):

“Because fixed traffic allocations are optimal for reaching statistical significance, MAB-driven experiments generally take longer to find winners and losers than A/B tests.”

VWO says it from the other direction: “A/B tests are still the fastest way to statistical significance even though you might lose some conversions in the process” (read 26 Jul 2026).

A bandit gets you earning sooner. It does not get you knowing sooner — every visitor it diverts from the weak arm is a visitor whose data you no longer have for that arm. If you’re shopping for bandits because your tests take too long, this is the wrong fix.

When a bandit is the wrong choice

Four conditions, every one of them from a vendor that sells bandits.

1. You need a clean read on a single hypothesis. VWO: bandit experiments “are not the best choice when you want to get a statistically robust winner” (26 Jul 2026). If the answer has to survive a conversation — a price change, a shipping threshold, anything finance will ask about — a bandit hands you a decision without the evidence for it.

2. Your traffic is thin. The vendors disagree, so read both. AB Tasty lists low traffic as a reason to use dynamic allocation: “When you have really low traffic on the page you want to test, but still want to do some optimizations.” VWO argues the opposite: “Ideally, run MAB tests on high-traffic pages. Small traffic volumes may prolong the time to reach statistical significance” (both 26 Jul 2026). Take VWO’s side. Give a bandit too little data and it locks onto the arm that got lucky on Tuesday, then sends most of your traffic there.

Neither method manufactures statistical power. Two arms, a 3% purchase rate, a +10% relative lift:

Days to complete a two-arm test at a 3% purchase rate and a +10% relative lift
Sessions per day Days to complete
1,000 107
3,000 36
10,000 11

Computed 26 Jul 2026 — two-proportion z-test, unpooled variance, α = .05 two-sided, 80% power: 53,208 sessions per arm, 106,416 total. Not a vendor figure; reproduce it against any published calculator.

At a thousand sessions a day you can’t resolve a 10% lift on purchase rate inside a quarter. That’s arithmetic, not tooling. Test a bigger change, or test higher in the funnel — the same target on a 12% add-to-cart baseline needs 12,001 sessions per arm, not 53,208.

3. Your metric is slow to mature. AB Tasty recommends dynamic allocation “When you want to optimize micro conversions that are expected to occur within a short period after the user has been exposed to a variation” (26 Jul 2026). That condition does a lot of work. If the first visit is Monday and the order lands on Sunday, an hourly-updating algorithm spends the week optimising against a signal that hasn’t happened yet.

4. Your traffic mix changes mid-test. AB Tasty, explicitly on when not to use it: “We recommend not using dynamic allocation in case you want to optimize a website/application where the visitors will be different over time” (26 Jul 2026). For a DTC store that’s the 40,000-person email on day three, the creator post, the cold Meta campaign switching on. The bandit reads the swing as evidence about your variants when it’s evidence about who showed up. A fixed split absorbs that. A bandit banks it.

And the operational cost nobody puts in the pitch. AB Tasty publishes what you give up the moment you launch: you cannot “Change the primary goal,” cannot “Add, remove or duplicate a variation within the test,” and cannot “Switch back to static allocation.” Their only remedy — “you need to pause the test and duplicate it.” Optimizely locks the same door: “After you start your MAB, you cannot change the primary metric,” and “do not change variations mid-experiment” (both 26 Jul 2026). Ask any vendor you shortlist whether theirs works the same way.

When a bandit is the right choice

A promotion with a hard end date. A ten-day sale, a Black Friday hero. By the time an A/B test concludes, the thing you were testing is over. Take the conversions.

More than six variants and no hypothesis. AB Tasty: “When you have a lot of variations to test (more than six), dynamic allocation enables you to quickly identify the lowest-performing variations” (26 Jul 2026). Six arms need three times the total sample of a two-arm test at the same per-arm precision, and hand you five comparisons when all you wanted was the best hero image.

An always-on slot with no end state. The recommendation strip, the merchandising block, the order of a category page. Those aren’t decisions, they’re settings.

A metric that fires inside the session. Add-to-cart rate, email capture, quiz completion.

What those four share: the cost of being wrong is low, and so is the value of knowing.

Which tools have one, and on which plan

Every tool here ships a bandit. What differs is the plan it sits on — the row rival comparison pages get wrong most often.

Which testing tools include a multi-armed bandit, and on which plan
Tool Bandit sits on
Optimizely Web Experimentation, Personalization, Feature Experimentation, Full Stack (Legacy). The support article names no tier. No price published — every plan is “individually packaged”
VWO (rebranding to Wingify) Web bandit on all three tiers, including Growth. Server-side and personalisation bandits start at Pro. No price published
AB Tasty Called “dynamic allocation.” No plan matrix published, and no free trial — a sales-run proof of concept instead
Convert Pro and Enterprise. Not on Growth. Pro is $599/mo billed monthly ($420 annual) at the 100K tested-users rung
GrowthBook Pro ($40/seat/mo) and Enterprise. Not on the free Starter tier, and not in the free self-hosted build
Statsig Pro ($150/mo) and Enterprise. Not on the free Developer tier
Mtrix Included, with no cap on how many experiments run. No plan matrix published

Every cell read from that vendor’s own pricing page or docs, 26 Jul 2026.

Planning to run one for free? GrowthBook and Statsig have genuine free tiers and both exclude bandits on their own pricing pages. Convert has no free tier at all — a 15-day trial, no credit card. Optimizely’s free Feature Experimentation plan is “Totally free, no credit card needed” but runs “one concurrent experiment at a time” and never mentions bandits, so reachability there is simply unpublished. Mtrix has no free tier either: one month free, no credit card, then paid — flat across three tiers, with no meter on sessions, orders or profiles.

On Shopify and haven’t bought anything yet? Since 5 June 2026, Shopify Rollouts can A/B test whole themes and whole checkout and customer-account configurations, included in the plan fee, with experiments gated to Grow or higher (changelog 5 Jun 2026; help centre, read 26 Jul 2026). It swaps one whole configuration for another on a fixed split — not element-level testing, not audience targeting. Exactly where it stops is mapped in what server-side testing is and every checkout-testing option. Start there. When it stops answering your questions, book a demo and bring the test you couldn’t run.

Where Mtrix fits

Mtrix runs bandit allocation that favours winners while still exploring, with no cap on the number of experiments. As the table shows, the bandit itself is table stakes. The differences are in what surrounds it.

Many tests at once, read two ways. Mtrix runs several tests together and reads the result either as the winning combination or as each test on its own — the headline test, the hero-image test and the badge test go live in the same week, and you can still ask which headline won. See how it reads both. The nearest equivalent elsewhere is multivariate testing, which is usually tiered: VWO’s own matrix puts multivariate on Pro and Enterprise, Convert’s on Pro and above (both 26 Jul 2026).

Overlapping tests that don’t collide. Audience targeting buckets on the server and makes overlapping experiments mutually exclusive as standard — what makes an uncapped test count safe rather than dangerous. VWO’s pricing matrix puts mutually exclusive groups, up to ten user groups, on Enterprise only. Shopify includes it in the plan fee — “Each experiment gets its own mutually exclusive segment of visitor traffic” — but scoped to one resource at a time, a theme or a checkout configuration (help centre, 26 Jul 2026), where Mtrix’s covers any experiment on the storefront. And “no cap on tests” is a licensing fact, not a statistical one: run more tests than your traffic supports and you get confident-looking noise, whichever allocation method you picked.

Assignment that sticks. The standard complaint about bandits is that a returning shopper sees a different arm than yesterday. Mtrix decides the arm in its backend and stores it against your visitor ID — a 30-day server-side lookback, not a coin flip in the browser. What that changes, and what it costs.

Ending on your terms. Experiments auto-complete on a visitor count, a date, or a revenue threshold, and report probability-to-be-best on revenue per visitor, conversion rate and add-to-cart rate. That’s a ranking, not a confidence interval: which arm is most likely ahead, not how much better. You can change the traffic split mid-flight, add a new variant after launch, and change the allocation at any time — the statistical analysis reshapes accordingly. And the reallocation runs both ways: if the results call for it, the bandit can move back toward an even split. In practice it rarely does, but nothing locks it.

When an arm loses, the sessions are still there. Filter session replay by experiment and by variant arm and watch the visits behind the number — every session, captured at 100%, not sampled. A bandit costs you the estimate for the weak arm. It doesn’t have to cost you the reason.

What vendor pages leave out

Mtrix doesn’t publish its allocation algorithm. Optimizely publishes its down to the procedure: Thompson Sampling for binary metrics, Epsilon Greedy for numeric ones, and “The MAB model is updated hourly” (26 Jul 2026). AB Tasty publishes neither — its dynamic-allocation doc names no algorithm and no interval, only that the allocation changes “periodically… following a statistical formula” (26 Jul 2026). If your team will ask which algorithm is running and how often it reweights, Optimizely’s transparency is a real reason to pick them. What Mtrix does state is the behaviour, not the procedure: continuous reallocation, shaped by the results as they accumulate — no named algorithm, no published cadence.

Mtrix is web only. No native iOS or Android SDK. Optimizely, VWO and AB Tasty all ship mobile SDKs. If you’re testing inside an app as well as a storefront, that’s decisive, and Mtrix has no answer to it.

If you already have a data warehouse, look at GrowthBook. It computes results where your data already lives, its free tier carries unlimited experiments and unlimited traffic, and its core is open source under MIT — bandits sit in the paid tiers. The premise is that you have a data function. Most DTC teams of one to twenty don’t.

FAQ

When should you use a multi-armed bandit instead of an A/B test?
When the window closes before an A/B test could finish, when you have more than about six variants, when the metric fires inside the session, and when you'll never have to explain the answer. Keep the fixed-split A/B test when the decision is permanent, when someone will ask you to justify it, when you care about more than one metric, or when you'll segment the result afterwards.
Is multi-armed bandit testing faster than A/B testing?
Not at finding a winner. Optimizely's own documentation: "MAB-driven experiments generally take longer to find winners and losers than A/B tests" (26 Jul 2026). A bandit is faster at earning, not at knowing. And it doesn't repeal the arithmetic: at 1,000 sessions a day, a 10% lift on a 3% purchase rate takes about 107 days to resolve either way.
Does a multi-armed bandit give you statistical significance?
Not on Optimizely's implementation: MABs "do not generate statistical significance and do not use a control or baseline experience" (26 Jul 2026). Vendors that report a Bayesian probability-to-be-best give you a ranking instead — which arm is most likely ahead, not how much better. If you need to state the size of the win, run a fixed-split A/B test.
Which testing tools include a multi-armed bandit?
All the major ones, rarely on the cheapest plan they sell. Convert puts it on Pro and above, GrowthBook on Pro ($40/seat/mo), Statsig on Pro ($150/mo) — and neither GrowthBook's nor Statsig's free tier includes it. Convert has no free tier at all, only a 15-day trial. VWO's web bandit is on all three tiers. Optimizely and AB Tasty publish no plan matrix. Mtrix includes it with no cap on how many experiments run (all read 26 Jul 2026).
Ada Kern Experimentation & analytics Writes about experimentation and analytics at Mtrix, and about what a test result does and does not justify.
Build with Mtrix