← Back to microservices patterns map
Microservices Pattern

A/B Testing Rollout

Show two versions to two groups of users and measure which one works better.

experiment
Lesson

A/B testing: which version is better?

Not a release-safety pattern but a product experiment: show two versions to two groups of users and measure which is better. This lesson explains how to split users correctly and why small numbers lie - then you run a real significance test.

01

The idea in short

Two versions, two groups of users, one business question.

An A/B test shows version A to one group of users and version B to another, at the same time, and compares a business result: sign-ups, purchases, clicks, return visits. It is not about safety - both versions are already safe. It asks: which version is BETTER?

Remember it as: split users, measure one goal, wait for enough data.

At a glance
QuestionIs B better than A for the business?
SplitBy user - each user always sees the same version.
MeasuredConversions, revenue, retention - not errors.
LengthDays or weeks, decided in advance.
NeedsEnough users and a statistics check.
02

An everyday picture: two shop windows

Two displays on similar streets, and a count of who walks in.

A shop with two similar branches tries a new window display in one and keeps the old display in the other. For a month it counts how many passers-by come in. Only if the new display brings in clearly more people - not just a few on one sunny Saturday - does it change every window.

03

Words you need

Six words used in this lesson.

These words come from experiment tools and statistics.

Small dictionary
VariantOne version in the test: A (control, the current one) or B (the new one).
ConversionThe goal action, such as buying. Conversion rate = buyers / users.
AssignmentDeciding which variant a user gets - stable, by user id.
Feature flagA switch in the code that shows A or B without a new deployment.
p-valueThe chance of seeing a difference this big if there were really no difference.
SignificantUsually p below 0.05: the difference is very unlikely to be luck.
04

How it works, step by step

Use Next to walk through real runs.

Step through the diagram. The true conversion rates are fixed in the lab - A 10%, B 12% - and the same experiment is run with more and more users.

request flowAn A/B test: is B really better?step 1 / 4

1 - Split users, not requests

The router hashes each user id, so the same user always sees the same version. user-42 asked three times and got v1 three times.

split
50 / 50 by user id
same user, 3 visits
same version
measure
who buys
not
errors or speed

Real runs: the router splits users 50/50 by user id. Version B converts 12% of users, A 10%.

The steps
1. PlanOne goal metric, the smallest effect that matters, and the number of users needed.
2. AssignSplit users randomly but stably - usually by hashing the user id.
3. RunShow each user their variant; log variant and goal events.
4. WaitUntil the planned number of users - do not stop early.
5. DecideCheck significance; keep the winner, or keep A if there is no clear winner.
05

Why teams use it - and what it costs

Decisions based on evidence, not opinions.

Teams argue about designs, prices and texts. An A/B test settles it with real behaviour. It also protects you from changes that feel better but perform worse.

The costs: it needs a lot of users (often thousands per group), weeks of patience, careful logging, and some statistics. A test that is too small does not just fail to answer - it gives confident wrong answers.

Benefits and costs
BenefitReal user behaviour decides.
BenefitMeasures what matters to the business.
BenefitCan test many ideas over time.
CostNeeds many users and time.
CostEasy to fool yourself with small samples or early stopping.
CostTwo variants to maintain until the decision.
06

Detail 1: split users, not requests

Each user must see one version, every time.

If each request chose A or B at random, a user would see the new checkout on one page and the old one on the next - confusing, and it spoils the measurement: whom do you credit with the purchase? Assign users, not requests, by hashing a stable user id. In the lab, user-42 visited three times and saw version A every time.

Careful with the hash. Code that used parseInt(userId) % 2 looks fine with numeric ids, but with ids like user-1 it put all 1,000 test users into the same group: parseInt("user-1") is not a number. Hash the id with a real hash function instead (shown in "In the real world" below). Adding the experiment name to the hash gives every experiment a different split, so the same users are not always the test group.

07

Detail 2: small samples lie

The lab ran the same experiment at three sizes.

The true rates never changed: A 10%, B 12%. With 200 users, B looked more than 40% better (14.4% against 10.1%) - but the p-value was 0.361: a difference like that happens by luck very often with so few users. With 2,000 users: still not enough (p 0.405). With 20,000 users: clear (p below 0.001).

Before you start, calculate how many users you need. The standard formula says: to detect a change from 10% to 12% with 95% confidence and 80% power, you need about 3,838 users in EACH group. Then run until you have them.

Same experiment, three sizes - measured
200 usersA 10.1%, B 14.4%. p = 0.361 - could be luck.
2,000 usersA 10.2%, B 11.4%. p = 0.405 - could be luck.
20,000 usersA 10.2%, B 11.8%. p < 0.001 - B is really better.
A/A, 2,000 usersSame version twice: 10.2% vs 9.7%, p = 0.687 - correctly no difference.

Watch out: Do not stop the test the first time the numbers look good. Checking often and stopping at the first "significant" moment finds false winners far more often than 5% of the time. Fix the sample size first.

08

Detail 3: plan it like an experiment

Write the plan down before you start.

Choose ONE main metric and the expected effect. Calculate the sample size. Run an A/A test once to check that your setup does not create differences by itself - in the lab it correctly found none. Run for whole weeks, because people behave differently on weekends. Watch a few guard metrics - errors, page speed, refunds - so a "winner" is not secretly harming something else.

09

In the real world

Feature flags and experiment platforms.

Most A/B tests do not deploy two services: they use a feature flag in one deployment, and the code shows A or B per user. Experiment platforms such as LaunchDarkly, Statsig and Optimizely handle assignment, logging and the statistics. The routing function below is the part of such a system that assigns users; it was tested here: 1,000 ids like user-1 split 505 / 495, and user-42 got the same variant every time.

Assign a user to A or B - hash the id, never parseInt
import { createHash } from "node:crypto"; // Fixed: hash the user id (plus the experiment name) into a bucket 0-99 function getVariant(userId, experiment = "checkout-test") { const bucket = createHash("md5").update(experiment + ":" + userId).digest().readUInt32BE(0) % 100; return bucket < 50 ? "A" : "B"; }
Using it in an Express app (example)
app.get("/checkout", (req, res) => { const variant = getVariant(req.user.id, "checkout-test"); analytics.track(req.user.id, "checkout_viewed", { variant }); // log it! res.render(variant === "A" ? "checkout-old" : "checkout-new"); });
10

When to use it, and when not

For product questions with enough users.

An A/B test needs a clear question and enough traffic.

Decide
Good fitCheckout and pricing pages, sign-up flows, texts and layouts.
Good fitRecommendation or ranking changes judged by user behaviour.
Poor fitLow traffic - you may never reach the needed sample.
Poor fitChecking whether a release is stable - use a canary.
Poor fitChanges you must ship anyway, for legal or security reasons.
11

Common mistakes

And how to avoid each one.

Most bad A/B decisions come from statistics, not code.

Mistake -> fix
Deciding with a small sampleCalculate the size first (lab: 200 users looked like +40%, p 0.361).
Stopping at the first good resultFix the length and wait.
Splitting by requestSplit by a stable user id.
parseInt(userId) % 2Use a real hash (old code put 1,000 of 1,000 users in one group).
Many metrics, no main oneChoose one goal before you start.
12

Compared with the other patterns

A/B asks "better?"; the others ask "safe?".

A/B testing is often confused with canary releases.

A/B and its neighbours
A/B testIs it better? Business metrics, days to weeks, ends with a decision.
CanaryIs it safe? Errors and latency, minutes to hours, ends at 100%.
ShadowIs it correct? Real traffic, no user sees it.
TogetherCanary the build to prove it is safe, then A/B test the feature inside it.
13

Interview questions

Short answers you can give in your own words.

What is A/B testing? Randomly but stably assigning users to two versions at the same time and comparing a predefined business metric, using statistics to decide whether the difference is real.

How is it different from a canary release? A canary checks operational safety with a small share of traffic and is promoted to everyone; an A/B test measures user behaviour over a fixed period to choose between versions.

Why does sample size matter? Small samples produce large random differences; without enough users per group you will often pick a false winner. Calculate it in advance from the baseline rate and the smallest effect worth detecting.

Remember

  • Show A to one group and B to another at the same time; measure one business goal.
  • Split by user, not by request - with a real hash, never parseInt(userId) % 2.
  • Calculate the number of users first (10% -> 12% needs ~3,838 per group).
  • Small samples lie: 200 users showed B +40%, but p was 0.361.
  • Do not stop early; run an A/A test once to check your setup.
  • A/B asks "is it better?"; canary asks "is it safe?".

Check yourself

Answer in your head first, then open the answer.

What does an A/B test measure, unlike a canary?

Business results - conversions, revenue, retention - not errors and latency.

Why must the split be by user and not by request?

So each user sees one version consistently, and results can be credited correctly.

What was wrong with parseInt(userId) % 2?

With ids like user-1, parseInt returns NaN, so every user landed in the same group.

With 200 users, B looked 40% better. Why was that not enough?

The p-value was 0.361 - a difference like that is common by luck with so few users.

What does an A/A test check?

That the setup itself does not create differences - it should find none.

Hands-on lab

Build it: an A/B test with a real significance check

The deployment lab kit's router splits simulated users 50/50, and a script measures conversions and runs a significance test, at three sizes and as an A/A check. Plain Node.js, no Docker. Every output is from a real run (Node 22).

1

Set up the lab kit

Nothing to install - only Node.js.

This lab uses the deployment lab kit's router.js and app.js, shown in full below, plus experiment.js. Version A is app.js with VERSION=v1, version B is VERSION=v2. The router's sticky mode assigns each user to one of them by a hash of their user id.

Terminal
mkdir deploy-lab && cd deploy-lab npm init -y npm pkg set type=module # no packages to install - only Node.js itself
app.js - one version of the service
// app.js - one version of the service. // Run: VERSION=v1 PORT=4001 node app.js import http from "node:http"; const VERSION = process.env.VERSION ?? "v1"; const PORT = Number(process.env.PORT ?? 4001); const ERROR_RATE = Number(process.env.ERROR_RATE ?? 0); // 0.2 = 20% of requests fail (a bad release) const STARTUP_MS = Number(process.env.STARTUP_MS ?? 0); // how long the app needs before it can serve const wait = (ms) => new Promise((resolve) => setTimeout(resolve, ms)); let inFlight = 0; const server = http.createServer(async (req, res) => { if (req.url === "/health") { res.writeHead(200); return res.end("ok"); } inFlight++; if (req.url.startsWith("/slow")) await wait(2000); // some requests take 2 seconds inFlight--; const failed = Math.random() < ERROR_RATE; res.writeHead(failed ? 500 : 200, { "content-type": "application/json" }); res.end(JSON.stringify(failed ? { version: VERSION, error: "bug in this release" } : { version: VERSION })); }); // A graceful stop: stop taking new requests, finish the ones in progress, then exit function shutdown() { console.log(VERSION, "stopping - finishing", inFlight, "requests in progress"); server.close(() => process.exit(0)); } process.on("SIGTERM", shutdown); process.on("SIGINT", shutdown); setTimeout(() => server.listen(PORT, () => console.log(VERSION, "ready on port", PORT)), STARTUP_MS);
router.js - the lab kit's load balancer
// router.js - a tiny load balancer on port 4000. Users only ever talk to this. import { createHash } from "node:crypto"; import http from "node:http"; // Where traffic goes. Change it while running with POST /admin/config. let config = { stable: ["http://localhost:4001"], // the version most users get (one or more copies) canary: [], // the new version, for some users (Canary lesson) canaryPercent: 0, // 0 - 100 sticky: false, // true = the same user always gets the same version }; let next = 0; function pickTarget(req) { const userId = req.headers["x-user-id"] ?? ""; // sticky: turn the user id into a number 0-99 - the same number every time for the same user const bucket = config.sticky ? createHash("md5").update(userId).digest().readUInt32BE(0) % 100 : Math.random() * 100; const pool = config.canary.length && bucket < config.canaryPercent ? config.canary : config.stable; return pool[next++ % pool.length]; // round robin inside the pool } const server = http.createServer((req, res) => { if (req.url === "/admin/config") { if (req.method === "GET") return res.end(JSON.stringify(config)); let body = ""; req.on("data", (chunk) => (body += chunk)); return req.on("end", () => { config = { ...config, ...JSON.parse(body) }; console.log("config:", JSON.stringify(config)); res.end(JSON.stringify(config)); }); } const target = pickTarget(req); const upstream = http.request(target + req.url, { method: req.method, headers: req.headers }, (answer) => { res.writeHead(answer.statusCode, answer.headers); answer.pipe(res); }); upstream.on("error", () => { // the instance is down or vanished mid-request if (!res.headersSent) res.writeHead(502, { "content-type": "application/json" }); res.end(JSON.stringify({ error: "upstream unavailable", target })); }); req.pipe(upstream); }); server.listen(4000, () => console.log("router on http://localhost:4000"));
2

The experiment script

Simulated users, real routing, a real significance test.

experiment.js sends one request per user through the router and records which version the user got. Then it decides whether the user "buys", using the conversion rate of that version: 10% for A and 12% for B by default. The decision comes from a hash of the user id, so every reader gets the same numbers. At the end it runs a two-proportion z-test and prints the p-value.

experiment.js
// experiment.js - an A/B test through the router: version A (v1) or B (v2), 50/50 by user. // Run with the router (sticky, 50%), v1 on 4001 and v2 on 4002: // node experiment.js <users> <A conversion> <B conversion> import { createHash } from "node:crypto"; const USERS = Number(process.argv[2] ?? 2000); const RATE = { v1: Number(process.argv[3] ?? 0.1), v2: Number(process.argv[4] ?? 0.12) }; // A simulated user: buys or not, with the chance of their version. A hash of the user id // makes it repeatable - everyone running this gets the same numbers. const chance = (id) => createHash("md5").update("buy:" + id).digest().readUInt32BE(0) / 2 ** 32; const result = { v1: { users: 0, bought: 0 }, v2: { users: 0, bought: 0 } }; for (let start = 0; start < USERS; start += 200) { // 200 users at a time await Promise.all( Array.from({ length: Math.min(200, USERS - start) }, async (_, k) => { const id = "user-" + (start + k); const { version } = await (await fetch("http://localhost:4000/", { headers: { "x-user-id": id } })).json(); result[version].users++; if (chance(id) < RATE[version]) result[version].bought++; }), ); } // Is the difference real, or luck? A two-proportion z-test. const a = result.v1, b = result.v2; const pA = a.bought / a.users, pB = b.bought / b.users; const pooled = (a.bought + b.bought) / (a.users + b.users); const z = (pB - pA) / Math.sqrt(pooled * (1 - pooled) * (1 / a.users + 1 / b.users)); const normal = (x) => { // standard normal CDF const t = 1 / (1 + 0.2316419 * Math.abs(x)); const d = 0.3989423 * Math.exp((-x * x) / 2); const p = d * t * (0.3193815 + t * (-0.3565638 + t * (1.781478 + t * (-1.821256 + t * 1.330274)))); return x > 0 ? 1 - p : p; }; const pValue = 2 * (1 - normal(Math.abs(z))); console.log(`A (v1): ${a.users} users, ${a.bought} bought = ${(pA * 100).toFixed(1)}%`); console.log(`B (v2): ${b.users} users, ${b.bought} bought = ${(pB * 100).toFixed(1)}%`); console.log(`p-value ${pValue.toFixed(3)} -> ${pValue < 0.05 ? "the difference is real (95% confidence)" : "could be luck - not enough evidence"}`);
Terminals
node router.js # terminal 1 VERSION=v1 PORT=4001 node app.js # terminal 2 - version A VERSION=v2 PORT=4002 node app.js # terminal 3 - version B # terminal 4: 50/50, sticky by user curl -X POST localhost:4000/admin/config \ -d '{"stable":["http://localhost:4001"],"canary":["http://localhost:4002"],"canaryPercent":50,"sticky":true}'
3

Run it at three sizes

The same true difference - three very different conclusions.

Run the experiment with 200, 2,000 and 20,000 users. The true rates never change. Only at 20,000 users is the evidence strong enough.

Terminal 4
node experiment.js 200 0.10 0.12 node experiment.js 2000 0.10 0.12 node experiment.js 20000 0.10 0.12
Output - measured
200 users A (v1): 89 users, 9 bought = 10.1% B (v2): 111 users, 16 bought = 14.4% p-value 0.361 -> could be luck - not enough evidence 2,000 users A (v1): 998 users, 102 bought = 10.2% B (v2): 1002 users, 114 bought = 11.4% p-value 0.405 -> could be luck - not enough evidence 20,000 users A (v1): 9934 users, 1009 bought = 10.2% B (v2): 10066 users, 1192 bought = 11.8% p-value 0.000 -> the difference is real (95% confidence)
4

Check your setup with an A/A test

Two identical versions should show no real difference.

Give both versions the same conversion rate. A good setup should find no difference. It did: 10.2% against 9.7%, p-value 0.687. Then check that users stay on one version.

Terminal 4
node experiment.js 2000 0.10 0.10 for i in 1 2 3; do curl -s -H "x-user-id: user-42" localhost:4000/; done
Output - measured
A (v1): 998 users, 102 bought = 10.2% B (v2): 1002 users, 97 bought = 9.7% p-value 0.687 -> could be luck - not enough evidence {"version":"v1"}{"version":"v1"}{"version":"v1"} <- user-42 always gets A
A/B test, measured
200 usersB looked 40% better; p = 0.361 - not real evidence.
2,000 usersp = 0.405 - still not enough.
20,000 usersp < 0.001 - B is really better.
A/A testp = 0.687 - no false difference.
Sticky assignmentuser-42 got the same version 3 times out of 3.

Practice on your own

  1. 1.

    Run experiment.js with 4,000 users, roughly the size the formula recommends. Is the result significant?

    Hint

    About 2,000 users per group - a bit above half of the recommended 3,838 per group. It may or may not be.

  2. 2.

    Run the A/A test 10 times with different user id prefixes (change "user-" to "u1-", "u2-" ...). How often does it wrongly say the difference is real?

    Hint

    At 95% confidence, about 1 time in 20 is expected.

  3. 3.

    Make B worse (0.10 against 0.08). Does the test detect a loss as well as a win?

    Hint

    The test is two-sided: it checks both directions.

  4. 4.

    Change the split to 90/10 (canaryPercent 10). How many users do you now need in total for the same confidence?

    Hint

    The smaller group limits the test - you need more users overall.

Comments

Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.

Loading comments...