A/B Testing Rollout
Show two versions to two groups of users and measure which one works better.
A/B testing: which version is better?
Not a release-safety pattern but a product experiment: show two versions to two groups of users and measure which is better. This lesson explains how to split users correctly and why small numbers lie - then you run a real significance test.
The idea in short
Two versions, two groups of users, one business question.
An A/B test shows version A to one group of users and version B to another, at the same time, and compares a business result: sign-ups, purchases, clicks, return visits. It is not about safety - both versions are already safe. It asks: which version is BETTER?
Remember it as: split users, measure one goal, wait for enough data.
QuestionIs B better than A for the business?SplitBy user - each user always sees the same version.MeasuredConversions, revenue, retention - not errors.LengthDays or weeks, decided in advance.NeedsEnough users and a statistics check.An everyday picture: two shop windows
Two displays on similar streets, and a count of who walks in.
A shop with two similar branches tries a new window display in one and keeps the old display in the other. For a month it counts how many passers-by come in. Only if the new display brings in clearly more people - not just a few on one sunny Saturday - does it change every window.
Words you need
Six words used in this lesson.
These words come from experiment tools and statistics.
VariantOne version in the test: A (control, the current one) or B (the new one).ConversionThe goal action, such as buying. Conversion rate = buyers / users.AssignmentDeciding which variant a user gets - stable, by user id.Feature flagA switch in the code that shows A or B without a new deployment.p-valueThe chance of seeing a difference this big if there were really no difference.SignificantUsually p below 0.05: the difference is very unlikely to be luck.How it works, step by step
Use Next to walk through real runs.
Step through the diagram. The true conversion rates are fixed in the lab - A 10%, B 12% - and the same experiment is run with more and more users.
1 - Split users, not requests
The router hashes each user id, so the same user always sees the same version. user-42 asked three times and got v1 three times.
Real runs: the router splits users 50/50 by user id. Version B converts 12% of users, A 10%.
1. PlanOne goal metric, the smallest effect that matters, and the number of users needed.2. AssignSplit users randomly but stably - usually by hashing the user id.3. RunShow each user their variant; log variant and goal events.4. WaitUntil the planned number of users - do not stop early.5. DecideCheck significance; keep the winner, or keep A if there is no clear winner.Why teams use it - and what it costs
Decisions based on evidence, not opinions.
Teams argue about designs, prices and texts. An A/B test settles it with real behaviour. It also protects you from changes that feel better but perform worse.
The costs: it needs a lot of users (often thousands per group), weeks of patience, careful logging, and some statistics. A test that is too small does not just fail to answer - it gives confident wrong answers.
BenefitReal user behaviour decides.BenefitMeasures what matters to the business.BenefitCan test many ideas over time.CostNeeds many users and time.CostEasy to fool yourself with small samples or early stopping.CostTwo variants to maintain until the decision.Detail 1: split users, not requests
Each user must see one version, every time.
If each request chose A or B at random, a user would see the new checkout on one page and the old one on the next - confusing, and it spoils the measurement: whom do you credit with the purchase? Assign users, not requests, by hashing a stable user id. In the lab, user-42 visited three times and saw version A every time.
Careful with the hash. Code that used parseInt(userId) % 2 looks fine with numeric ids, but with ids like user-1 it put all 1,000 test users into the same group: parseInt("user-1") is not a number. Hash the id with a real hash function instead (shown in "In the real world" below). Adding the experiment name to the hash gives every experiment a different split, so the same users are not always the test group.
Detail 2: small samples lie
The lab ran the same experiment at three sizes.
The true rates never changed: A 10%, B 12%. With 200 users, B looked more than 40% better (14.4% against 10.1%) - but the p-value was 0.361: a difference like that happens by luck very often with so few users. With 2,000 users: still not enough (p 0.405). With 20,000 users: clear (p below 0.001).
Before you start, calculate how many users you need. The standard formula says: to detect a change from 10% to 12% with 95% confidence and 80% power, you need about 3,838 users in EACH group. Then run until you have them.
200 usersA 10.1%, B 14.4%. p = 0.361 - could be luck.2,000 usersA 10.2%, B 11.4%. p = 0.405 - could be luck.20,000 usersA 10.2%, B 11.8%. p < 0.001 - B is really better.A/A, 2,000 usersSame version twice: 10.2% vs 9.7%, p = 0.687 - correctly no difference.Watch out: Do not stop the test the first time the numbers look good. Checking often and stopping at the first "significant" moment finds false winners far more often than 5% of the time. Fix the sample size first.
Detail 3: plan it like an experiment
Write the plan down before you start.
Choose ONE main metric and the expected effect. Calculate the sample size. Run an A/A test once to check that your setup does not create differences by itself - in the lab it correctly found none. Run for whole weeks, because people behave differently on weekends. Watch a few guard metrics - errors, page speed, refunds - so a "winner" is not secretly harming something else.
In the real world
Feature flags and experiment platforms.
Most A/B tests do not deploy two services: they use a feature flag in one deployment, and the code shows A or B per user. Experiment platforms such as LaunchDarkly, Statsig and Optimizely handle assignment, logging and the statistics. The routing function below is the part of such a system that assigns users; it was tested here: 1,000 ids like user-1 split 505 / 495, and user-42 got the same variant every time.
import { createHash } from "node:crypto";
// Fixed: hash the user id (plus the experiment name) into a bucket 0-99
function getVariant(userId, experiment = "checkout-test") {
const bucket = createHash("md5").update(experiment + ":" + userId).digest().readUInt32BE(0) % 100;
return bucket < 50 ? "A" : "B";
}app.get("/checkout", (req, res) => {
const variant = getVariant(req.user.id, "checkout-test");
analytics.track(req.user.id, "checkout_viewed", { variant }); // log it!
res.render(variant === "A" ? "checkout-old" : "checkout-new");
});When to use it, and when not
For product questions with enough users.
An A/B test needs a clear question and enough traffic.
Good fitCheckout and pricing pages, sign-up flows, texts and layouts.Good fitRecommendation or ranking changes judged by user behaviour.Poor fitLow traffic - you may never reach the needed sample.Poor fitChecking whether a release is stable - use a canary.Poor fitChanges you must ship anyway, for legal or security reasons.Common mistakes
And how to avoid each one.
Most bad A/B decisions come from statistics, not code.
Deciding with a small sampleCalculate the size first (lab: 200 users looked like +40%, p 0.361).Stopping at the first good resultFix the length and wait.Splitting by requestSplit by a stable user id.parseInt(userId) % 2Use a real hash (old code put 1,000 of 1,000 users in one group).Many metrics, no main oneChoose one goal before you start.Compared with the other patterns
A/B asks "better?"; the others ask "safe?".
A/B testing is often confused with canary releases.
A/B testIs it better? Business metrics, days to weeks, ends with a decision.CanaryIs it safe? Errors and latency, minutes to hours, ends at 100%.ShadowIs it correct? Real traffic, no user sees it.TogetherCanary the build to prove it is safe, then A/B test the feature inside it.Interview questions
Short answers you can give in your own words.
What is A/B testing? Randomly but stably assigning users to two versions at the same time and comparing a predefined business metric, using statistics to decide whether the difference is real.
How is it different from a canary release? A canary checks operational safety with a small share of traffic and is promoted to everyone; an A/B test measures user behaviour over a fixed period to choose between versions.
Why does sample size matter? Small samples produce large random differences; without enough users per group you will often pick a false winner. Calculate it in advance from the baseline rate and the smallest effect worth detecting.
Remember
- Show A to one group and B to another at the same time; measure one business goal.
- Split by user, not by request - with a real hash, never parseInt(userId) % 2.
- Calculate the number of users first (10% -> 12% needs ~3,838 per group).
- Small samples lie: 200 users showed B +40%, but p was 0.361.
- Do not stop early; run an A/A test once to check your setup.
- A/B asks "is it better?"; canary asks "is it safe?".
Check yourself
Answer in your head first, then open the answer.
What does an A/B test measure, unlike a canary?
Business results - conversions, revenue, retention - not errors and latency.
Why must the split be by user and not by request?
So each user sees one version consistently, and results can be credited correctly.
What was wrong with parseInt(userId) % 2?
With ids like user-1, parseInt returns NaN, so every user landed in the same group.
With 200 users, B looked 40% better. Why was that not enough?
The p-value was 0.361 - a difference like that is common by luck with so few users.
What does an A/A test check?
That the setup itself does not create differences - it should find none.
Build it: an A/B test with a real significance check
The deployment lab kit's router splits simulated users 50/50, and a script measures conversions and runs a significance test, at three sizes and as an A/A check. Plain Node.js, no Docker. Every output is from a real run (Node 22).
Set up the lab kit
Nothing to install - only Node.js.
This lab uses the deployment lab kit's router.js and app.js, shown in full below, plus experiment.js. Version A is app.js with VERSION=v1, version B is VERSION=v2. The router's sticky mode assigns each user to one of them by a hash of their user id.
mkdir deploy-lab && cd deploy-lab
npm init -y
npm pkg set type=module
# no packages to install - only Node.js itself// app.js - one version of the service.
// Run: VERSION=v1 PORT=4001 node app.js
import http from "node:http";
const VERSION = process.env.VERSION ?? "v1";
const PORT = Number(process.env.PORT ?? 4001);
const ERROR_RATE = Number(process.env.ERROR_RATE ?? 0); // 0.2 = 20% of requests fail (a bad release)
const STARTUP_MS = Number(process.env.STARTUP_MS ?? 0); // how long the app needs before it can serve
const wait = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
let inFlight = 0;
const server = http.createServer(async (req, res) => {
if (req.url === "/health") {
res.writeHead(200);
return res.end("ok");
}
inFlight++;
if (req.url.startsWith("/slow")) await wait(2000); // some requests take 2 seconds
inFlight--;
const failed = Math.random() < ERROR_RATE;
res.writeHead(failed ? 500 : 200, { "content-type": "application/json" });
res.end(JSON.stringify(failed ? { version: VERSION, error: "bug in this release" } : { version: VERSION }));
});
// A graceful stop: stop taking new requests, finish the ones in progress, then exit
function shutdown() {
console.log(VERSION, "stopping - finishing", inFlight, "requests in progress");
server.close(() => process.exit(0));
}
process.on("SIGTERM", shutdown);
process.on("SIGINT", shutdown);
setTimeout(() => server.listen(PORT, () => console.log(VERSION, "ready on port", PORT)), STARTUP_MS);// router.js - a tiny load balancer on port 4000. Users only ever talk to this.
import { createHash } from "node:crypto";
import http from "node:http";
// Where traffic goes. Change it while running with POST /admin/config.
let config = {
stable: ["http://localhost:4001"], // the version most users get (one or more copies)
canary: [], // the new version, for some users (Canary lesson)
canaryPercent: 0, // 0 - 100
sticky: false, // true = the same user always gets the same version
};
let next = 0;
function pickTarget(req) {
const userId = req.headers["x-user-id"] ?? "";
// sticky: turn the user id into a number 0-99 - the same number every time for the same user
const bucket = config.sticky
? createHash("md5").update(userId).digest().readUInt32BE(0) % 100
: Math.random() * 100;
const pool = config.canary.length && bucket < config.canaryPercent ? config.canary : config.stable;
return pool[next++ % pool.length]; // round robin inside the pool
}
const server = http.createServer((req, res) => {
if (req.url === "/admin/config") {
if (req.method === "GET") return res.end(JSON.stringify(config));
let body = "";
req.on("data", (chunk) => (body += chunk));
return req.on("end", () => {
config = { ...config, ...JSON.parse(body) };
console.log("config:", JSON.stringify(config));
res.end(JSON.stringify(config));
});
}
const target = pickTarget(req);
const upstream = http.request(target + req.url, { method: req.method, headers: req.headers }, (answer) => {
res.writeHead(answer.statusCode, answer.headers);
answer.pipe(res);
});
upstream.on("error", () => { // the instance is down or vanished mid-request
if (!res.headersSent) res.writeHead(502, { "content-type": "application/json" });
res.end(JSON.stringify({ error: "upstream unavailable", target }));
});
req.pipe(upstream);
});
server.listen(4000, () => console.log("router on http://localhost:4000"));The experiment script
Simulated users, real routing, a real significance test.
experiment.js sends one request per user through the router and records which version the user got. Then it decides whether the user "buys", using the conversion rate of that version: 10% for A and 12% for B by default. The decision comes from a hash of the user id, so every reader gets the same numbers. At the end it runs a two-proportion z-test and prints the p-value.
// experiment.js - an A/B test through the router: version A (v1) or B (v2), 50/50 by user.
// Run with the router (sticky, 50%), v1 on 4001 and v2 on 4002:
// node experiment.js <users> <A conversion> <B conversion>
import { createHash } from "node:crypto";
const USERS = Number(process.argv[2] ?? 2000);
const RATE = { v1: Number(process.argv[3] ?? 0.1), v2: Number(process.argv[4] ?? 0.12) };
// A simulated user: buys or not, with the chance of their version. A hash of the user id
// makes it repeatable - everyone running this gets the same numbers.
const chance = (id) => createHash("md5").update("buy:" + id).digest().readUInt32BE(0) / 2 ** 32;
const result = { v1: { users: 0, bought: 0 }, v2: { users: 0, bought: 0 } };
for (let start = 0; start < USERS; start += 200) { // 200 users at a time
await Promise.all(
Array.from({ length: Math.min(200, USERS - start) }, async (_, k) => {
const id = "user-" + (start + k);
const { version } = await (await fetch("http://localhost:4000/", { headers: { "x-user-id": id } })).json();
result[version].users++;
if (chance(id) < RATE[version]) result[version].bought++;
}),
);
}
// Is the difference real, or luck? A two-proportion z-test.
const a = result.v1, b = result.v2;
const pA = a.bought / a.users, pB = b.bought / b.users;
const pooled = (a.bought + b.bought) / (a.users + b.users);
const z = (pB - pA) / Math.sqrt(pooled * (1 - pooled) * (1 / a.users + 1 / b.users));
const normal = (x) => { // standard normal CDF
const t = 1 / (1 + 0.2316419 * Math.abs(x));
const d = 0.3989423 * Math.exp((-x * x) / 2);
const p = d * t * (0.3193815 + t * (-0.3565638 + t * (1.781478 + t * (-1.821256 + t * 1.330274))));
return x > 0 ? 1 - p : p;
};
const pValue = 2 * (1 - normal(Math.abs(z)));
console.log(`A (v1): ${a.users} users, ${a.bought} bought = ${(pA * 100).toFixed(1)}%`);
console.log(`B (v2): ${b.users} users, ${b.bought} bought = ${(pB * 100).toFixed(1)}%`);
console.log(`p-value ${pValue.toFixed(3)} -> ${pValue < 0.05 ? "the difference is real (95% confidence)" : "could be luck - not enough evidence"}`);node router.js # terminal 1
VERSION=v1 PORT=4001 node app.js # terminal 2 - version A
VERSION=v2 PORT=4002 node app.js # terminal 3 - version B
# terminal 4: 50/50, sticky by user
curl -X POST localhost:4000/admin/config \
-d '{"stable":["http://localhost:4001"],"canary":["http://localhost:4002"],"canaryPercent":50,"sticky":true}'Run it at three sizes
The same true difference - three very different conclusions.
Run the experiment with 200, 2,000 and 20,000 users. The true rates never change. Only at 20,000 users is the evidence strong enough.
node experiment.js 200 0.10 0.12
node experiment.js 2000 0.10 0.12
node experiment.js 20000 0.10 0.12200 users
A (v1): 89 users, 9 bought = 10.1%
B (v2): 111 users, 16 bought = 14.4%
p-value 0.361 -> could be luck - not enough evidence
2,000 users
A (v1): 998 users, 102 bought = 10.2%
B (v2): 1002 users, 114 bought = 11.4%
p-value 0.405 -> could be luck - not enough evidence
20,000 users
A (v1): 9934 users, 1009 bought = 10.2%
B (v2): 10066 users, 1192 bought = 11.8%
p-value 0.000 -> the difference is real (95% confidence)Check your setup with an A/A test
Two identical versions should show no real difference.
Give both versions the same conversion rate. A good setup should find no difference. It did: 10.2% against 9.7%, p-value 0.687. Then check that users stay on one version.
node experiment.js 2000 0.10 0.10
for i in 1 2 3; do curl -s -H "x-user-id: user-42" localhost:4000/; doneA (v1): 998 users, 102 bought = 10.2%
B (v2): 1002 users, 97 bought = 9.7%
p-value 0.687 -> could be luck - not enough evidence
{"version":"v1"}{"version":"v1"}{"version":"v1"} <- user-42 always gets A200 usersB looked 40% better; p = 0.361 - not real evidence.2,000 usersp = 0.405 - still not enough.20,000 usersp < 0.001 - B is really better.A/A testp = 0.687 - no false difference.Sticky assignmentuser-42 got the same version 3 times out of 3.Practice on your own
- 1.
Run experiment.js with 4,000 users, roughly the size the formula recommends. Is the result significant?
Hint
About 2,000 users per group - a bit above half of the recommended 3,838 per group. It may or may not be.
- 2.
Run the A/A test 10 times with different user id prefixes (change "user-" to "u1-", "u2-" ...). How often does it wrongly say the difference is real?
Hint
At 95% confidence, about 1 time in 20 is expected.
- 3.
Make B worse (0.10 against 0.08). Does the test detect a loss as well as a win?
Hint
The test is two-sided: it checks both directions.
- 4.
Change the split to 90/10 (canaryPercent 10). How many users do you now need in total for the same confidence?
Hint
The smaller group limits the test - you need more users overall.
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...
AI
System Design
Backend
- GraphQL8 modules · 69 lessons planned
- Core Python13 modules · 75 lessons planned
- FastAPI5 sections · 20 lessons
- Node.js14 modules · 206 lessons planned
- Node.js Performance7 chapters · 36 topics
- Event Loop Lifecycle6 phases · 3 scenarios
- Docker & Containerization11 modules · 144 lessons planned
- AWS for Developers14 modules · 219 lessons planned
- CI/CD & DevOps Automation10 modules · 134 lessons planned