Canary Deployment
Release to a small percentage of traffic, monitor, then gradually increase.
Canary: let a few users find the bugs
Give the new version to a few users first, and widen only while the numbers are good. This lesson explains the loop, how to decide with enough data, and how to set it up - then you build an automatic canary.
The idea in short
A few users first. Watch. Widen, or roll back.
A canary release sends a small share of real users - for example 5% - to the new version. Everyone else stays on the old version. You compare the two: errors, speed, and business numbers. If the new version is as good, you give it more users step by step: 25%, 50%, 100%. If it is worse, you send those few users back, and most users never noticed anything.
Remember it as: small, watch, widen.
DowntimeNone.Users affected by a bugOnly the current canary share.RollbackFast - set the share back to 0%.CostNo second environment; a few extra copies.NeedsWeighted routing and good metrics.An everyday picture: a new recipe in a restaurant
Serve it at a few tables first.
A restaurant changes the recipe of its most popular dish. It does not change it for the whole room on a busy Saturday. It serves the new version at two or three tables and watches the plates that come back. If people like it, more tables get it. If plates come back half-eaten, the kitchen goes back to the old recipe - and only three tables had a bad meal.
The name comes from coal mines: miners took a canary underground because it reacted to dangerous gas before people did. Your few canary users are the early warning.
Words you need
Five words used in this lesson.
These words appear in every canary tool.
CanaryThe new version, serving a small share of traffic.Stable (baseline)The current version, serving everyone else.Weight (traffic split)The percentage of requests or users sent to the canary.PromoteGive the canary more traffic, finally 100%.Canary analysisComparing the canary's metrics with the stable version's to decide: promote, wait or roll back.How it works, step by step
Use Next to walk through a real run with a buggy release.
Step through the diagram. It is a real run of the lab's automatic canary: v2 had a bug that failed 20% of requests.
1 - 5% of traffic goes to v2
The router sends 5% of requests to v2. In 4 seconds that was only 19 requests - fewer than the 20 the script needs before it may judge. So it cannot decide yet, and moves on to the next step.
A real run of canary.js with a bad v2 that fails 20% of requests.
1. DeployRun v2 beside v1. It gets 0% of traffic.2. Shift a littleSend a small share (1-5%) to v2.3. CompareErrors, speed and business numbers: v2 against v1, at the same time.4. DecideGood: widen to the next step. Bad: back to 0%.5. FinishAt 100%, v2 becomes the new stable version.Why teams use it - and what it costs
Real users find bugs - but only a few of them.
Tests never cover everything real users do. A canary lets real traffic find the problems your tests missed, while limiting how many people meet them. The lab shows the difference: releasing a buggy v2 to everyone at once gave 77 failed requests in 4 seconds - and they would have continued. The canary stopped after 26.
The cost is time and tooling. A careful canary takes minutes to hours. And it only works if you can measure each version separately and trust the numbers.
BenefitA bug reaches only the canary share.BenefitDecisions are based on real traffic, not guesses.BenefitNo second full environment.CostNeeds weighted routing and per-version metrics.CostOld and new versions run together, so they must be compatible.CostSlower than switching everyone at once.Detail 1: decide with enough data
Five requests prove nothing.
At 5%, the lab's canary received only 19 requests in 4 seconds. Five of them failed - alarming, but 19 requests is too few to be sure. So the script requires at least 20 before it may decide, and it only judged at 25%, with 112 requests. In production, small steps need longer waits, or more traffic.
Compare the canary with the stable version over the SAME period, not with yesterday. If the whole site is slow because of a database problem, both versions are slow, and that is not the canary's fault.
Detail 2: decide the rules before you start
Write down what "good" and "bad" mean.
Decide the steps and the limits before the release, not while you watch the graphs. A common ladder is 1% -> 5% -> 10% -> 25% -> 50% -> 100%, with a pause at each step. Common limits: the error rate may not rise more than about 0.1 to 1 percentage point, the slowest requests (p99 latency) may not get more than about 10-20% slower, CPU and memory stay normal, and business numbers - orders, sign-ups, payments - do not drop.
Error rateShare of failed requests, canary against stable.Latency (p95, p99)How slow the slowest 5% or 1% of requests are.SaturationCPU and memory of the canary copies.Business numbersOrders, sign-ups, payments per minute. A release can answer 200 and still break checkout.Detail 3: random or sticky?
Should one user see both versions?
The simplest split picks v1 or v2 at random for every request. In the lab at 25%, 19 of 20 users saw both versions, jumping between them. For a website that is confusing. A sticky split hashes the user id, so each user always lands on the same version: 0 of 20 users saw both.
Two things from the lab: use a real hash - a home-made one put user-0 to user-9 into buckets 70-79, so at 25% nobody got the canary - and remember that with only a few users the share is uneven (11 of 20 users landed in the 25%).
In the real world
NGINX, Kubernetes with Argo Rollouts, and service meshes.
NGINX can split traffic with split_clients, which hashes a value - here the client address - into buckets, so a user's requests stay on one side. In Kubernetes, Argo Rollouts and Flagger run the whole canary for you: the steps, the pauses, and the metric checks that promote or roll back automatically. Service meshes such as Istio split by exact percentages.
These configuration examples show the shape. They were not run in this lesson's lab.
upstream stable { server 127.0.0.1:3001; }
upstream canary { server 127.0.0.1:3002; }
# hash the client address into a bucket: 5% canary, the rest stable
split_clients "${remote_addr}" $pool {
5% canary;
* stable;
}
server {
listen 80;
location / {
proxy_pass http://$pool;
}
}apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: shop
spec:
replicas: 10
strategy:
canary:
steps:
- setWeight: 5
- pause: { duration: 10m } # watch the metrics
- setWeight: 25
- pause: { duration: 10m }
- setWeight: 50
- pause: { duration: 10m } # then 100%
selector:
matchLabels: { app: shop }
template:
metadata:
labels: { app: shop }
spec:
containers:
- name: shop
image: shop:v2When to use it, and when not
Most useful where many users and good monitoring meet.
A canary pays off when a bug would hurt many users and you can measure each version.
Good fitHigh-traffic APIs and shops; new algorithms or models; risky code changes.Good fitTeams that already have monitoring per version (Prometheus, Datadog).Poor fitVery low traffic - it takes too long to collect enough data.Poor fitChanges where old and new cannot run together.Poor fitNo per-version metrics - you would be guessing.Common mistakes
And how to avoid each one.
Most canary failures are failures of measurement, not of routing.
Judging on too few requestsSet a minimum number per step (lab: 19 were too few).Comparing with yesterdayCompare canary and stable at the same time.Watching only errorsAdd latency and business numbers.Random split on a websiteUse a sticky split by user id.Deciding the limits during the releaseWrite them down before you start.Compared with the other patterns
Canary checks safety; A/B checks value.
Canary is often confused with A/B testing because both send some users to a new version.
Blue-greenTests first, then everyone at once.CanaryA few real users first, then more - asks "is it safe?".A/B testTwo versions for weeks - asks "is it better?" using business numbers.ShadowReal traffic, but no user ever gets the new version's answer.Interview questions
Short answers you can give in your own words.
What is a canary release? Routing a small, growing share of production traffic to a new version, comparing its metrics with the stable version at each step, and promoting or rolling back based on that comparison.
How do you decide whether to promote? With limits set in advance on errors, latency and business metrics, compared against the stable version over the same period, and only after enough requests have reached the canary.
How do you keep a user on the same version? Route by a hash of a stable user identifier instead of choosing randomly for each request.
Remember
- Send a small share of users to the new version; everyone else stays on the old one.
- Compare canary and stable at the same time: errors, latency, business numbers.
- Widen step by step; roll back to 0% the moment it is worse.
- Do not judge on a handful of requests - set a minimum per step.
- Decide the steps and limits before you start.
- Use a sticky split (hash of the user id) for websites.
Check yourself
Answer in your head first, then open the answer.
What share of traffic does a canary usually start with?
A small share - often 1% to 5%.
In the lab, why could the canary not decide at 5%?
Only 19 requests reached v2 - fewer than the minimum of 20.
A bad v2 (20% errors): how many requests failed with the canary, and with everyone at once?
26 with the canary; 77 in just 4 seconds with everyone at once.
Why compare the canary with the stable version at the same time?
Problems that affect both, like a slow database, would otherwise be blamed on the canary.
How do you keep one user on one version?
Hash the user id into a bucket and route by the bucket.
Build it: an automatic canary that rolls back a bad release
The same plain Node.js lab kit - no Docker, nothing to install - plus a canary script that widens v2 step by step and rolls it back if its error rate is worse than v1's. You release a good v2 and a bad one, and compare with releasing to everyone at once. Every output is from a real run (Node 22).
Set up the lab kit
Three files, no packages, no Docker.
The lab uses three small files and nothing to install. app.js is the service: run it with VERSION=v1 or v2 and a PORT. It can be told to fail some requests (ERROR_RATE) or to start slowly (STARTUP_MS), and on Ctrl+C or a normal kill it finishes the requests it is working on before it exits.
router.js plays the load balancer (NGINX, an AWS load balancer, a Kubernetes Service). Users only ever talk to it, on port 4000. You change where it sends traffic while it runs, with POST /admin/config. traffic.js sends a steady stream of requests from 20 users and prints, every half second, how many answers came from v1, from v2, and how many failed. It ends with PASS or FAIL - that is how you check your work in every lab.
mkdir deploy-lab && cd deploy-lab
npm init -y
npm pkg set type=module
# no packages to install - the lab uses only Node.js itself// app.js - one version of the service.
// Run: VERSION=v1 PORT=4001 node app.js
import http from "node:http";
const VERSION = process.env.VERSION ?? "v1";
const PORT = Number(process.env.PORT ?? 4001);
const ERROR_RATE = Number(process.env.ERROR_RATE ?? 0); // 0.2 = 20% of requests fail (a bad release)
const STARTUP_MS = Number(process.env.STARTUP_MS ?? 0); // how long the app needs before it can serve
const wait = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
let inFlight = 0;
const server = http.createServer(async (req, res) => {
if (req.url === "/health") {
res.writeHead(200);
return res.end("ok");
}
inFlight++;
if (req.url.startsWith("/slow")) await wait(2000); // some requests take 2 seconds
inFlight--;
const failed = Math.random() < ERROR_RATE;
res.writeHead(failed ? 500 : 200, { "content-type": "application/json" });
res.end(JSON.stringify(failed ? { version: VERSION, error: "bug in this release" } : { version: VERSION }));
});
// A graceful stop: stop taking new requests, finish the ones in progress, then exit
function shutdown() {
console.log(VERSION, "stopping - finishing", inFlight, "requests in progress");
server.close(() => process.exit(0));
}
process.on("SIGTERM", shutdown);
process.on("SIGINT", shutdown);
setTimeout(() => server.listen(PORT, () => console.log(VERSION, "ready on port", PORT)), STARTUP_MS);// router.js - a tiny load balancer on port 4000. Users only ever talk to this.
import { createHash } from "node:crypto";
import http from "node:http";
// Where traffic goes. Change it while running with POST /admin/config.
let config = {
stable: ["http://localhost:4001"], // the version most users get (one or more copies)
canary: [], // the new version, for some users (Canary lesson)
canaryPercent: 0, // 0 - 100
sticky: false, // true = the same user always gets the same version
};
let next = 0;
function pickTarget(req) {
const userId = req.headers["x-user-id"] ?? "";
// sticky: turn the user id into a number 0-99 - the same number every time for the same user
const bucket = config.sticky
? createHash("md5").update(userId).digest().readUInt32BE(0) % 100
: Math.random() * 100;
const pool = config.canary.length && bucket < config.canaryPercent ? config.canary : config.stable;
return pool[next++ % pool.length]; // round robin inside the pool
}
const server = http.createServer((req, res) => {
if (req.url === "/admin/config") {
if (req.method === "GET") return res.end(JSON.stringify(config));
let body = "";
req.on("data", (chunk) => (body += chunk));
return req.on("end", () => {
config = { ...config, ...JSON.parse(body) };
console.log("config:", JSON.stringify(config));
res.end(JSON.stringify(config));
});
}
const target = pickTarget(req);
const upstream = http.request(target + req.url, { method: req.method, headers: req.headers }, (answer) => {
res.writeHead(answer.statusCode, answer.headers);
answer.pipe(res);
});
upstream.on("error", () => { // the instance is down or vanished mid-request
if (!res.headersSent) res.writeHead(502, { "content-type": "application/json" });
res.end(JSON.stringify({ error: "upstream unavailable", target }));
});
req.pipe(upstream);
});
server.listen(4000, () => console.log("router on http://localhost:4000"));// traffic.js - steady traffic through the router, like real users.
// Run: node traffic.js <seconds> <requests per second> [share of slow requests]
const seconds = Number(process.argv[2] ?? 5);
const perSecond = Number(process.argv[3] ?? 50);
const slowShare = Number(process.argv[4] ?? 0);
const buckets = []; // one row per half second
const usersSeen = new Map(); // user -> set of versions that user saw
const started = Date.now();
const pending = [];
const timer = setInterval(() => {
const user = "user-" + Math.floor(Math.random() * 20);
const path = Math.random() < slowShare ? "/slow" : "/";
const sentAt = Date.now();
pending.push(
fetch("http://localhost:4000" + path, { headers: { "x-user-id": user } })
.then(async (r) => ({ status: r.status, body: await r.json() }))
.catch(() => ({ status: 0, body: {} }))
.then(({ status, body }) => {
const slot = Math.floor((sentAt - started) / 500);
const row = (buckets[slot] ??= { v1: 0, v2: 0, errors: 0 });
if (status === 200) {
row[body.version] = (row[body.version] ?? 0) + 1;
if (!usersSeen.has(user)) usersSeen.set(user, new Set());
usersSeen.get(user).add(body.version);
} else row.errors++;
}),
);
}, 1000 / perSecond);
setTimeout(async () => {
clearInterval(timer);
await Promise.all(pending);
let total = { v1: 0, v2: 0, errors: 0 };
buckets.forEach((row, i) => {
if (!row) return;
console.log(`${(i / 2).toFixed(1).padStart(4)}s v1 ${String(row.v1).padStart(3)} v2 ${String(row.v2).padStart(3)} errors ${row.errors}`);
total = { v1: total.v1 + row.v1, v2: total.v2 + row.v2, errors: total.errors + row.errors };
});
const mixed = [...usersSeen.values()].filter((versions) => versions.size > 1).length;
console.log("total:", JSON.stringify(total), "| users who saw both versions:", mixed, "of", usersSeen.size);
console.log(total.errors === 0 ? "PASS - no request failed" : `FAIL - ${total.errors} requests failed`);
process.exit(total.errors === 0 ? 0 : 1);
}, seconds * 1000);The automatic canary
Widen v2 step by step; roll back if it is worse than v1.
canary.js does what a release tool does. For each step - 5%, 25%, 50%, 100% - it sets v2's share in the router, sends 100 requests per second for 4 seconds, and counts errors per version. If at least 20 requests reached v2 and v2's error rate is more than 2 points above v1's, it sets v2's share to 0 and stops with exit code 1. If v2 survives every step, it makes v2 the stable version for everyone.
// canary.js - an automatic canary: widen v2 step by step, roll back if it is worse than v1.
// Run with the router, v1 on 4001 and v2 on 4002: node canary.js
const STEPS = [5, 25, 50, 100]; // percent of traffic for v2
const STEP_SECONDS = 4;
const PER_SECOND = 100;
const MIN_CANARY_REQUESTS = 20; // do not judge on too few requests
const MAX_EXTRA_ERRORS = 0.02; // v2 may be at most 2 points worse than v1
const setConfig = (config) =>
fetch("http://localhost:4000/admin/config", { method: "POST", body: JSON.stringify(config) });
async function measure(seconds) {
const counts = { v1: { ok: 0, failed: 0 }, v2: { ok: 0, failed: 0 } };
const requests = [];
for (let i = 0; i < seconds * PER_SECOND; i++) {
requests.push(
fetch("http://localhost:4000/", { headers: { "x-user-id": "user-" + (i % 50) } })
.then(async (r) => ({ ok: r.ok, body: await r.json() }))
.then(({ ok, body }) => counts[body.version][ok ? "ok" : "failed"]++),
);
await new Promise((resolve) => setTimeout(resolve, 1000 / PER_SECOND));
}
await Promise.all(requests);
return counts;
}
const rate = (c) => c.failed / Math.max(1, c.ok + c.failed);
await setConfig({ stable: ["http://localhost:4001"], canary: ["http://localhost:4002"], canaryPercent: 0 });
let customersHitByV2Errors = 0;
for (const percent of STEPS) {
await setConfig({ canaryPercent: percent });
const counts = await measure(STEP_SECONDS);
const v2Total = counts.v2.ok + counts.v2.failed;
customersHitByV2Errors += counts.v2.failed;
console.log(
`${String(percent).padStart(3)}% to v2 | v1 errors ${(rate(counts.v1) * 100).toFixed(1)}% of ${counts.v1.ok + counts.v1.failed}`,
`| v2 errors ${(rate(counts.v2) * 100).toFixed(1)}% of ${v2Total}`,
);
if (v2Total >= MIN_CANARY_REQUESTS && rate(counts.v2) > rate(counts.v1) + MAX_EXTRA_ERRORS) {
await setConfig({ canaryPercent: 0 });
console.log(`ROLLED BACK at ${percent}% - v2 is worse than v1. Requests that hit v2's bug: ${customersHitByV2Errors}`);
process.exit(1);
}
}
await setConfig({ stable: ["http://localhost:4002"], canary: [], canaryPercent: 0 });
console.log(`PROMOTED - v2 now serves everyone. Requests that hit v2 errors on the way: ${customersHitByV2Errors}`);Release a good v2
It passes every step and is promoted.
Start the router, v1 and a healthy v2, then run the canary.
node router.js # terminal 1
VERSION=v1 PORT=4001 node app.js # terminal 2
VERSION=v2 PORT=4002 node app.js # terminal 3
node canary.js # terminal 4 5% to v2 | v1 errors 0.0% of 366 | v2 errors 0.0% of 34
25% to v2 | v1 errors 0.0% of 294 | v2 errors 0.0% of 106
50% to v2 | v1 errors 0.0% of 199 | v2 errors 0.0% of 201
100% to v2 | v1 errors 0.0% of 0 | v2 errors 0.0% of 400
PROMOTED - v2 now serves everyone. Requests that hit v2 errors on the way: 0Release a bad v2
20% of its requests fail. The canary stops it.
Restart v2 with ERROR_RATE=0.2 and run the canary again. At 5%, 19 requests reached v2 - too few to judge. At 25%, 112 requests showed 18.8% errors against 0.0%, and the canary rolled back. 26 requests hit the bug.
For comparison, we sent the same bad v2 to everyone at once for 4 seconds: 77 requests failed - and in a real release they would keep failing until someone noticed.
VERSION=v2 PORT=4002 ERROR_RATE=0.2 node app.js # terminal 3 (restart v2)
node canary.js # terminal 4 5% to v2 | v1 errors 0.0% of 381 | v2 errors 26.3% of 19 <- too few to judge
25% to v2 | v1 errors 0.0% of 288 | v2 errors 18.8% of 112
ROLLED BACK at 25% - v2 is worse than v1. Requests that hit v2's bug: 26
the same bad v2 released to everyone at once, 4 s at 100 requests/s:
total: {"v1":0,"v2":289,"errors":77}CanaryRolled back at 25%; 26 requests hit the bug.Everyone at once77 requests failed in the first 4 seconds alone.Random split or sticky split
Should one user see both versions?
Set a 25% canary and send traffic from 20 users, first with a random split, then with sticky: true. The router's sticky mode uses an MD5 hash of the x-user-id header, so each user always lands in the same bucket.
curl -X POST localhost:4000/admin/config \
-d '{"stable":["http://localhost:4001"],"canary":["http://localhost:4002"],"canaryPercent":25,"sticky":false}'
node traffic.js 4 50
curl -X POST localhost:4000/admin/config -d '{"sticky":true}'
node traffic.js 4 50random split, 25%:
total: {"v1":145,"v2":45,"errors":0} | users who saw both versions: 19 of 20
sticky split, 25%:
total: {"v1":77,"v2":113,"errors":0} | users who saw both versions: 0 of 20
(11 of the 20 users hash under 25, so v2 got more than 25% of this small group)Watch out: Use a real hash function for sticky routing. Our first attempt, a simple character sum, put user-0 to user-9 into buckets 70-79 - at 25% nobody got the canary. MD5 spread 10,000 users to 25.69%.
Practice on your own
- 1.
Change MAX_EXTRA_ERRORS to 0.3 and release the bad v2 again. Does the canary still catch it? What does this tell you about choosing thresholds?
Hint
18.8% is less than 0% + 30 points.
- 2.
Make STEP_SECONDS 8 and run the bad release again. Can the canary judge at 5% now?
Hint
Twice as long means about twice as many requests reach v2.
- 3.
Add latency to the canary check: make v2 slow with a small change to app.js, and roll back if v2's average response time is 50% higher than v1's.
Hint
Measure Date.now() before and after each fetch, per version.
- 4.
Change the sticky hash to use the user id plus a release name, so each release picks a different group of users. Why might that be fairer?
Hint
Otherwise the same users are always the guinea pigs.
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...