Rolling Update
Replace old instances with new instances gradually.
Rolling update: one copy at a time
Kubernetes' default way to release: replace the copies of your service one at a time. This lesson explains how, why readiness checks matter so much, and why old and new versions must work together - then you build it.
The idea in short
Replace the copies of your service a few at a time.
Most services run several copies (in Kubernetes: pods) behind a load balancer. A rolling update replaces them gradually: start a new copy, wait until it is ready, add it, remove one old copy, and repeat until all copies run the new version. The service keeps serving the whole time, and you need no second environment. That is why it is Kubernetes' default.
The price: for a while, old and new versions serve users at the same time. Remember it as: one in, one out, until done.
DowntimeNone - if new copies are ready before they get traffic.RollbackAnother rolling update, backwards - minutes, not seconds.CostNo second environment; maybe one extra copy.RiskOld and new versions run together.Best forStateless APIs on Kubernetes.An everyday picture: changing the tyres one at a time
The car never stops standing on four wheels.
With only one jack, you change tyres one at a time: lift one corner, swap that tyre, lower it, move to the next. The car always stands on four wheels. For a while it has two old tyres and two new ones - so the new tyres must fit and drive well together with the old ones.
Words you need
Five words used in this lesson.
These are the Kubernetes words for a rolling update.
Copy (instance, pod)One running process of your service.Readiness check (readinessProbe)A test such as GET /health that says "this copy can serve users now".Liveness check (livenessProbe)A test that says "this copy is still alive" - if it fails, the copy is restarted.maxSurgeHow many extra copies may run during the update.maxUnavailableHow many copies may be missing during the update.How it works, step by step
Use Next to walk through a real run.
Step through the diagram: four copies of v1 replaced by v2, one at a time, while 40 requests per second kept arriving. v2 needed 1.5 seconds to start.
1 - Four copies of v1
The load balancer spreads users over four copies of v1.
A real run: four copies of v1 replaced by v2, one at a time, while 40 requests per second kept arriving.
1. StartStart one copy of v2.2. WaitWait until its readiness check answers.3. AddAdd it to the load balancer.4. RemoveRemove one v1 copy from the load balancer.5. Drain and stopLet the old copy finish its requests, then stop it.Why teams use it - and what it costs
The everyday default: no extra servers, no downtime.
A rolling update needs no second environment, keeps the service up, and is built into Kubernetes - you get it by changing the image version. In the lab it replaced four copies in about 7 seconds with 0 errors.
The costs: both versions serve users during the update (in the lab, all 20 users saw both), so changes must be backward-compatible; and rolling back is another rolling update, which takes minutes with many copies.
BenefitNo extra environment.BenefitNo downtime, when readiness checks are right.BenefitBuilt into Kubernetes.CostMixed versions during the update.CostSlower rollback than blue-green.CostA bad version reaches users copy by copy, unless you add checks.Detail 1: readiness - no users before a copy is ready
The most important check in a rolling update.
A new copy is not ready when its process starts: it may still load configuration, open database connections or warm a cache. Send users to it too early and their requests fail. In the lab, waiting for each copy's /health check gave 0 errors. Not waiting gave 62 errors - about a quarter of requests failed during each 1.5-second start-up, because one of four copies could not answer.
A readiness check decides whether a copy gets traffic. A liveness check is different: it restarts a copy that has stopped responding. Do not mix them up - a liveness check that fails during a slow start causes endless restarts.
Watch out: A /health that answers "ok" before the app can really work is worse than none. Make it check what the app needs - for example its database connection.
Detail 2: speed and capacity - maxSurge and maxUnavailable
How fast, and how much capacity you keep.
maxSurge is how many copies above the normal number may exist; maxUnavailable is how many may be missing. The lab used surge 1, unavailable 0: start the new copy first, then remove an old one, so there are never fewer than four ready copies. Faster settings either use more servers or serve users with fewer copies for a while.
surge 1, unavailable 0Safest. Never fewer than 4 ready copies; needs room for a 5th. (The lab.)surge 0, unavailable 1No extra servers; 3 copies serve users during the update.surge 25%, unavailable 25%Kubernetes' default: 1 extra and 1 missing at a time, with 4 copies.surge 100%All new copies first - close to blue-green, double capacity.Detail 3: mixed versions need compatibility
During the update, the same user meets both versions.
For several seconds in the lab, both versions answered, and every user saw both. So v2 must work together with v1: an API change must keep old fields, a new message format must still be readable by v1, and database changes follow the expand-contract steps from the blue-green lesson. If two versions truly cannot run together, use recreate instead.
Rollback works the same way: kubectl rollout undo starts another rolling update back to the previous version.
In the real world
A Kubernetes Deployment and the commands you will use.
In Kubernetes you rarely write a rolling update yourself: you change the image in a Deployment, and Kubernetes does the rounds described above. These examples show the settings and commands; they were not run in this lesson's lab, which plays Kubernetes with a small Node.js script.
apiVersion: apps/v1
kind: Deployment
metadata:
name: shop
spec:
replicas: 4
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 1 # at most 1 extra copy
maxUnavailable: 0 # never fewer than 4 ready copies
selector:
matchLabels: { app: shop }
template:
metadata:
labels: { app: shop }
spec:
containers:
- name: shop
image: shop:v2
readinessProbe: # no traffic until this answers
httpGet: { path: /health, port: 3000 }
periodSeconds: 2kubectl set image deployment/shop shop=shop:v2 # start the rolling update
kubectl rollout status deployment/shop # watch it
kubectl rollout pause deployment/shop # stop half-way to look
kubectl rollout resume deployment/shop
kubectl rollout undo deployment/shop # roll back (another rolling update)When to use it, and when not
The default for stateless services.
Most web APIs fit a rolling update well.
Good fitStateless REST or GraphQL APIs with several copies.Good fitKubernetes workloads and teams without budget for a second environment.Poor fitVersions that cannot run together - use recreate.Poor fitYou need rollback in seconds - use blue-green.Poor fitA risky change that only a few users should meet first - use canary.Common mistakes
And how to avoid each one.
The first one is responsible for most failed rolling updates.
No readiness checkAdd one that checks real dependencies (lab: 62 errors vs 0).Liveness check used as readinessUse both, for their own jobs.Breaking API or data changesMake v2 compatible with v1, or use recreate.Killing old copies without drainingGive them a grace period to finish requests.Rolling all copies at onceKeep maxUnavailable low enough to keep capacity.Compared with the other patterns
Rolling sits between blue-green and canary.
Each pattern makes a different trade between cost, speed of rollback and risk.
Rolling updateCopies replaced one by one. Cheap. Mixed versions. Slow rollback.Blue-greenTwo environments. Instant rollback. Double cost.CanaryControlled share of users, decided by metrics.RecreateStop all, then start all. Downtime, no mixed versions.Interview questions
Short answers you can give in your own words.
What is a rolling update? Replacing the instances of a service gradually - start a new one, wait until it is ready, remove an old one - so the service keeps serving throughout, without a second environment.
What do maxSurge and maxUnavailable control? How many extra instances may exist and how many may be missing during the rollout: the trade-off between speed, extra resources and capacity.
What is the main risk? Old and new versions serve traffic together, so changes must be backward-compatible; and instances must not receive traffic before they are ready.
Remember
- Replace copies a few at a time: start new, wait for ready, add it, remove an old one.
- Never send users to a copy before its readiness check passes (lab: 0 errors vs 62).
- Readiness decides traffic; liveness decides restarts.
- maxSurge and maxUnavailable trade speed against extra servers and capacity.
- Old and new versions serve users together - changes must be compatible.
- Rollback is another rolling update: kubectl rollout undo.
Check yourself
Answer in your head first, then open the answer.
What must happen before a new copy gets traffic?
Its readiness check must pass.
In the lab, how many requests failed without readiness checks?
62. With readiness checks: 0.
What does maxUnavailable: 0 guarantee?
Ready capacity never drops below the normal number of copies.
Why must v2 be compatible with v1 in a rolling update?
Both versions serve users at the same time during the update.
Which command rolls a Kubernetes Deployment back?
kubectl rollout undo deployment/<name>.
Build it: replace four copies while users keep working
The same plain Node.js lab kit - no Docker, no Kubernetes - plus a script that does what Kubernetes does: replace four copies of v1 with v2, one at a time. You run it with and without readiness checks and measure the difference. Every output is from a real run (Node 22).
Set up the lab kit
Three files, no packages, no Docker.
The lab uses three small files and nothing to install. app.js is the service: run it with VERSION=v1 or v2 and a PORT. It can be told to fail some requests (ERROR_RATE) or to start slowly (STARTUP_MS), and on Ctrl+C or a normal kill it finishes the requests it is working on before it exits.
router.js plays the load balancer (NGINX, an AWS load balancer, a Kubernetes Service). Users only ever talk to it, on port 4000. You change where it sends traffic while it runs, with POST /admin/config. traffic.js sends a steady stream of requests from 20 users and prints, every half second, how many answers came from v1, from v2, and how many failed. It ends with PASS or FAIL - that is how you check your work in every lab.
mkdir deploy-lab && cd deploy-lab
npm init -y
npm pkg set type=module
# no packages to install - the lab uses only Node.js itself// app.js - one version of the service.
// Run: VERSION=v1 PORT=4001 node app.js
import http from "node:http";
const VERSION = process.env.VERSION ?? "v1";
const PORT = Number(process.env.PORT ?? 4001);
const ERROR_RATE = Number(process.env.ERROR_RATE ?? 0); // 0.2 = 20% of requests fail (a bad release)
const STARTUP_MS = Number(process.env.STARTUP_MS ?? 0); // how long the app needs before it can serve
const wait = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
let inFlight = 0;
const server = http.createServer(async (req, res) => {
if (req.url === "/health") {
res.writeHead(200);
return res.end("ok");
}
inFlight++;
if (req.url.startsWith("/slow")) await wait(2000); // some requests take 2 seconds
inFlight--;
const failed = Math.random() < ERROR_RATE;
res.writeHead(failed ? 500 : 200, { "content-type": "application/json" });
res.end(JSON.stringify(failed ? { version: VERSION, error: "bug in this release" } : { version: VERSION }));
});
// A graceful stop: stop taking new requests, finish the ones in progress, then exit
function shutdown() {
console.log(VERSION, "stopping - finishing", inFlight, "requests in progress");
server.close(() => process.exit(0));
}
process.on("SIGTERM", shutdown);
process.on("SIGINT", shutdown);
setTimeout(() => server.listen(PORT, () => console.log(VERSION, "ready on port", PORT)), STARTUP_MS);// router.js - a tiny load balancer on port 4000. Users only ever talk to this.
import { createHash } from "node:crypto";
import http from "node:http";
// Where traffic goes. Change it while running with POST /admin/config.
let config = {
stable: ["http://localhost:4001"], // the version most users get (one or more copies)
canary: [], // the new version, for some users (Canary lesson)
canaryPercent: 0, // 0 - 100
sticky: false, // true = the same user always gets the same version
};
let next = 0;
function pickTarget(req) {
const userId = req.headers["x-user-id"] ?? "";
// sticky: turn the user id into a number 0-99 - the same number every time for the same user
const bucket = config.sticky
? createHash("md5").update(userId).digest().readUInt32BE(0) % 100
: Math.random() * 100;
const pool = config.canary.length && bucket < config.canaryPercent ? config.canary : config.stable;
return pool[next++ % pool.length]; // round robin inside the pool
}
const server = http.createServer((req, res) => {
if (req.url === "/admin/config") {
if (req.method === "GET") return res.end(JSON.stringify(config));
let body = "";
req.on("data", (chunk) => (body += chunk));
return req.on("end", () => {
config = { ...config, ...JSON.parse(body) };
console.log("config:", JSON.stringify(config));
res.end(JSON.stringify(config));
});
}
const target = pickTarget(req);
const upstream = http.request(target + req.url, { method: req.method, headers: req.headers }, (answer) => {
res.writeHead(answer.statusCode, answer.headers);
answer.pipe(res);
});
upstream.on("error", () => { // the instance is down or vanished mid-request
if (!res.headersSent) res.writeHead(502, { "content-type": "application/json" });
res.end(JSON.stringify({ error: "upstream unavailable", target }));
});
req.pipe(upstream);
});
server.listen(4000, () => console.log("router on http://localhost:4000"));// traffic.js - steady traffic through the router, like real users.
// Run: node traffic.js <seconds> <requests per second> [share of slow requests]
const seconds = Number(process.argv[2] ?? 5);
const perSecond = Number(process.argv[3] ?? 50);
const slowShare = Number(process.argv[4] ?? 0);
const buckets = []; // one row per half second
const usersSeen = new Map(); // user -> set of versions that user saw
const started = Date.now();
const pending = [];
const timer = setInterval(() => {
const user = "user-" + Math.floor(Math.random() * 20);
const path = Math.random() < slowShare ? "/slow" : "/";
const sentAt = Date.now();
pending.push(
fetch("http://localhost:4000" + path, { headers: { "x-user-id": user } })
.then(async (r) => ({ status: r.status, body: await r.json() }))
.catch(() => ({ status: 0, body: {} }))
.then(({ status, body }) => {
const slot = Math.floor((sentAt - started) / 500);
const row = (buckets[slot] ??= { v1: 0, v2: 0, errors: 0 });
if (status === 200) {
row[body.version] = (row[body.version] ?? 0) + 1;
if (!usersSeen.has(user)) usersSeen.set(user, new Set());
usersSeen.get(user).add(body.version);
} else row.errors++;
}),
);
}, 1000 / perSecond);
setTimeout(async () => {
clearInterval(timer);
await Promise.all(pending);
let total = { v1: 0, v2: 0, errors: 0 };
buckets.forEach((row, i) => {
if (!row) return;
console.log(`${(i / 2).toFixed(1).padStart(4)}s v1 ${String(row.v1).padStart(3)} v2 ${String(row.v2).padStart(3)} errors ${row.errors}`);
total = { v1: total.v1 + row.v1, v2: total.v2 + row.v2, errors: total.errors + row.errors };
});
const mixed = [...usersSeen.values()].filter((versions) => versions.size > 1).length;
console.log("total:", JSON.stringify(total), "| users who saw both versions:", mixed, "of", usersSeen.size);
console.log(total.errors === 0 ? "PASS - no request failed" : `FAIL - ${total.errors} requests failed`);
process.exit(total.errors === 0 ? 0 : 1);
}, seconds * 1000);The rolling update script
Start a new copy, wait for it, swap it in, stop an old one gracefully.
rolling.js plays Kubernetes. It starts four copies of v1 on ports 4011 to 4014 and gives them to the router. Then, for each copy: it starts a v2 copy on a new port (v2 needs 1.5 seconds to start), waits for its /health check, swaps it into the router in place of one old copy, and stops the old copy gracefully. With --no-ready it skips the wait, so you can see why the wait matters.
// rolling.js - replace 4 copies of v1 with v2, one at a time, while users keep using the app.
// Run with the router on 4000: node rolling.js (waits for each new copy to be ready)
// node rolling.js --no-ready (does not wait - see what breaks)
import { spawn } from "node:child_process";
const WAIT_FOR_READY = process.argv[2] !== "--no-ready";
const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
const started = Date.now();
const log = (text) => console.log(`${((Date.now() - started) / 1000).toFixed(1).padStart(5)}s ${text}`);
function startCopy(version, port) {
// v2 needs 1.5 s to start - like a real app loading config and opening connections
const env = { ...process.env, VERSION: version, PORT: String(port), STARTUP_MS: version === "v2" ? "1500" : "0" };
return { version, port, process: spawn("node", ["app.js"], { env, stdio: "ignore" }) };
}
async function waitUntilReady(port) {
for (;;) {
try {
if ((await fetch(`http://localhost:${port}/health`)).ok) return;
} catch {}
await sleep(100);
}
}
const usePool = (copies) =>
fetch("http://localhost:4000/admin/config", {
method: "POST",
body: JSON.stringify({ stable: copies.map((c) => `http://localhost:${c.port}`) }),
});
// 1. Four copies of v1 serve the users
const copies = [4011, 4012, 4013, 4014].map((port) => startCopy("v1", port));
await Promise.all(copies.map((c) => waitUntilReady(c.port)));
await usePool(copies);
log("serving: " + copies.map((c) => c.version).join(" "));
await sleep(2000);
// 2. Replace them one at a time: start a new copy, add it, remove an old one
for (let i = 0; i < copies.length; i++) {
const fresh = startCopy("v2", 4021 + i);
if (WAIT_FOR_READY) await waitUntilReady(fresh.port); // the readiness check
const old = copies[i];
copies[i] = fresh;
await usePool(copies); // new copy in, old copy out
old.process.kill("SIGTERM"); // let it finish its requests, then stop
log("serving: " + copies.map((c) => c.version).join(" "));
if (!WAIT_FOR_READY) await sleep(1600);
}
log("done - v2 serves everyone. Press Ctrl+C to stop the copies.");Roll with readiness checks
Four copies replaced, 0 errors.
Start the router, then rolling.js, and send traffic while it works. The script's log shows each copy changing; the traffic timeline shows v2's share growing in four steps - and no failed request.
node router.js # terminal 1
node rolling.js # terminal 2
node traffic.js 11 40 # terminal 3, right after starting rolling.jsrolling.js
0.1s serving: v1 v1 v1 v1
3.8s serving: v2 v1 v1 v1
5.5s serving: v2 v2 v1 v1
7.1s serving: v2 v2 v2 v1
8.8s serving: v2 v2 v2 v2
traffic.js
2.0s v1 19 v2 0 errors 0
3.0s v1 14 v2 5 errors 0
5.0s v1 9 v2 10 errors 0
6.5s v1 5 v2 15 errors 0
7.5s v1 0 v2 19 errors 0
total: {"v1":195,"v2":227,"errors":0} | users who saw both versions: 20 of 20
PASS - no request failedRoll without readiness checks
The same update, sending users to copies that are still starting.
Run it again with --no-ready. Each new copy joins the router the moment it is started, while it still needs 1.5 seconds before it can answer. During each of those windows, about one request in four went to a copy that could not answer. 62 requests failed.
node rolling.js --no-ready 1.0s v1 14 v2 0 errors 5
2.5s v1 10 v2 4 errors 6
4.0s v1 7 v2 8 errors 5
5.5s v1 3 v2 12 errors 4
7.0s v1 0 v2 16 errors 3
7.5s v1 0 v2 19 errors 0
total: {"v1":128,"v2":233,"errors":62} | users who saw both versions: 20 of 20
FAIL - 62 requests failedWait for /health0 errors. About 7 seconds for 4 copies.Do not wait62 errors - one copy in four could not answer while starting.Mixed versionsBoth runs: all 20 users saw v1 and v2 during the update.Practice on your own
- 1.
Change rolling.js to replace two copies at a time. How long does the update take now, and how many copies serve users at the worst moment?
Hint
Start two v2 copies, wait for both, then swap two old ones.
- 2.
Make v2 broken: start it with ERROR_RATE=0.5. Add a check to rolling.js that sends 20 test requests to each new copy and stops the rollout (keeping the rest on v1) if more than 2 fail.
Hint
This is a small canary check inside a rolling update - what Argo Rollouts adds to Kubernetes.
- 3.
Make app.js v2 answer { version, amount } instead of { version }, and write a client that reads body.version and body.amount. What does the client see during the rolling update?
Hint
Some answers come from v1 and do not have amount.
- 4.
Write the Kubernetes Deployment settings for this lab: 4 replicas, maxSurge 1, maxUnavailable 0, and a readinessProbe on /health.
Hint
The reference section above shows the YAML shape.
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...
AI
System Design
Backend
- GraphQL8 modules · 69 lessons planned
- Core Python13 modules · 75 lessons planned
- FastAPI5 sections · 20 lessons
- Node.js14 modules · 206 lessons planned
- Node.js Performance7 chapters · 36 topics
- Event Loop Lifecycle6 phases · 3 scenarios
- Docker & Containerization11 modules · 144 lessons planned
- AWS for Developers14 modules · 219 lessons planned
- CI/CD & DevOps Automation10 modules · 134 lessons planned