← Back to microservices patterns map
Microservices Pattern

Rolling Update

Replace old instances with new instances gradually.

deploy
Lesson

Rolling update: one copy at a time

Kubernetes' default way to release: replace the copies of your service one at a time. This lesson explains how, why readiness checks matter so much, and why old and new versions must work together - then you build it.

01

The idea in short

Replace the copies of your service a few at a time.

Most services run several copies (in Kubernetes: pods) behind a load balancer. A rolling update replaces them gradually: start a new copy, wait until it is ready, add it, remove one old copy, and repeat until all copies run the new version. The service keeps serving the whole time, and you need no second environment. That is why it is Kubernetes' default.

The price: for a while, old and new versions serve users at the same time. Remember it as: one in, one out, until done.

At a glance
DowntimeNone - if new copies are ready before they get traffic.
RollbackAnother rolling update, backwards - minutes, not seconds.
CostNo second environment; maybe one extra copy.
RiskOld and new versions run together.
Best forStateless APIs on Kubernetes.
02

An everyday picture: changing the tyres one at a time

The car never stops standing on four wheels.

With only one jack, you change tyres one at a time: lift one corner, swap that tyre, lower it, move to the next. The car always stands on four wheels. For a while it has two old tyres and two new ones - so the new tyres must fit and drive well together with the old ones.

03

Words you need

Five words used in this lesson.

These are the Kubernetes words for a rolling update.

Small dictionary
Copy (instance, pod)One running process of your service.
Readiness check (readinessProbe)A test such as GET /health that says "this copy can serve users now".
Liveness check (livenessProbe)A test that says "this copy is still alive" - if it fails, the copy is restarted.
maxSurgeHow many extra copies may run during the update.
maxUnavailableHow many copies may be missing during the update.
04

How it works, step by step

Use Next to walk through a real run.

Step through the diagram: four copies of v1 replaced by v2, one at a time, while 40 requests per second kept arriving. v2 needed 1.5 seconds to start.

workflowRolling update: one copy at a timestep 1 / 4

1 - Four copies of v1

The load balancer spreads users over four copies of v1.

copies
v1 v1 v1 v1
errors
0
extra servers
none
users mixing versions
not yet

A real run: four copies of v1 replaced by v2, one at a time, while 40 requests per second kept arriving.

One round, repeated for every copy
1. StartStart one copy of v2.
2. WaitWait until its readiness check answers.
3. AddAdd it to the load balancer.
4. RemoveRemove one v1 copy from the load balancer.
5. Drain and stopLet the old copy finish its requests, then stop it.
05

Why teams use it - and what it costs

The everyday default: no extra servers, no downtime.

A rolling update needs no second environment, keeps the service up, and is built into Kubernetes - you get it by changing the image version. In the lab it replaced four copies in about 7 seconds with 0 errors.

The costs: both versions serve users during the update (in the lab, all 20 users saw both), so changes must be backward-compatible; and rolling back is another rolling update, which takes minutes with many copies.

Benefits and costs
BenefitNo extra environment.
BenefitNo downtime, when readiness checks are right.
BenefitBuilt into Kubernetes.
CostMixed versions during the update.
CostSlower rollback than blue-green.
CostA bad version reaches users copy by copy, unless you add checks.
06

Detail 1: readiness - no users before a copy is ready

The most important check in a rolling update.

A new copy is not ready when its process starts: it may still load configuration, open database connections or warm a cache. Send users to it too early and their requests fail. In the lab, waiting for each copy's /health check gave 0 errors. Not waiting gave 62 errors - about a quarter of requests failed during each 1.5-second start-up, because one of four copies could not answer.

A readiness check decides whether a copy gets traffic. A liveness check is different: it restarts a copy that has stopped responding. Do not mix them up - a liveness check that fails during a slow start causes endless restarts.

Watch out: A /health that answers "ok" before the app can really work is worse than none. Make it check what the app needs - for example its database connection.

07

Detail 2: speed and capacity - maxSurge and maxUnavailable

How fast, and how much capacity you keep.

maxSurge is how many copies above the normal number may exist; maxUnavailable is how many may be missing. The lab used surge 1, unavailable 0: start the new copy first, then remove an old one, so there are never fewer than four ready copies. Faster settings either use more servers or serve users with fewer copies for a while.

Settings for 4 copies
surge 1, unavailable 0Safest. Never fewer than 4 ready copies; needs room for a 5th. (The lab.)
surge 0, unavailable 1No extra servers; 3 copies serve users during the update.
surge 25%, unavailable 25%Kubernetes' default: 1 extra and 1 missing at a time, with 4 copies.
surge 100%All new copies first - close to blue-green, double capacity.
08

Detail 3: mixed versions need compatibility

During the update, the same user meets both versions.

For several seconds in the lab, both versions answered, and every user saw both. So v2 must work together with v1: an API change must keep old fields, a new message format must still be readable by v1, and database changes follow the expand-contract steps from the blue-green lesson. If two versions truly cannot run together, use recreate instead.

Rollback works the same way: kubectl rollout undo starts another rolling update back to the previous version.

09

In the real world

A Kubernetes Deployment and the commands you will use.

In Kubernetes you rarely write a rolling update yourself: you change the image in a Deployment, and Kubernetes does the rounds described above. These examples show the settings and commands; they were not run in this lesson's lab, which plays Kubernetes with a small Node.js script.

Kubernetes Deployment - a rolling update with a readiness check
apiVersion: apps/v1 kind: Deployment metadata: name: shop spec: replicas: 4 strategy: type: RollingUpdate rollingUpdate: maxSurge: 1 # at most 1 extra copy maxUnavailable: 0 # never fewer than 4 ready copies selector: matchLabels: { app: shop } template: metadata: labels: { app: shop } spec: containers: - name: shop image: shop:v2 readinessProbe: # no traffic until this answers httpGet: { path: /health, port: 3000 } periodSeconds: 2
The commands
kubectl set image deployment/shop shop=shop:v2 # start the rolling update kubectl rollout status deployment/shop # watch it kubectl rollout pause deployment/shop # stop half-way to look kubectl rollout resume deployment/shop kubectl rollout undo deployment/shop # roll back (another rolling update)
10

When to use it, and when not

The default for stateless services.

Most web APIs fit a rolling update well.

Decide
Good fitStateless REST or GraphQL APIs with several copies.
Good fitKubernetes workloads and teams without budget for a second environment.
Poor fitVersions that cannot run together - use recreate.
Poor fitYou need rollback in seconds - use blue-green.
Poor fitA risky change that only a few users should meet first - use canary.
11

Common mistakes

And how to avoid each one.

The first one is responsible for most failed rolling updates.

Mistake -> fix
No readiness checkAdd one that checks real dependencies (lab: 62 errors vs 0).
Liveness check used as readinessUse both, for their own jobs.
Breaking API or data changesMake v2 compatible with v1, or use recreate.
Killing old copies without drainingGive them a grace period to finish requests.
Rolling all copies at onceKeep maxUnavailable low enough to keep capacity.
12

Compared with the other patterns

Rolling sits between blue-green and canary.

Each pattern makes a different trade between cost, speed of rollback and risk.

Rolling and its neighbours
Rolling updateCopies replaced one by one. Cheap. Mixed versions. Slow rollback.
Blue-greenTwo environments. Instant rollback. Double cost.
CanaryControlled share of users, decided by metrics.
RecreateStop all, then start all. Downtime, no mixed versions.
13

Interview questions

Short answers you can give in your own words.

What is a rolling update? Replacing the instances of a service gradually - start a new one, wait until it is ready, remove an old one - so the service keeps serving throughout, without a second environment.

What do maxSurge and maxUnavailable control? How many extra instances may exist and how many may be missing during the rollout: the trade-off between speed, extra resources and capacity.

What is the main risk? Old and new versions serve traffic together, so changes must be backward-compatible; and instances must not receive traffic before they are ready.

Remember

  • Replace copies a few at a time: start new, wait for ready, add it, remove an old one.
  • Never send users to a copy before its readiness check passes (lab: 0 errors vs 62).
  • Readiness decides traffic; liveness decides restarts.
  • maxSurge and maxUnavailable trade speed against extra servers and capacity.
  • Old and new versions serve users together - changes must be compatible.
  • Rollback is another rolling update: kubectl rollout undo.

Check yourself

Answer in your head first, then open the answer.

What must happen before a new copy gets traffic?

Its readiness check must pass.

In the lab, how many requests failed without readiness checks?

62. With readiness checks: 0.

What does maxUnavailable: 0 guarantee?

Ready capacity never drops below the normal number of copies.

Why must v2 be compatible with v1 in a rolling update?

Both versions serve users at the same time during the update.

Which command rolls a Kubernetes Deployment back?

kubectl rollout undo deployment/<name>.

Hands-on lab

Build it: replace four copies while users keep working

The same plain Node.js lab kit - no Docker, no Kubernetes - plus a script that does what Kubernetes does: replace four copies of v1 with v2, one at a time. You run it with and without readiness checks and measure the difference. Every output is from a real run (Node 22).

1

Set up the lab kit

Three files, no packages, no Docker.

The lab uses three small files and nothing to install. app.js is the service: run it with VERSION=v1 or v2 and a PORT. It can be told to fail some requests (ERROR_RATE) or to start slowly (STARTUP_MS), and on Ctrl+C or a normal kill it finishes the requests it is working on before it exits.

router.js plays the load balancer (NGINX, an AWS load balancer, a Kubernetes Service). Users only ever talk to it, on port 4000. You change where it sends traffic while it runs, with POST /admin/config. traffic.js sends a steady stream of requests from 20 users and prints, every half second, how many answers came from v1, from v2, and how many failed. It ends with PASS or FAIL - that is how you check your work in every lab.

Terminal
mkdir deploy-lab && cd deploy-lab npm init -y npm pkg set type=module # no packages to install - the lab uses only Node.js itself
app.js - one version of the service
// app.js - one version of the service. // Run: VERSION=v1 PORT=4001 node app.js import http from "node:http"; const VERSION = process.env.VERSION ?? "v1"; const PORT = Number(process.env.PORT ?? 4001); const ERROR_RATE = Number(process.env.ERROR_RATE ?? 0); // 0.2 = 20% of requests fail (a bad release) const STARTUP_MS = Number(process.env.STARTUP_MS ?? 0); // how long the app needs before it can serve const wait = (ms) => new Promise((resolve) => setTimeout(resolve, ms)); let inFlight = 0; const server = http.createServer(async (req, res) => { if (req.url === "/health") { res.writeHead(200); return res.end("ok"); } inFlight++; if (req.url.startsWith("/slow")) await wait(2000); // some requests take 2 seconds inFlight--; const failed = Math.random() < ERROR_RATE; res.writeHead(failed ? 500 : 200, { "content-type": "application/json" }); res.end(JSON.stringify(failed ? { version: VERSION, error: "bug in this release" } : { version: VERSION })); }); // A graceful stop: stop taking new requests, finish the ones in progress, then exit function shutdown() { console.log(VERSION, "stopping - finishing", inFlight, "requests in progress"); server.close(() => process.exit(0)); } process.on("SIGTERM", shutdown); process.on("SIGINT", shutdown); setTimeout(() => server.listen(PORT, () => console.log(VERSION, "ready on port", PORT)), STARTUP_MS);
router.js - a tiny load balancer
// router.js - a tiny load balancer on port 4000. Users only ever talk to this. import { createHash } from "node:crypto"; import http from "node:http"; // Where traffic goes. Change it while running with POST /admin/config. let config = { stable: ["http://localhost:4001"], // the version most users get (one or more copies) canary: [], // the new version, for some users (Canary lesson) canaryPercent: 0, // 0 - 100 sticky: false, // true = the same user always gets the same version }; let next = 0; function pickTarget(req) { const userId = req.headers["x-user-id"] ?? ""; // sticky: turn the user id into a number 0-99 - the same number every time for the same user const bucket = config.sticky ? createHash("md5").update(userId).digest().readUInt32BE(0) % 100 : Math.random() * 100; const pool = config.canary.length && bucket < config.canaryPercent ? config.canary : config.stable; return pool[next++ % pool.length]; // round robin inside the pool } const server = http.createServer((req, res) => { if (req.url === "/admin/config") { if (req.method === "GET") return res.end(JSON.stringify(config)); let body = ""; req.on("data", (chunk) => (body += chunk)); return req.on("end", () => { config = { ...config, ...JSON.parse(body) }; console.log("config:", JSON.stringify(config)); res.end(JSON.stringify(config)); }); } const target = pickTarget(req); const upstream = http.request(target + req.url, { method: req.method, headers: req.headers }, (answer) => { res.writeHead(answer.statusCode, answer.headers); answer.pipe(res); }); upstream.on("error", () => { // the instance is down or vanished mid-request if (!res.headersSent) res.writeHead(502, { "content-type": "application/json" }); res.end(JSON.stringify({ error: "upstream unavailable", target })); }); req.pipe(upstream); }); server.listen(4000, () => console.log("router on http://localhost:4000"));
traffic.js - steady traffic, like real users
// traffic.js - steady traffic through the router, like real users. // Run: node traffic.js <seconds> <requests per second> [share of slow requests] const seconds = Number(process.argv[2] ?? 5); const perSecond = Number(process.argv[3] ?? 50); const slowShare = Number(process.argv[4] ?? 0); const buckets = []; // one row per half second const usersSeen = new Map(); // user -> set of versions that user saw const started = Date.now(); const pending = []; const timer = setInterval(() => { const user = "user-" + Math.floor(Math.random() * 20); const path = Math.random() < slowShare ? "/slow" : "/"; const sentAt = Date.now(); pending.push( fetch("http://localhost:4000" + path, { headers: { "x-user-id": user } }) .then(async (r) => ({ status: r.status, body: await r.json() })) .catch(() => ({ status: 0, body: {} })) .then(({ status, body }) => { const slot = Math.floor((sentAt - started) / 500); const row = (buckets[slot] ??= { v1: 0, v2: 0, errors: 0 }); if (status === 200) { row[body.version] = (row[body.version] ?? 0) + 1; if (!usersSeen.has(user)) usersSeen.set(user, new Set()); usersSeen.get(user).add(body.version); } else row.errors++; }), ); }, 1000 / perSecond); setTimeout(async () => { clearInterval(timer); await Promise.all(pending); let total = { v1: 0, v2: 0, errors: 0 }; buckets.forEach((row, i) => { if (!row) return; console.log(`${(i / 2).toFixed(1).padStart(4)}s v1 ${String(row.v1).padStart(3)} v2 ${String(row.v2).padStart(3)} errors ${row.errors}`); total = { v1: total.v1 + row.v1, v2: total.v2 + row.v2, errors: total.errors + row.errors }; }); const mixed = [...usersSeen.values()].filter((versions) => versions.size > 1).length; console.log("total:", JSON.stringify(total), "| users who saw both versions:", mixed, "of", usersSeen.size); console.log(total.errors === 0 ? "PASS - no request failed" : `FAIL - ${total.errors} requests failed`); process.exit(total.errors === 0 ? 0 : 1); }, seconds * 1000);
2

The rolling update script

Start a new copy, wait for it, swap it in, stop an old one gracefully.

rolling.js plays Kubernetes. It starts four copies of v1 on ports 4011 to 4014 and gives them to the router. Then, for each copy: it starts a v2 copy on a new port (v2 needs 1.5 seconds to start), waits for its /health check, swaps it into the router in place of one old copy, and stops the old copy gracefully. With --no-ready it skips the wait, so you can see why the wait matters.

rolling.js
// rolling.js - replace 4 copies of v1 with v2, one at a time, while users keep using the app. // Run with the router on 4000: node rolling.js (waits for each new copy to be ready) // node rolling.js --no-ready (does not wait - see what breaks) import { spawn } from "node:child_process"; const WAIT_FOR_READY = process.argv[2] !== "--no-ready"; const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms)); const started = Date.now(); const log = (text) => console.log(`${((Date.now() - started) / 1000).toFixed(1).padStart(5)}s ${text}`); function startCopy(version, port) { // v2 needs 1.5 s to start - like a real app loading config and opening connections const env = { ...process.env, VERSION: version, PORT: String(port), STARTUP_MS: version === "v2" ? "1500" : "0" }; return { version, port, process: spawn("node", ["app.js"], { env, stdio: "ignore" }) }; } async function waitUntilReady(port) { for (;;) { try { if ((await fetch(`http://localhost:${port}/health`)).ok) return; } catch {} await sleep(100); } } const usePool = (copies) => fetch("http://localhost:4000/admin/config", { method: "POST", body: JSON.stringify({ stable: copies.map((c) => `http://localhost:${c.port}`) }), }); // 1. Four copies of v1 serve the users const copies = [4011, 4012, 4013, 4014].map((port) => startCopy("v1", port)); await Promise.all(copies.map((c) => waitUntilReady(c.port))); await usePool(copies); log("serving: " + copies.map((c) => c.version).join(" ")); await sleep(2000); // 2. Replace them one at a time: start a new copy, add it, remove an old one for (let i = 0; i < copies.length; i++) { const fresh = startCopy("v2", 4021 + i); if (WAIT_FOR_READY) await waitUntilReady(fresh.port); // the readiness check const old = copies[i]; copies[i] = fresh; await usePool(copies); // new copy in, old copy out old.process.kill("SIGTERM"); // let it finish its requests, then stop log("serving: " + copies.map((c) => c.version).join(" ")); if (!WAIT_FOR_READY) await sleep(1600); } log("done - v2 serves everyone. Press Ctrl+C to stop the copies.");
3

Roll with readiness checks

Four copies replaced, 0 errors.

Start the router, then rolling.js, and send traffic while it works. The script's log shows each copy changing; the traffic timeline shows v2's share growing in four steps - and no failed request.

Terminals
node router.js # terminal 1 node rolling.js # terminal 2 node traffic.js 11 40 # terminal 3, right after starting rolling.js
Output - measured (selected lines)
rolling.js 0.1s serving: v1 v1 v1 v1 3.8s serving: v2 v1 v1 v1 5.5s serving: v2 v2 v1 v1 7.1s serving: v2 v2 v2 v1 8.8s serving: v2 v2 v2 v2 traffic.js 2.0s v1 19 v2 0 errors 0 3.0s v1 14 v2 5 errors 0 5.0s v1 9 v2 10 errors 0 6.5s v1 5 v2 15 errors 0 7.5s v1 0 v2 19 errors 0 total: {"v1":195,"v2":227,"errors":0} | users who saw both versions: 20 of 20 PASS - no request failed
4

Roll without readiness checks

The same update, sending users to copies that are still starting.

Run it again with --no-ready. Each new copy joins the router the moment it is started, while it still needs 1.5 seconds before it can answer. During each of those windows, about one request in four went to a copy that could not answer. 62 requests failed.

Terminal 2
node rolling.js --no-ready
Output - measured (selected lines)
1.0s v1 14 v2 0 errors 5 2.5s v1 10 v2 4 errors 6 4.0s v1 7 v2 8 errors 5 5.5s v1 3 v2 12 errors 4 7.0s v1 0 v2 16 errors 3 7.5s v1 0 v2 19 errors 0 total: {"v1":128,"v2":233,"errors":62} | users who saw both versions: 20 of 20 FAIL - 62 requests failed
Rolling update of 4 copies, measured
Wait for /health0 errors. About 7 seconds for 4 copies.
Do not wait62 errors - one copy in four could not answer while starting.
Mixed versionsBoth runs: all 20 users saw v1 and v2 during the update.

Practice on your own

  1. 1.

    Change rolling.js to replace two copies at a time. How long does the update take now, and how many copies serve users at the worst moment?

    Hint

    Start two v2 copies, wait for both, then swap two old ones.

  2. 2.

    Make v2 broken: start it with ERROR_RATE=0.5. Add a check to rolling.js that sends 20 test requests to each new copy and stops the rollout (keeping the rest on v1) if more than 2 fail.

    Hint

    This is a small canary check inside a rolling update - what Argo Rollouts adds to Kubernetes.

  3. 3.

    Make app.js v2 answer { version, amount } instead of { version }, and write a client that reads body.version and body.amount. What does the client see during the rolling update?

    Hint

    Some answers come from v1 and do not have amount.

  4. 4.

    Write the Kubernetes Deployment settings for this lab: 4 replicas, maxSurge 1, maxUnavailable 0, and a readinessProbe on /health.

    Hint

    The reference section above shows the YAML shape.

Comments

Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.

Loading comments...