Recreate Deployment
Stop old version first, then start the new version.
Recreate: planned downtime, no mixed versions
Stop the old version, then start the new one. Simple, with planned downtime - and the right choice when two versions cannot run together. This lesson explains when and how - then you measure it against a rolling update.
The idea in short
Stop the old version completely. Then start the new one.
Recreate is the simplest deployment: stop every copy of the old version, then start the new version. Between the two, nothing serves users - planned downtime. In return, old and new versions never run at the same time.
Remember it as: stop all, start all. Kubernetes supports it directly with strategy type Recreate.
DowntimeYes - from the last old copy stopping to the first new copy being ready.Mixed versionsNever.CostNo extra servers.ComplexityThe lowest of all patterns.Best forIncompatible changes, internal tools, maintenance windows.An everyday picture: renovating a shop
Close for the afternoon, with a sign on the door.
A shop that rebuilds its shelves and floor closes for an afternoon, with a sign: "Closed for renovation, back at 5 pm". Rebuilding around customers would be slower and dangerous. And a sign on the door is much better than customers finding a locked door with no explanation.
Words you need
Four words used in this lesson.
These words come up whenever you plan downtime.
DowntimeTime when the service cannot serve users.Maintenance windowA planned, announced time for downtime, usually when traffic is low.Incompatible changeA change the old version cannot work with, such as a new data format.Maintenance pageA small answer shown during downtime: HTTP 503, a message, and Retry-After.How it works, step by step
Use Next to walk through a real run.
Step through the diagram. It is a real run of the lab: v2 changes the format of shared data when it starts, so v1 and v2 cannot run together.
1 - v1 serves; the data is in the old format
Four copies of v1 serve users. They all read data.json in its old format, {"price": 499}.
A real run: v2 changes the format of shared data, so v1 and v2 cannot run at the same time.
1. PrepareGet everything ready for v2 while v1 still serves (images, configuration).2. Show maintenanceOptional but kind: point users to a maintenance page.3. Stop allStop every v1 copy - let each finish its requests.4. Start allStart the v2 copies; run any data migration.5. OpenWhen v2 is ready, send users to it.Why teams use it - and what it costs
Simplicity and safety from mixed versions, paid for with downtime.
Every other pattern runs old and new versions together, at least for a moment. That is only safe when they are compatible. Recreate never mixes them. It also needs no extra servers and almost no tooling.
The cost is downtime: in the lab, about 1.6 seconds and 65 failed requests. With a slow start-up or a long data migration, it can be minutes - and users notice.
BenefitNo mixed versions - no compatibility problems.BenefitNo extra infrastructure.BenefitVery simple to understand and to run.CostPlanned downtime.CostThe downtime grows with start-up and migration time.CostRollback means another recreate - more downtime.Detail 1: when versions cannot run together
The real reason to choose recreate.
The lab's v2 converts shared data to a new format when it starts. First we tried a rolling update: as soon as the first v2 copy converted the data, the remaining v1 copies could not read it - 156 requests failed with "unknown data format". With recreate, 65 requests failed during the short gap, and after that, none.
Other cases: a job that must never run twice at the same time, a licence that allows only one running copy, a big framework or runtime upgrade where old and new cannot share data. If you can make the change compatible instead (expand-contract), you usually should - then other patterns become possible.
Detail 2: make the downtime short and polite
A maintenance page instead of a connection error.
During the gap, users get whatever the load balancer returns when nothing is behind it - in the lab, 502 "upstream unavailable". Pointing the load balancer at a small maintenance page first is much better: HTTP 503, a Retry-After header, and a clear message. In the lab, the same 65 requests all got "We are updating the site. Please try again in a few seconds."
To keep the gap short: make start-up fast, prepare everything before stopping v1, run long migrations ahead of time if possible, choose a low-traffic time, and tell users in advance.
In the real world
Kubernetes and a maintenance switch.
In Kubernetes, setting the strategy to Recreate makes the Deployment terminate all old pods before creating any new ones. For the maintenance page, many teams switch the load balancer or ingress to a static page during the release. These examples show the shape; they were not run in this lesson's lab.
apiVersion: apps/v1
kind: Deployment
metadata:
name: reports
spec:
replicas: 4
strategy:
type: Recreate # stop all old pods, then start new ones
selector:
matchLabels: { app: reports }
template:
metadata:
labels: { app: reports }
spec:
containers:
- name: reports
image: reports:v2server {
listen 80;
location / {
if (-f /etc/nginx/maintenance.on) {
return 503;
}
proxy_pass http://app;
}
error_page 503 /maintenance.html;
location = /maintenance.html {
root /usr/share/nginx/html;
add_header Retry-After 300 always;
}
}When to use it, and when not
Fine for many systems; wrong for anything always-on.
Ask one question: can our users accept a short, planned outage?
Good fitDevelopment, test and staging environments.Good fitIncompatible data or schema changes that cannot be split.Good fitInternal tools, batch systems, planned maintenance windows.Poor fitCheckout, login, public APIs - anything users need every minute.Poor fitServices with an uptime promise (SLA) of 99.9% or more.Common mistakes
And how to avoid each one.
Recreate is simple, but the downtime must be managed.
No maintenance pageShow 503 with a message and Retry-After.Preparing after stopping v1Do everything possible before the stop.A long migration inside the downtimeRun it ahead of time, or in smaller steps.No announcementTell users when, and for how long.Choosing recreate for a compatible changeUse rolling or blue-green and avoid the downtime.Compared with the other patterns
The only pattern that guarantees one version at a time.
Recreate trades availability for simplicity.
RecreateDowntime, never mixed versions, no extra cost.Rolling updateNo downtime, but mixed versions.Blue-greenNo downtime and no mixing for users, but double infrastructure and a shared database.CanaryNo downtime, mixed versions on purpose.Interview questions
Short answers you can give in your own words.
What is the recreate strategy? Terminating all instances of the old version before starting the new one, accepting downtime so that two versions never run together.
When would you choose it? When versions are incompatible - a data or schema change the old version cannot handle, singleton workloads - and some downtime is acceptable.
How do you reduce its impact? Serve a maintenance response with Retry-After, keep start-up fast, prepare everything before stopping the old version, and schedule it in a low-traffic window.
Remember
- Stop every old copy, then start the new version - planned downtime.
- Old and new versions never run together.
- Choose it when versions cannot run together (lab: rolling gave 156 errors, recreate 65 in a short gap).
- Show a maintenance page (503 + Retry-After) during the gap.
- Shorten the gap: prepare first, start fast, migrate ahead of time.
- If you can make the change compatible, prefer a pattern without downtime.
Check yourself
Answer in your head first, then open the answer.
What happens between stopping v1 and v2 being ready?
Nothing serves users - that is the planned downtime.
Why did the rolling update fail in the lab?
v2 converted the shared data, and the remaining v1 copies could not read it: 156 errors.
What should users see during the gap?
A maintenance page: HTTP 503, a clear message, and a Retry-After header.
Which Kubernetes setting gives this behaviour?
strategy: type: Recreate.
Name a system where recreate is a poor choice.
Checkout, login, or any public API that users need every minute.
Build it: release a v2 that changes the data format
A service whose v2 converts shared data to a new format, released three ways - rolling, recreate, and recreate with a maintenance page - in plain Node.js with no Docker. Every output is from a real run (Node 22).
Set up the project
Nothing to install - only Node.js.
This lab reuses router.js from the deployment lab kit (the blue-green and canary lessons show it in full), and adds four small files: a service that reads shared data, a maintenance page, a script that performs the release, and a client that shows what users got.
mkdir deploy-lab && cd deploy-lab
npm init -y
npm pkg set type=module
# no packages to install - only Node.js itself// router.js - a tiny load balancer on port 4000. Users only ever talk to this.
import { createHash } from "node:crypto";
import http from "node:http";
// Where traffic goes. Change it while running with POST /admin/config.
let config = {
stable: ["http://localhost:4001"], // the version most users get (one or more copies)
canary: [], // the new version, for some users (Canary lesson)
canaryPercent: 0, // 0 - 100
sticky: false, // true = the same user always gets the same version
};
let next = 0;
function pickTarget(req) {
const userId = req.headers["x-user-id"] ?? "";
// sticky: turn the user id into a number 0-99 - the same number every time for the same user
const bucket = config.sticky
? createHash("md5").update(userId).digest().readUInt32BE(0) % 100
: Math.random() * 100;
const pool = config.canary.length && bucket < config.canaryPercent ? config.canary : config.stable;
return pool[next++ % pool.length]; // round robin inside the pool
}
const server = http.createServer((req, res) => {
if (req.url === "/admin/config") {
if (req.method === "GET") return res.end(JSON.stringify(config));
let body = "";
req.on("data", (chunk) => (body += chunk));
return req.on("end", () => {
config = { ...config, ...JSON.parse(body) };
console.log("config:", JSON.stringify(config));
res.end(JSON.stringify(config));
});
}
const target = pickTarget(req);
const upstream = http.request(target + req.url, { method: req.method, headers: req.headers }, (answer) => {
res.writeHead(answer.statusCode, answer.headers);
answer.pipe(res);
});
upstream.on("error", () => { // the instance is down or vanished mid-request
if (!res.headersSent) res.writeHead(502, { "content-type": "application/json" });
res.end(JSON.stringify({ error: "upstream unavailable", target }));
});
req.pipe(upstream);
});
server.listen(4000, () => console.log("router on http://localhost:4000"));A service whose v2 changes the data format
v1 cannot read what v2 writes.
store-app.js answers with a course price that it reads from data.json. v1 understands {"price": 499}. v2, when it starts, converts the file to {"price": {"amount": 49900, "currency": "INR"}}. From that moment, v1 answers 500 "unknown data format". maintenance.js answers 503 with a Retry-After header.
// store-app.js - answers with the course price, which it reads from a shared file, data.json.
// v1 understands {"price": 499}
// v2 changes it to {"price": {"amount": 49900, "currency": "INR"}} when it starts
// Run: VERSION=v1 PORT=4011 node store-app.js
import fs from "node:fs";
import http from "node:http";
const VERSION = process.env.VERSION ?? "v1";
const PORT = Number(process.env.PORT ?? 4011);
const STARTUP_MS = Number(process.env.STARTUP_MS ?? 0);
const read = () => JSON.parse(fs.readFileSync("data.json", "utf8"));
if (VERSION === "v2") { // v2's data migration, on start-up
const data = read();
if (typeof data.price === "number") {
fs.writeFileSync("data.json", JSON.stringify({ price: { amount: data.price * 100, currency: "INR" } }));
}
}
const server = http.createServer((req, res) => {
const data = read();
res.setHeader("content-type", "application/json");
if (VERSION === "v1" && typeof data.price !== "number") {
res.writeHead(500); // v1 cannot read the new format
return res.end(JSON.stringify({ version: VERSION, error: "unknown data format" }));
}
const price = VERSION === "v1" ? data.price : data.price.amount / 100;
res.end(JSON.stringify({ version: VERSION, price }));
});
function shutdown() {
server.close(() => process.exit(0));
}
process.on("SIGTERM", shutdown);
process.on("SIGINT", shutdown);
setTimeout(() => server.listen(PORT, () => console.log(VERSION, "ready on port", PORT)), STARTUP_MS);// maintenance.js - port 4099. A polite "we are updating" answer while nothing else can serve.
import http from "node:http";
http
.createServer((req, res) => {
res.writeHead(503, { "content-type": "application/json", "retry-after": "5" });
res.end(JSON.stringify({ error: "We are updating the site. Please try again in a few seconds." }));
})
.listen(4099, () => console.log("maintenance page on port 4099"));The release script and the user's view
Recreate, recreate with a maintenance page, or rolling - for comparison.
recreate.js writes the old data format, starts four v1 copies and routes users to them. Then it releases v2: by default it stops all v1 copies (each finishes its requests) and then starts four v2 copies; with --maintenance it first routes users to the maintenance page; with --rolling it does a rolling update instead. downtime.js sends 40 requests per second and shows, every half second, what users got.
// recreate.js - stop ALL copies of v1, then start the copies of v2.
// Run with the router on 4000: node recreate.js (plain recreate)
// node recreate.js --maintenance (show a maintenance page meanwhile)
// node recreate.js --rolling (for comparison: a rolling update)
import fs from "node:fs";
import { spawn } from "node:child_process";
const mode = process.argv[2] ?? "";
const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
const started = Date.now();
const log = (text) => console.log(`${((Date.now() - started) / 1000).toFixed(1).padStart(5)}s ${text}`);
const startCopy = (version, port) => ({
port,
process: spawn("node", ["store-app.js"], {
env: { ...process.env, VERSION: version, PORT: String(port), STARTUP_MS: version === "v2" ? "1500" : "0" },
stdio: "ignore",
}),
});
async function waitUntilReady(port) {
for (;;) {
try {
await fetch(`http://localhost:${port}/`);
return;
} catch {}
await sleep(100);
}
}
const stopped = (copy) => new Promise((resolve) => { copy.process.on("exit", resolve); copy.process.kill("SIGTERM"); });
const route = (ports) =>
fetch("http://localhost:4000/admin/config", {
method: "POST",
body: JSON.stringify({ stable: ports.map((p) => `http://localhost:${p}`) }),
});
fs.writeFileSync("data.json", JSON.stringify({ price: 499 })); // the old format
const copies = [4011, 4012, 4013, 4014].map((port) => startCopy("v1", port));
await Promise.all(copies.map((c) => waitUntilReady(c.port)));
await route(copies.map((c) => c.port));
log("4 x v1 serving");
await sleep(2000);
if (mode === "--rolling") {
for (let i = 0; i < 4; i++) {
const fresh = startCopy("v2", 4021 + i);
await waitUntilReady(fresh.port);
const old = copies[i];
copies[i] = fresh;
await route(copies.map((c) => c.port));
await stopped(old);
log(`copy ${i + 1} is v2`);
}
} else {
if (mode === "--maintenance") await route([4099]); // users see the maintenance page
await Promise.all(copies.map(stopped)); // 1. stop everything
log("all v1 copies stopped");
const fresh = [4021, 4022, 4023, 4024].map((port) => startCopy("v2", port)); // 2. start v2
await Promise.all(fresh.map((c) => waitUntilReady(c.port)));
await route(fresh.map((c) => c.port));
log("4 x v2 serving");
}
log("done. Press Ctrl+C to stop.");// downtime.js - steady traffic for <seconds>; shows what users got every half second
const seconds = Number(process.argv[2] ?? 8);
const rows = [];
const started = Date.now();
const pending = [];
const label = { 200: "ok", 500: "app error", 502: "nothing listening", 503: "maintenance page" };
const timer = setInterval(() => {
const slot = Math.floor((Date.now() - started) / 500);
pending.push(
fetch("http://localhost:4000/")
.then((r) => r.status)
.catch(() => 0)
.then((status) => {
const row = (rows[slot] ??= {});
row[status] = (row[status] ?? 0) + 1;
}),
);
}, 25);
setTimeout(async () => {
clearInterval(timer);
await Promise.all(pending);
const total = {};
rows.forEach((row, i) => {
const parts = Object.entries(row ?? {}).map(([s, n]) => `${label[s] ?? s} ${n}`);
console.log(`${(i / 2).toFixed(1).padStart(4)}s ${parts.join(", ")}`);
for (const [s, n] of Object.entries(row ?? {})) total[label[s] ?? s] = (total[label[s] ?? s] ?? 0) + n;
});
const without = rows.filter((row) => row && !row[200]).length / 2;
console.log("total:", JSON.stringify(total), `| time with no successful answer: ${without} s`);
}, seconds * 1000);Try a rolling update first
156 failed requests - the versions cannot run together.
Run the rolling update. As soon as the first v2 copy started and converted the data, the remaining v1 copies could not read it. For about six seconds, a large share of requests failed with "unknown data format" - 156 in total.
node router.js # terminal 1
node maintenance.js # terminal 2
node recreate.js --rolling # terminal 3
node downtime.js 8 # terminal 4, right after 0.5s ok 19
1.0s ok 9, app error 10
1.5s app error 20
3.0s ok 5, app error 14
5.0s ok 10, app error 9
7.5s ok 17, app error 2
total: {"ok":151,"app error":156}Now recreate - plain, and with a maintenance page
65 failed requests in a short gap; then a polite version of the same gap.
Restart the router and run node recreate.js. All v1 copies stopped at 2.1 s and v2 served at 3.8 s. Users got "nothing listening" (502) for 65 requests - about 1.6 seconds at 40 requests per second - and then every request worked, with no format errors.
Run node recreate.js --maintenance. The gap is the same 65 requests, but every one of them got the maintenance page: HTTP 503, Retry-After: 5, and a clear message.
2.1s all v1 copies stopped
3.8s 4 x v2 serving
1.0s ok 5, nothing listening 14
1.5s nothing listening 20
2.0s nothing listening 19
2.5s ok 7, nothing listening 12
3.0s ok 19
total: {"ok":242,"nothing listening":65} 1.0s ok 5, maintenance page 14
1.5s maintenance page 19
2.0s maintenance page 19
2.5s ok 7, maintenance page 13
3.0s ok 19
total: {"ok":242,"maintenance page":65}Rolling update156 app errors over several seconds.Recreate65 failed requests (~1.6 s), then 0 errors.Recreate + maintenance65 clear 503 answers, then 0 errors.Practice on your own
- 1.
Make v2 start slower (STARTUP_MS 4000 in recreate.js). How many requests fail now? What does that tell you about start-up time?
Hint
The gap lasts until v2 is ready.
- 2.
Change recreate.js to start the v2 copies BEFORE stopping v1, but only send traffic to them after v1 is stopped. Does that shorten the gap? What goes wrong?
Hint
v2 migrates the data when it starts - while v1 is still serving.
- 3.
Redesign the data change with expand-contract so that a rolling update works: v2 writes both formats and v1 keeps working.
Hint
Keep price as a number and add priceAmount and currency next to it.
- 4.
Add a /status page to maintenance.js that says when the update started, and show it in the 503 message.
Hint
Record Date.now() when maintenance.js starts.
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...
AI
System Design
Backend
- GraphQL8 modules · 69 lessons planned
- Core Python13 modules · 75 lessons planned
- FastAPI5 sections · 20 lessons
- Node.js14 modules · 206 lessons planned
- Node.js Performance7 chapters · 36 topics
- Event Loop Lifecycle6 phases · 3 scenarios
- Docker & Containerization11 modules · 144 lessons planned
- AWS for Developers14 modules · 219 lessons planned
- CI/CD & DevOps Automation10 modules · 134 lessons planned