← Back to architecture map
Architecture Pattern

Blue-Green Deploy

Two identical environments. Switch traffic instantly for zero-downtime releases.

deploy
Lesson

Blue-green: switch everyone, keep the old one ready

Two identical environments and one switch: the safest way to release with no downtime, and the fastest way back. This lesson explains how it works, the three details that break it in real life, and how to set it up - then you build and measure it.

01

The idea in short

Two copies of production. Switch everyone at once. Keep the old one for a fast undo.

Blue-green deployment means you run two identical copies of your production system. One copy, called Blue, serves all users. You put the new version on the other copy, called Green, and test it there. When Green is ready, the load balancer sends every user to Green in one move. Blue keeps running for a while, so if something is wrong you can send everyone back to Blue just as fast.

Remember it as three words: prepare, switch, keep. Prepare Green, switch all traffic, keep Blue ready.

At a glance
DowntimeNone - the switch is one routing change.
RollbackSeconds - switch back to Blue.
CostTwo full copies while you release.
RiskEveryone moves at once.
Best forPayments, login, critical APIs.
02

An everyday picture: two stages in a theatre

Build the next scene behind the curtain, then move the spotlight.

Some theatres have two stages. While the audience watches stage A, the crew builds the next scene on stage B behind a curtain and checks every light. When the scene changes, the spotlight simply moves to stage B. The audience never sees an empty stage. If stage B has a problem, the spotlight goes back to stage A, which is still standing.

The spotlight is your load balancer. The two stages are Blue and Green. Building behind the curtain is testing Green before any user sees it.

03

Words you need

Five words used in this lesson.

You will see these words again in every deployment lesson.

Small dictionary
Load balancerThe component that receives every user request and decides which server answers it - NGINX, an AWS load balancer, a Kubernetes Service.
EnvironmentOne complete copy of your system: servers, configuration, connections.
Switch (cut-over)Changing the load balancer so traffic goes to the other environment.
RollbackGoing back to the previous version.
DrainingLetting an old server finish the requests it is working on before you stop it.
04

How it works, step by step

Four steps. Use Next to walk through a real run.

Step through the diagram. Every number in it was measured in the lab at the end of this page: 40 requests per second, some of them slow, while the router switched from Blue to Green.

request flowBlue-green: switch everyone in one movestep 1 / 4

1 - Blue is live; Green is deployed beside it

All users are on Blue. Green, the new version, is already running on its own port. You test it directly - curl localhost:4002 answered v2 - while no user can reach it yet.

users on
Blue (v1)
Green
running, tested privately
errors
0
cost
two full setups

A real run: 40 requests per second, some taking 2 seconds. The router switches from Blue to Green.

The four steps
1. PrepareDeploy v2 to Green. Blue still serves everyone.
2. TestTest Green directly, on its own address. No user can reach it yet.
3. SwitchPoint the load balancer at Green. Everyone moves at once.
4. KeepLeave Blue running until you are sure. Then drain it and stop it.
05

Why teams use it - and what it costs

Zero downtime and the fastest undo, for the price of a second copy.

You get three things. No downtime: the switch is a routing change, not a restart - the lab switched 40 requests per second with 0 errors. A full test before users arrive: Green runs on real servers with the real configuration. And the fastest rollback of any pattern: the same one change, backwards - 0 errors in the lab.

You pay for it with capacity: two full environments while the release is in progress. And because everyone moves at once, any bug your tests missed reaches every user at once - until you switch back.

Benefits and costs
BenefitNo downtime.
BenefitRollback in seconds, as long as Blue is still running.
BenefitGreen is tested on real infrastructure first.
CostDouble infrastructure during the release.
CostAll users meet a missed bug at the same moment.
CostDatabase changes must work for both versions.
06

Detail 1: drain Blue before you stop it

The switch only moves NEW requests.

When you switch, requests that started on Blue a moment earlier are still running there. In the lab, 26 requests were still in progress on Blue at the moment of the switch.

Stop Blue immediately and those requests break: we measured 29 errors. Stop it gracefully - stop accepting new work, finish the current work, then exit - and nothing breaks: 0 errors. Load balancers call this connection draining (AWS: deregistration delay). Kubernetes gives each pod a grace period before it is killed, for the same reason.

Watch out: The most common blue-green mistake: shutting down Blue the moment you switch. Drain it, and keep it until you are sure you will not need to roll back.

07

Detail 2: the database must work for both versions

Expand first, contract later.

Blue and Green usually share one database. If v2 renames a column, Blue breaks the moment the database changes - and your instant rollback breaks with it. So you change the database in small steps that keep both versions working. This is called expand-contract.

Expand: add the new column and keep the old one; v2 writes both. Migrate: copy the old data into the new column. Contract: in a LATER release, when no running version reads the old column, remove it.

Renaming price to amount, safely
Release 1 - expandAdd column amount. v2 writes price AND amount. v1 still works.
Background jobCopy price into amount for old rows.
Release 2Code reads amount only, but still writes both - rollback still works.
Release 3 - contractRemove price. No running version uses it any more.
08

Detail 3: sessions and other state

Anything kept inside Blue's memory is lost at the switch.

If logged-in sessions, carts or caches live in Blue's memory, users lose them when they move to Green - they are logged out, or their cart is empty. Keep this state outside the servers, in a shared store such as Redis or the database, so both environments can read it.

The same applies to background jobs and scheduled tasks: make sure only one environment runs them, or they run twice.

09

In the real world

The same switch with NGINX, Kubernetes, and managed tools.

With NGINX, Blue and Green are two upstreams and the switch is one line plus a reload; nginx -s reload starts new worker processes with the new configuration and lets the old ones finish their requests, so it drains for you. In Kubernetes, the switch is usually the selector of a Service: change version: blue to version: green and the Service sends traffic to the other set of pods. Cloud tools such as AWS CodeDeploy (blue/green for EC2 and ECS) and Argo Rollouts (blueGreen strategy) automate the test, switch and wait.

These configuration examples show the shape. They were not run in this lesson's lab, which uses a small Node.js router instead.

NGINX - switch by changing one line, then: nginx -s reload
upstream backend { server 127.0.0.1:3001; # Blue (live) # server 127.0.0.1:3002; # Green - swap the comment to switch } server { listen 80; location / { proxy_pass http://backend; } }
Kubernetes - the Service selector decides which pods get traffic
apiVersion: v1 kind: Service metadata: name: shop spec: selector: app: shop version: blue # change to green to switch everyone ports: - port: 80 targetPort: 3000
Switch with one command (Kubernetes)
kubectl patch service shop -p '{"spec":{"selector":{"app":"shop","version":"green"}}}' # roll back kubectl patch service shop -p '{"spec":{"selector":{"app":"shop","version":"blue"}}}'
10

When to use it, and when not

Critical services with a budget for two copies.

Use blue-green when downtime or a slow rollback is expensive, and you can afford a second environment for the length of a release.

Decide
Good fitPayments, login and identity, public APIs, regulated systems.
Good fitReleases that must happen at an exact time.
Poor fitInfrastructure that is very expensive to double.
Poor fitDatabase changes that cannot be made backward-compatible.
Poor fitYou want only a few users to try the risky change first - use canary.
11

Common mistakes

And how to avoid each one.

Each of these breaks the two promises of blue-green: no downtime and an instant rollback.

Mistake -> fix
Stopping Blue right after the switchDrain it first, and keep it until you are sure (lab: 29 errors vs 0).
A database change that breaks BlueUse expand-contract: add first, remove in a later release.
Sessions in server memoryKeep them in Redis or the database.
Switching without testing GreenTest Green on its own address before the switch.
Both environments running the same scheduled jobsRun jobs in one environment only.
12

Compared with the other patterns

Where blue-green sits.

Blue-green moves everyone at once after testing. The others move users more gradually, or not at all.

Blue-green and its neighbours
Blue-greenEveryone at once. Rollback in seconds. Double infrastructure.
CanaryA few users first, then more. Needs good metrics.
Rolling updateCopies replaced one by one. No extra environment, slower rollback.
RecreateStop old, start new. Downtime, but very simple.
13

Interview questions

Short answers you can give in your own words.

What is blue-green deployment? Running two identical production environments and switching all traffic from the live one to the one with the new version in a single routing change, keeping the old one ready for an instant rollback.

What are the main risks? Doubling infrastructure; database changes that break the old version (solved with expand-contract); lost in-memory sessions; and in-flight requests broken by stopping the old environment too early (solved with draining).

How is it different from canary? Blue-green moves everyone at once after testing. Canary moves a small share of real users first and widens step by step based on metrics.

Remember

  • Two identical environments: Blue serves users, Green gets the new version.
  • Test Green privately, then switch everyone with one routing change.
  • Rollback is the same change backwards - only while Blue is still running.
  • Drain Blue before stopping it: 29 errors when killed at once, 0 when drained.
  • Database changes must work for both versions: expand first, contract later.
  • Keep sessions and other state outside the servers.

Check yourself

Answer in your head first, then open the answer.

What is the one action that moves users from Blue to Green?

Changing the load balancer (or Kubernetes Service) so it sends traffic to Green.

Why keep Blue running after the switch?

So you can roll back instantly by switching back.

In the lab, what happened when Blue was killed right after the switch?

29 requests that were still running on Blue failed. With a graceful stop: 0.

Why can renaming a database column break blue-green?

Blue and Green share the database; Blue still expects the old column. Use expand-contract.

Blue-green or canary: which one exposes a missed bug to every user?

Blue-green - everyone moves at once.

Hands-on lab

Build it: switch, roll back, and drain

A tiny load balancer, two versions of a service and a traffic generator, all in plain Node.js - no Docker, nothing to install. You switch users from v1 to v2 while traffic flows, roll back, and measure what stopping the old version too early costs. Every output is from a real run (Node 22); your timings will differ a little.

1

Set up the lab kit

Three files, no packages, no Docker.

The lab uses three small files and nothing to install. app.js is the service: run it with VERSION=v1 or v2 and a PORT. It can be told to fail some requests (ERROR_RATE) or to start slowly (STARTUP_MS), and on Ctrl+C or a normal kill it finishes the requests it is working on before it exits.

router.js plays the load balancer (NGINX, an AWS load balancer, a Kubernetes Service). Users only ever talk to it, on port 4000. You change where it sends traffic while it runs, with POST /admin/config. traffic.js sends a steady stream of requests from 20 users and prints, every half second, how many answers came from v1, from v2, and how many failed. It ends with PASS or FAIL - that is how you check your work in every lab.

Terminal
mkdir deploy-lab && cd deploy-lab npm init -y npm pkg set type=module # no packages to install - the lab uses only Node.js itself
app.js - one version of the service
// app.js - one version of the service. // Run: VERSION=v1 PORT=4001 node app.js import http from "node:http"; const VERSION = process.env.VERSION ?? "v1"; const PORT = Number(process.env.PORT ?? 4001); const ERROR_RATE = Number(process.env.ERROR_RATE ?? 0); // 0.2 = 20% of requests fail (a bad release) const STARTUP_MS = Number(process.env.STARTUP_MS ?? 0); // how long the app needs before it can serve const wait = (ms) => new Promise((resolve) => setTimeout(resolve, ms)); let inFlight = 0; const server = http.createServer(async (req, res) => { if (req.url === "/health") { res.writeHead(200); return res.end("ok"); } inFlight++; if (req.url.startsWith("/slow")) await wait(2000); // some requests take 2 seconds inFlight--; const failed = Math.random() < ERROR_RATE; res.writeHead(failed ? 500 : 200, { "content-type": "application/json" }); res.end(JSON.stringify(failed ? { version: VERSION, error: "bug in this release" } : { version: VERSION })); }); // A graceful stop: stop taking new requests, finish the ones in progress, then exit function shutdown() { console.log(VERSION, "stopping - finishing", inFlight, "requests in progress"); server.close(() => process.exit(0)); } process.on("SIGTERM", shutdown); process.on("SIGINT", shutdown); setTimeout(() => server.listen(PORT, () => console.log(VERSION, "ready on port", PORT)), STARTUP_MS);
router.js - a tiny load balancer
// router.js - a tiny load balancer on port 4000. Users only ever talk to this. import { createHash } from "node:crypto"; import http from "node:http"; // Where traffic goes. Change it while running with POST /admin/config. let config = { stable: ["http://localhost:4001"], // the version most users get (one or more copies) canary: [], // the new version, for some users (Canary lesson) canaryPercent: 0, // 0 - 100 sticky: false, // true = the same user always gets the same version }; let next = 0; function pickTarget(req) { const userId = req.headers["x-user-id"] ?? ""; // sticky: turn the user id into a number 0-99 - the same number every time for the same user const bucket = config.sticky ? createHash("md5").update(userId).digest().readUInt32BE(0) % 100 : Math.random() * 100; const pool = config.canary.length && bucket < config.canaryPercent ? config.canary : config.stable; return pool[next++ % pool.length]; // round robin inside the pool } const server = http.createServer((req, res) => { if (req.url === "/admin/config") { if (req.method === "GET") return res.end(JSON.stringify(config)); let body = ""; req.on("data", (chunk) => (body += chunk)); return req.on("end", () => { config = { ...config, ...JSON.parse(body) }; console.log("config:", JSON.stringify(config)); res.end(JSON.stringify(config)); }); } const target = pickTarget(req); const upstream = http.request(target + req.url, { method: req.method, headers: req.headers }, (answer) => { res.writeHead(answer.statusCode, answer.headers); answer.pipe(res); }); upstream.on("error", () => { // the instance is down or vanished mid-request if (!res.headersSent) res.writeHead(502, { "content-type": "application/json" }); res.end(JSON.stringify({ error: "upstream unavailable", target })); }); req.pipe(upstream); }); server.listen(4000, () => console.log("router on http://localhost:4000"));
traffic.js - steady traffic, like real users
// traffic.js - steady traffic through the router, like real users. // Run: node traffic.js <seconds> <requests per second> [share of slow requests] const seconds = Number(process.argv[2] ?? 5); const perSecond = Number(process.argv[3] ?? 50); const slowShare = Number(process.argv[4] ?? 0); const buckets = []; // one row per half second const usersSeen = new Map(); // user -> set of versions that user saw const started = Date.now(); const pending = []; const timer = setInterval(() => { const user = "user-" + Math.floor(Math.random() * 20); const path = Math.random() < slowShare ? "/slow" : "/"; const sentAt = Date.now(); pending.push( fetch("http://localhost:4000" + path, { headers: { "x-user-id": user } }) .then(async (r) => ({ status: r.status, body: await r.json() })) .catch(() => ({ status: 0, body: {} })) .then(({ status, body }) => { const slot = Math.floor((sentAt - started) / 500); const row = (buckets[slot] ??= { v1: 0, v2: 0, errors: 0 }); if (status === 200) { row[body.version] = (row[body.version] ?? 0) + 1; if (!usersSeen.has(user)) usersSeen.set(user, new Set()); usersSeen.get(user).add(body.version); } else row.errors++; }), ); }, 1000 / perSecond); setTimeout(async () => { clearInterval(timer); await Promise.all(pending); let total = { v1: 0, v2: 0, errors: 0 }; buckets.forEach((row, i) => { if (!row) return; console.log(`${(i / 2).toFixed(1).padStart(4)}s v1 ${String(row.v1).padStart(3)} v2 ${String(row.v2).padStart(3)} errors ${row.errors}`); total = { v1: total.v1 + row.v1, v2: total.v2 + row.v2, errors: total.errors + row.errors }; }); const mixed = [...usersSeen.values()].filter((versions) => versions.size > 1).length; console.log("total:", JSON.stringify(total), "| users who saw both versions:", mixed, "of", usersSeen.size); console.log(total.errors === 0 ? "PASS - no request failed" : `FAIL - ${total.errors} requests failed`); process.exit(total.errors === 0 ? 0 : 1); }, seconds * 1000);
2

Start Blue, deploy Green beside it

Users reach Blue through the router; Green is tested directly.

Start the router and Blue (v1). Then start Green (v2) on its own port. The router still sends everyone to Blue, but you can test Green directly on port 4002 before any user reaches it.

Four terminals
node router.js # terminal 1 VERSION=v1 PORT=4001 node app.js # terminal 2 - Blue VERSION=v2 PORT=4002 node app.js # terminal 3 - Green curl localhost:4002/ # terminal 4: test Green directly curl localhost:4000/ # what users get
Output - measured
{"version":"v2"} <- Green works, tested privately {"version":"v1"} <- users are still on Blue
3

Switch everyone - and roll back

One config change each way, while traffic flows.

Start traffic for 7 seconds. After 2.5 seconds, switch the router to Green; after 5 seconds, switch back to Blue - a rollback. Watch the timeline: the switch happens inside one half-second slot, and not one request fails.

Terminal 4 - traffic, then the two switches from terminal 5
node traffic.js 7 40 # terminal 5, while traffic runs: curl -X POST localhost:4000/admin/config -d '{"stable":["http://localhost:4002"]}' # switch to Green curl -X POST localhost:4000/admin/config -d '{"stable":["http://localhost:4001"]}' # roll back to Blue
Output - measured (selected lines)
2.0s v1 19 v2 1 errors 0 2.5s v1 0 v2 19 errors 0 <- switched to Green 3.0s v1 0 v2 19 errors 0 4.5s v1 0 v2 19 errors 0 5.0s v1 18 v2 1 errors 0 <- rolled back to Blue 5.5s v1 20 v2 0 errors 0 total: {"v1":171,"v2":98,"errors":0} PASS - no request failed
4

Stop Blue the wrong way, then the right way

29 errors from killing it; 0 from letting it drain.

Now send traffic where 30% of requests are slow (2 seconds). Switch to Green after 2 seconds, and immediately stop Blue. First with kill -9, which ends the process at once: every request still running on Blue broke - 29 errors. Then start Blue again and repeat with a normal kill (SIGTERM), which app.js handles gracefully: it stops taking new requests, finishes the 26 in progress, and exits. 0 errors.

The two ways to stop Blue (use its process id)
node traffic.js 5 40 0.3 # 30% slow requests # after switching to Green: kill -9 <blue-pid> # wrong: stops at once kill <blue-pid> # right: SIGTERM - finish, then stop
Output - measured
kill -9 right after the switch total: {"v1":46,"v2":117,"errors":29} FAIL - 29 requests failed kill (SIGTERM) right after the switch total: {"v1":75,"v2":117,"errors":0} PASS - no request failed blue log: v1 stopping - finishing 26 requests in progress
Blue-green, measured
Switch to Green0 errors; all traffic moved within one half second.
Roll back to Blue0 errors; one config change.
Kill Blue at once29 errors - requests still running on Blue broke.
Stop Blue gracefully0 errors - Blue finished 26 requests first.

Practice on your own

  1. 1.

    Start Green with ERROR_RATE=0.3 (a bad release). Switch to it during traffic, notice the errors in the timeline, and roll back. How many requests failed before you rolled back?

    Hint

    The rollback is the same curl command with port 4001.

  2. 2.

    Stop Blue, then try to roll back to it. What happens to traffic? What does this tell you about when to shut Blue down?

    Hint

    The router sends everyone to a port where nothing is listening.

  3. 3.

    Write a small switch.js that checks Green's /health first and only switches if it answers 200.

    Hint

    fetch the health URL, check response.ok, then POST the new config.

  4. 4.

    Plan the expand-contract steps to split a name column into first_name and last_name without breaking blue-green rollback.

    Hint

    Add the new columns first; remove the old one only in a later release.

Comments

Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.

Loading comments...