Bulkhead
Isolate services into pools. One pool crashing doesn't sink the others.
Bulkhead: keep one slow service from sinking the rest
Order Service calls Payment and Inventory. When Payment becomes slow, should stock checks slow down too? Without bulkheads, they do - we measured a 20 ms call taking 6 seconds. This lesson explains why, how compartments stop it, and how to size them.
What the pattern is
Separate compartments, so one failure cannot sink everything.
The name comes from ships. The hull of a ship is divided into watertight compartments by walls called bulkheads. If the hull is damaged, water fills one compartment - and stops there. The ship stays afloat. Without the walls, one hole sinks the whole ship.
In software, the "water" is a dependency that becomes slow or broken, and the compartments are your service's resources: connections, threads, worker slots, memory. The Bulkhead pattern gives each dependency, or each kind of work, its own limited share of those resources. When one dependency fails, it can use up only its own share. Everything else keeps working.
This lesson follows Order Service, which calls two other services: Payment (to charge an order) and Inventory (to check stock). You will see Payment become slow and drag Inventory down with it - and then build the walls that stop it.
The problem: one shared pool
A slow dependency uses up resources that a healthy one needs.
Every outgoing call needs a resource while it waits for an answer: a connection, a thread, a slot. These come from a pool of limited size. Most services, by default, use one pool for all their outgoing calls.
Step through the diagram. Payment becomes slow - 3 seconds per call. The orders that need Payment take every connection in the pool and hold them for 3 seconds. Then stock checks arrive. Inventory is completely healthy, but there is no free connection to reach it. In our real run, a stock check that normally takes 20 ms took about 6 seconds. To the user, Inventory looks broken. It is not - Payment is.
1 - Orders take every connection
20 orders arrive. The first 10 take all 10 connections and wait 3 seconds for the slow Payment. The other 10 orders queue for a connection.
A real run: Order Service has one pool of 10 connections for all its outgoing calls. Payment becomes slow.
The solution: one compartment per dependency
Give Payment and Inventory their own pools, and limit how much may wait.
With bulkheads, Order Service keeps a separate pool for each dependency: 5 connections for Payment, 5 for Inventory. Payment can still use up its own 5 - but it cannot touch Inventory's. Stock checks stayed at 30 - 53 ms while Payment was slow.
There is a second half to the pattern. A full compartment should not collect an endless queue of waiting work: with only separate pools, the last orders waited up to 12 seconds. So the Payment bulkhead also limits its waiting area - here, 5 running and 5 waiting. Anything beyond that is refused at once, in 16 ms, with a clear "busy, try again" answer. Failing fast is better than making a user wait 12 seconds for a failure.
1 - Payment's compartment fills up
The orders can only use Payment's own compartment: 5 calls run, 5 wait. That compartment is full - but it is the only one that is full.
The same load with a separate pool for each dependency, and a limit on how much work may wait for Payment.
What can be a bulkhead
The walls can be built at every level, from one function to whole servers.
A bulkhead is any limit that keeps one kind of work from using resources another kind needs. In the lab you build two kinds: separate connection pools, and a concurrency limit with a small queue. The same idea appears at every level of a system.
Connection pool per dependencyA separate HTTP agent or client pool for each service you call. Built in the lab.Concurrency limit (semaphore)At most N calls running at once, plus a small queue; the rest are refused. Built in the lab with cockatiel.Thread pool per dependencyCommon in Java: each dependency gets its own threads (Resilience4j ThreadPoolBulkhead).Database pool per workloadSlow reports use one pool, fast logins another, so a heavy report cannot block sign-ins.Queue and workers per job typeEmails, video encoding and payments each get their own queue, so one backlog does not delay the others.Separate instances per customer groupCritical or paying customers are served by their own copies of a service.Container resource limitsKubernetes CPU and memory limits, so one container cannot starve the others on the same machine.How big should a compartment be?
Use your normal traffic, then add headroom.
A useful rule, called Little's Law: the number of calls in progress at the same time equals how many calls arrive per second multiplied by how long each call takes. If Payment receives 20 calls per second and each takes 0.1 seconds, about 2 calls are in progress at any moment on a normal day.
Set the limit a few times above that - say 5 or 6 - so normal bursts fit, and keep the waiting area small. If the limit is too small, good requests are refused on a normal day. If it is too big, a slow dependency can still hold a large share of your resources before the wall stops it. Watch two numbers in production: how often the bulkhead is full, and how many requests it refuses.
Normal day20 calls/s x 0.1 s = 2 calls in progress. A limit of 5 leaves room for bursts.Payment slow (3 s)20 calls/s x 3 s = 60 calls in progress - unless a bulkhead stops it at 5.Limit too small (1)Normal bursts are refused - users see errors on a healthy day.Limit too big (50)A slow Payment can hold 50 connections before the wall helps.Bulkhead, timeout and circuit breaker together
Three patterns, three different limits.
Bulkheads are usually combined with the patterns from the circuit breaker lesson, because each one limits something different. A timeout limits how long one call may wait. A bulkhead limits how many calls may wait at the same time. A circuit breaker stops making calls at all when most of them fail.
In the circuit breaker lab, a burst of 20 orders arriving at the same moment all reached Payment, because the breaker had nothing to judge yet. A bulkhead is exactly what limits such a burst: it does not need to wait for failures - it only counts how many calls are in progress.
TimeoutHow long ONE call may take. Frees resources from calls that hang.BulkheadHow MANY calls may run or wait at once. Protects everything else from one dependency.Circuit breakerWHETHER to call at all. Stops calling a dependency that keeps failing.Rate limiterHow many calls per second may START. Protects the dependency from too much traffic.Trade-offs, and when not to use it
Walls cost capacity, and need sizing.
Bulkheads are not free. Resources split into compartments cannot be shared: if Inventory is idle and Payment is busy, Payment still cannot borrow Inventory's connections. You need to choose and watch the sizes. And refusing requests is only better than waiting if the caller handles the refusal well - a clear message, a retry later, a fallback.
You probably do not need bulkheads for a service with a single dependency, or for internal tools where a slow afternoon is acceptable. You do need them when one service calls several others and some of those calls matter more than others - checkout versus recommendations, login versus reports.
ProOne slow dependency cannot take down unrelated features.ProOverload becomes a fast, clear "busy" answer instead of long waits.ConCapacity is split: an idle compartment cannot lend to a busy one.ConSizes must be chosen, monitored and adjusted as traffic changes.Interview questions
Short answers you can give in your own words.
What is the Bulkhead pattern? Isolating resources - pools, threads, queues, instances - per dependency or per workload, so a failure in one part can use up only that part's resources and the rest of the system keeps working.
How is a bulkhead different from a circuit breaker? A circuit breaker decides whether to call a dependency, based on its recent failures. A bulkhead limits how many calls to it can be in progress, regardless of failures - so it also works against sudden bursts.
How do you size a bulkhead? Start from normal concurrency (Little's Law: arrival rate x latency), add headroom for bursts, keep the queue small, and adjust using the measured rate of full and rejected requests.
What happens to requests when the bulkhead is full? They are rejected immediately, and the caller should return a clear "busy" response, use a fallback, or retry later with backoff - not wait without limit.
Build it: one shared pool, then bulkheads, then a limit
Three small Node.js services run on your own computer - no Docker. You will make Payment slow, measure what it does to Inventory through a shared connection pool, then build the walls and measure again. Every output is from a real run (Node 22, Express 5.2.1, cockatiel 4.0.0); your times will differ a little.
Set up the project
Two packages and plain Node.js - no Docker.
You need Node.js 18 or newer. Install Express for the services and cockatiel, a resilience library for Node.js that includes a bulkhead. The project has three services: Payment on port 3002, Inventory on port 3003, and Order Service on port 3001, which calls both.
mkdir bulkhead-lab && cd bulkhead-lab
npm init -y
npm pkg set type=module
npm install express cockatielbulkhead-lab/
payment-service.js step 2 - can be switched to slow
inventory-service.js step 2 - always healthy
http-call.js step 2 - makes a request through a chosen pool
load.js step 3 - 20 orders, then 10 stock checks
order-service-shared.js step 3 - one pool for everything
order-service-bulkhead.js step 5 - one pool per dependency
order-service-limited.js step 6 - pools + a limit on waiting work
check.js step 7 - checks your bulkhead automaticallyTwo services, and a small helper
Payment can be made slow; Inventory is always fine.
Payment Service answers in 20 ms, or in 3 seconds when you switch it to "slow". It also records the largest number of payment calls it handled at the same time - that number shows whether a bulkhead is working. Inventory Service always answers in about 20 ms.
Order Service will make its calls through http-call.js. It uses Node's built-in http module because that lets you choose the connection pool for each request: the agent option. An Agent is Node's connection pool, and choosing which agent each call uses is exactly how you build the walls.
// payment-service.js - port 3002. Switch it to "slow" to make every payment take 3 seconds.
import express from "express";
const app = express();
let mode = "ok";
let inFlight = 0;
let maxInFlight = 0;
const wait = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
app.post("/payments/:orderId", async (req, res) => {
inFlight++;
maxInFlight = Math.max(maxInFlight, inFlight);
await wait(mode === "slow" ? 3000 : 20);
inFlight--;
res.json({ orderId: req.params.orderId, status: "PAID" });
});
app.post("/admin/mode/:mode", (req, res) => {
mode = req.params.mode;
res.json({ mode });
});
app.get("/admin/stats", (req, res) => res.json({ mode, inFlight, maxInFlight }));
app.listen(3002, () => console.log("payment-service on http://localhost:3002"));// inventory-service.js - port 3003. Always healthy: answers in about 20 ms.
import express from "express";
const app = express();
const wait = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
app.get("/stock/:item", async (req, res) => {
await wait(20);
res.json({ item: req.params.item, inStock: 7 });
});
app.listen(3003, () => console.log("inventory-service on http://localhost:3003"));// http-call.js - a tiny helper: make an HTTP request through a given connection pool (Agent)
import http from "node:http";
export function call(url, { method = "GET", agent }) {
return new Promise((resolve, reject) => {
const req = http.request(url, { method, agent }, (res) => {
let body = "";
res.on("data", (chunk) => (body += chunk));
res.on("end", () => resolve(JSON.parse(body)));
});
req.on("error", reject);
req.end();
});
}Order Service with one shared pool
Measure what one slow dependency does to a healthy one.
The first Order Service uses one Agent for every outgoing call, with at most 10 connections in total (maxTotalSockets). load.js sends 20 orders at once and, 100 ms later, 10 stock checks, then prints how long each kind took.
With Payment healthy, everything is fast: stock checks 27 - 29 ms. With Payment slow, the 20 orders filled the pool, and the stock checks - which only need the healthy Inventory - took about 6 seconds each. Payment's own counter shows 10 calls at once: it had the whole pool.
// order-service-shared.js - port 3001. ONE connection pool for every outgoing call.
import http from "node:http";
import express from "express";
import { call } from "./http-call.js";
// one pool: at most 10 open connections in total, to all services together
const sharedPool = new http.Agent({ maxTotalSockets: 10 });
const app = express();
app.post("/orders/:orderId", async (req, res) => {
res.json(await call(`http://localhost:3002/payments/${req.params.orderId}`, { method: "POST", agent: sharedPool }));
});
app.get("/stock/:item", async (req, res) => {
res.json(await call(`http://localhost:3003/stock/${req.params.item}`, { agent: sharedPool }));
});
app.listen(3001, () => console.log("order-service (shared pool) on http://localhost:3001"));// load.js - 20 orders (Payment) and, 100 ms later, 10 stock checks (Inventory), all at once
const wait = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
const orders = Number(process.argv[2] ?? 20);
const checks = Number(process.argv[3] ?? 10);
async function timed(url, method) {
const t = Date.now();
const response = await fetch(url, { method });
const body = await response.json();
return { ms: Date.now() - t, status: response.status, body };
}
const summary = (label, results) => {
const ms = results.map((r) => r.ms);
const codes = {};
for (const r of results) codes[r.status] = (codes[r.status] ?? 0) + 1;
console.log(label.padEnd(16), "fastest", String(Math.min(...ms)).padStart(5), "ms | slowest", String(Math.max(...ms)).padStart(5), "ms | HTTP", JSON.stringify(codes));
};
const orderCalls = Array.from({ length: orders }, (_, i) => timed(`http://localhost:3001/orders/o${i}`, "POST"));
await wait(100);
const stockCalls = Array.from({ length: checks }, (_, i) => timed(`http://localhost:3001/stock/item${i}`, "GET"));
summary(`${checks} stock checks`, await Promise.all(stockCalls));
summary(`${orders} orders`, await Promise.all(orderCalls));node payment-service.js # terminal 1
node inventory-service.js # terminal 2
node order-service-shared.js # terminal 3
node load.js # terminal 4: Payment healthy
curl -X POST localhost:3002/admin/mode/slow
node load.js # Payment slow
curl localhost:3002/admin/statsPayment healthy
10 stock checks fastest 27 ms | slowest 29 ms | HTTP {"200":10}
20 orders fastest 43 ms | slowest 64 ms | HTTP {"200":20}
Payment slow
10 stock checks fastest 5982 ms | slowest 6203 ms | HTTP {"200":10} <- Inventory is healthy!
20 orders fastest 3040 ms | slowest 6051 ms | HTTP {"200":20}
payment stats: {"maxInFlight":10} <- Payment had the whole poolA surprise: an idle pool can block too
With keep-alive, a shared pool hurt even when Payment was healthy.
Many services turn on keepAlive, which keeps connections open after a request so the next one is faster. We tried the shared pool with keepAlive: true and Payment completely healthy. The orders finished in about 60 ms - and then the stock checks took about 6 seconds anyway.
The reason: the 10 kept-open connections to Payment were idle, but they still counted towards the shared limit of 10. Inventory had to wait until they were closed, about 5 seconds later. Nothing was slow and nothing was broken - the shared pool alone made a healthy call 300 times slower. Separate pools do not have this problem, because Payment's idle connections never count against Inventory's.
const sharedPool = new http.Agent({ keepAlive: true, maxTotalSockets: 10 });10 stock checks fastest 5997 ms | slowest 6208 ms | HTTP {"200":10}
20 orders fastest 44 ms | slowest 61 ms | HTTP {"200":20}Watch out: In Node.js, maxTotalSockets on a keep-alive Agent counts idle connections too. A shared pool can delay a healthy dependency even when nothing is slow. One pool per dependency avoids it.
Build the walls: one pool per dependency
Payment gets 5 connections, Inventory gets 5 - and they share nothing.
The bulkhead version has two Agents: paymentPool and inventoryPool, each limited to 5 connections (maxSockets). Each route uses its own pool. That is the whole change.
With Payment slow, the stock checks stayed at 38 - 58 ms. But look at the orders: the slowest one waited 12 seconds. 20 slow payments shared 5 connections, so they went through in four rounds of 3 seconds. The wall protected Inventory, but Payment's compartment collected a long queue.
// order-service-bulkhead.js - port 3001. A SEPARATE pool for each dependency: the bulkheads.
import http from "node:http";
import express from "express";
import { call } from "./http-call.js";
const paymentPool = new http.Agent({ keepAlive: true, maxSockets: 5 }); // compartment 1
const inventoryPool = new http.Agent({ keepAlive: true, maxSockets: 5 }); // compartment 2
const app = express();
app.post("/orders/:orderId", async (req, res) => {
res.json(await call(`http://localhost:3002/payments/${req.params.orderId}`, { method: "POST", agent: paymentPool }));
});
app.get("/stock/:item", async (req, res) => {
res.json(await call(`http://localhost:3003/stock/${req.params.item}`, { agent: inventoryPool }));
});
app.listen(3001, () => console.log("order-service (bulkheads) on http://localhost:3001"));10 stock checks fastest 38 ms | slowest 58 ms | HTTP {"200":10} <- protected
20 orders fastest 3023 ms | slowest 12044 ms | HTTP {"200":20} <- a 12-second queueLimit the waiting area
5 running, 5 waiting - everything else is refused at once.
cockatiel's bulkhead(5, 5) allows 5 calls to run and 5 more to wait. When both are full, execute() throws BulkheadRejectedError at once, and the route answers HTTP 503 with a clear message. Inventory keeps its own pool as before.
With Payment slow: stock checks 30 - 53 ms. 10 orders were served (5 after 3 seconds, 5 after 6), and 10 were refused in 16 ms instead of waiting up to 12 seconds. Payment never had more than 5 calls at once. While it was busy, /admin/bulkhead showed 5 running and 0 free waiting places.
// order-service-limited.js - port 3001. Separate pools, PLUS a limit on waiting work for Payment.
import http from "node:http";
import express from "express";
import { bulkhead, BulkheadRejectedError } from "cockatiel";
import { call } from "./http-call.js";
const paymentPool = new http.Agent({ keepAlive: true, maxSockets: 5 });
const inventoryPool = new http.Agent({ keepAlive: true, maxSockets: 5 });
// at most 5 payment calls running, at most 5 more waiting - everything else is refused at once
const paymentBulkhead = bulkhead(5, 5);
const app = express();
app.post("/orders/:orderId", async (req, res) => {
try {
const result = await paymentBulkhead.execute(() =>
call(`http://localhost:3002/payments/${req.params.orderId}`, { method: "POST", agent: paymentPool }),
);
res.json(result);
} catch (error) {
if (error instanceof BulkheadRejectedError) {
return res.status(503).json({ error: "payments are busy, please try again shortly" });
}
res.status(502).json({ error: error.message });
}
});
app.get("/stock/:item", async (req, res) => {
res.json(await call(`http://localhost:3003/stock/${req.params.item}`, { agent: inventoryPool }));
});
app.get("/admin/bulkhead", (req, res) => {
res.json({ running: 5 - paymentBulkhead.executionSlots, freeQueueSlots: paymentBulkhead.queueSlots });
});
app.listen(3001, () => console.log("order-service (bulkheads + limit) on http://localhost:3001"));10 stock checks fastest 30 ms | slowest 53 ms | HTTP {"200":10}
20 orders fastest 16 ms | slowest 6032 ms | HTTP {"200":10,"503":10}
payment stats: {"maxInFlight":5}
GET /admin/bulkhead while busy -> {"running":5,"freeQueueSlots":0}One shared poolStock checks ~6 s. Orders 3 - 6 s. Payment got 10 at once.One pool per dependencyStock checks 38 - 58 ms. Orders up to 12 s in a queue.Pools + limited waitingStock checks 30 - 53 ms. 10 orders served, 10 refused in 16 ms. Payment never more than 5.Check your work
An automated check instead of judging by eye.
check.js makes Payment slow, sends 20 orders and 10 stock checks, and checks four things: stock checks stay under 300 ms, Payment never handles more than 5 calls at once, some orders are refused, and refused orders are answered in under 200 ms.
Against order-service-limited.js all four checks passed. Against order-service-shared.js three failed - including a stock check that took 6.2 seconds.
// check.js - run while payment-service, inventory-service and your order-service are running
const wait = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
let failed = 0;
const check = (ok, text) => {
console.log(ok ? " PASS" : " FAIL", text);
if (!ok) failed++;
};
async function timed(url, method) {
const t = Date.now();
const response = await fetch(url, { method });
return { ms: Date.now() - t, status: response.status };
}
console.log("Payment is slow. 20 orders arrive, then 10 stock checks.");
await fetch("http://localhost:3002/admin/mode/slow", { method: "POST" });
const orders = Array.from({ length: 20 }, (_, i) => timed(`http://localhost:3001/orders/c${i}`, "POST"));
await wait(100);
const stock = await Promise.all(Array.from({ length: 10 }, (_, i) => timed(`http://localhost:3001/stock/c${i}`, "GET")));
const done = await Promise.all(orders);
const payment = await (await fetch("http://localhost:3002/admin/stats")).json();
await fetch("http://localhost:3002/admin/mode/ok", { method: "POST" });
const slowestStock = Math.max(...stock.map((r) => r.ms));
const refused = done.filter((r) => r.status === 503);
check(slowestStock < 300, `slowest stock check took ${slowestStock} ms (should be under 300 ms)`);
check(payment.maxInFlight <= 5, `at most ${payment.maxInFlight} payments ran at once (should be 5 or fewer)`);
check(refused.length > 0, `${refused.length} orders were refused (some should be, instead of waiting)`);
check(refused.every((r) => r.ms < 200), `refused orders answered in at most ${Math.max(0, ...refused.map((r) => r.ms))} ms (should be under 200 ms)`);
console.log(failed === 0 ? "All checks passed." : `${failed} check(s) failed.`);
process.exit(failed === 0 ? 0 : 1); PASS slowest stock check took 52 ms (should be under 300 ms)
PASS at most 5 payments ran at once (should be 5 or fewer)
PASS 10 orders were refused (some should be, instead of waiting)
PASS refused orders answered in at most 16 ms (should be under 200 ms)
All checks passed. FAIL slowest stock check took 6200 ms (should be under 300 ms)
FAIL at most 10 payments ran at once (should be 5 or fewer)
FAIL 0 orders were refused (some should be, instead of waiting)
3 check(s) failed.Practice on your own
- 1.
Change the Payment bulkhead to bulkhead(5, 0) - no waiting area at all. Run load.js with Payment slow. How many orders are served, and how fast are the refusals?
Hint
With no queue, everything beyond the 5 running calls is refused immediately.
- 2.
Add a timeout of 1 second to the Payment call (cockatiel has timeout(), or use the circuit breaker lesson's approach). With Payment slow, what happens to the 5 running calls now?
Hint
A timeout frees a slot after 1 second instead of 3.
- 3.
Make Inventory slow instead of Payment (add the same mode switch to inventory-service.js). Do orders stay fast with the bulkhead version? With the shared version?
Hint
The walls work in both directions.
- 4.
Use Little's Law to size the Payment bulkhead for 50 calls per second at 40 ms each. What limit would you choose, and why?
Hint
50 x 0.04 = 2 calls in progress normally. Add room for bursts.
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...
AI
System Design
Backend
- GraphQL8 modules · 69 lessons planned
- Core Python13 modules · 75 lessons planned
- FastAPI5 sections · 20 lessons
- Node.js14 modules · 206 lessons planned
- Node.js Performance7 chapters · 36 topics
- Event Loop Lifecycle6 phases · 3 scenarios
- Docker & Containerization11 modules · 144 lessons planned
- AWS for Developers14 modules · 219 lessons planned
- CI/CD & DevOps Automation10 modules · 134 lessons planned