Circuit Breaker
Stops calls to failing downstreams to prevent cascade failures via CLOSED -> OPEN -> HALF-OPEN recovery probes.
What the pattern is
Stop calling a dependency that is failing, and fail fast instead of waiting.
The circuit breaker sits between a caller and a dependency and watches how calls turn out. While they succeed it stays out of the way. Once enough of them fail, it stops letting calls through at all for a while.
The borrowed name is exact: an electrical breaker does not repair the fault, it disconnects the circuit so the fault cannot burn the house down. A software breaker does not fix the dependency either - it stops the caller from destroying itself on a dependency that is already broken.
- Cascading failures, where one unhealthy service takes its callers down with it
- Resource exhaustion in the caller - threads, connections, memory held by doomed calls
- Network calls that have no chance of succeeding
- Request latency that climbs until everything upstream times out
- Extra load piled onto a service that is already struggling
One caller, one dependency. Everything that follows happens on this edge.
Request 1 goes out
The call leaves Order Service and takes a thread with it. That thread is now unavailable until Payment answers or the timeout expires.
Payment Service is unavailable. Each call still has to wait out its full 2 second timeout before the caller learns anything.
The same outage with a breaker
Three states, and the routing decision each one makes.
Put a breaker on that edge and the caller gets a third option besides waiting and succeeding: it can decline to make the call at all. Which of the three the breaker picks depends on what it has recently seen.
Step through the whole cycle below. It is worth noticing that the breaker never repairs anything - every improvement comes from calls that are not made.
1 · CLOSED, and traffic flows
Nothing is blocked. The breaker passes every call to Payment and records how each one turned out.
The same request arrives ten times. Where it ends up depends entirely on the breaker state at that moment.
Why it matters in a microservices system
One slow dependency, and how far the damage travels without a breaker.
A single service failing is a small problem. The reason this pattern exists is that in a system of services, a single failure does not stay single - it propagates up the dependency graph to everything that was waiting.
Tap any service in the diagram to see what it touches. Then step through what happens when the one at the bottom gets slow.
Payment sits behind Order, which sits behind the gateway. Every one of those edges is a place where waiting can spread.
Payment slows down
Responses that took 200ms now take 8 seconds. Every one of them still succeeds, so no error appears anywhere yet.
Payment has not gone down. It has just become slow - which is the more dangerous of the two.
CLOSED: counting, not blocking
Drive the breaker yourself and watch the failure rate decide when it trips.
In CLOSED every call goes through. The breaker is only observing - building up a window of outcomes and comparing the failure rate in that window against the threshold.
Send calls below and watch two things: the circuit will not trip before the minimum call count is reached no matter how bad the results are, and once it does trip, the rate is what decides it, not a streak.
Every call you send enters the sliding window. Try ten failures in a row, then try alternating - the difference tells you what the breaker is actually measuring.
next transition · 10 more calls before the failure rate is evaluated at all.
OPEN: failing fast, and what that buys
Both paths return an error. Only one of them costs three seconds.
An open circuit does not turn a broken request into a working one. What changes is how long the caller pays for the failure, and whether a thread is occupied while it waits.
That difference is the entire value of the pattern. At 500 requests per second, a 3 second timeout means 1,500 calls in flight at any moment; against an open circuit it means none.
Payment Service is down. Both calls end in an error - watch how long each one takes to get there. Timings illustrate a 3 second timeout rather than measuring a real system.
No circuit breaker
waits- Request sent
- Waiting on Payment
- Still waiting
- Timeout
- Error returned
Circuit breaker OPEN
fails fast- Request sent
- Breaker checked
- Rejected at the breaker
- Fallback returned
HALF-OPEN: testing without committing
The state that exists so recovery does not become the next outage.
A breaker that went straight from OPEN back to CLOSED would send the full load at a service the moment its timer expired - and a service that has just come back is the least able to absorb a thundering herd. HALF-OPEN is the ramp that prevents that.
In the playground above, drive a circuit to OPEN, skip the wait, and then send probes. One failure sends it straight back; the configured number of successes closes it and clears the window.
- Only the permitted number of probes reach the dependency; everything else still fails fast.
- A single failed probe reopens the circuit and restarts the full wait duration.
- Requiring several successes rather than one is what stops a flapping service from oscillating between states.
- Closing clears the failure window, so recovery starts from a clean slate.
Tuning the breaker
Six knobs. The defaults are a starting point, not a configuration.
Names differ between libraries but the concepts line up. The two that matter most are the minimum call count, which stops a single unlucky request from tripping the circuit, and the sliding window, which decides what counts as "recently".
The share of failed calls in the window that trips the circuit.
50% - half the recent calls failing is enough to stop trying.
How many calls must be recorded before the rate is calculated at all. Without it, one failure out of one call is a 100% failure rate.
10 calls before the breaker will consider opening.
What counts as recent. Count-based looks at the last N calls; time-based looks at the last N seconds. Time-based reacts better to low-traffic services.
Last 10 calls, or the last 60 seconds.
How long the circuit stays open before it is willing to probe again. Long enough for the dependency to actually recover, short enough that users are not stuck on a fallback.
30 seconds.
How many probes are allowed through to test recovery. More probes mean a more confident decision and more risk if the service is still broken.
3 probes; all must succeed to close the circuit.
The latency above which a successful call is still counted as a failure. This is what catches degradation that never returns an error.
Calls over 2 seconds count as failures; 50% slow trips the circuit.
Failure is not the only signal
A service answering HTTP 200 in twelve seconds is still unhealthy.
Counting only errors misses the most common real degradation: the dependency still answers, but far too slowly. Every one of those calls is a success by status code and a disaster by latency.
Most implementations can therefore treat a call slower than some threshold as a failure for breaker purposes. Resilience4j calls this the slow call rate; opossum gets there through its own timeout, which turns a slow call into a failed one.
Payment Service normally answers in 200ms. Nothing below has failed - each one is a success by status code. Move the slow-call threshold and watch how many of them the breaker counts as failures anyway.
A breaker counting only error codes would see a perfectly healthy dependency here. Reveal more calls, or lower the threshold.
What to serve while the circuit is open
Failing fast is half the pattern. The other half is what the user sees.
An open circuit gives the caller something it did not have before: the knowledge that the call will fail, early enough to do something else. What that something is depends entirely on the operation - and getting it wrong is how a resilience feature turns into a correctness bug.
Slightly stale data is better than none - listings, prices, recommendations.
{
"products": [ /* last known good */ ],
"stale": true,
"asOf": "2026-09-23T10:14:00Z",
"message": "Showing recently cached results."
}watch out · Say it is stale. A silently old price is worse than an honest error.
The worked example
Ten calls, one trip, and a recovery you drive yourself.
This is the sequence from the classic example: an Order Service calling Payment with a minimum of 10 calls, a 50% threshold, a 30 second open duration, and 3 permitted probes.
Step through the ten calls one at a time and watch the window fill. The circuit cannot trip before call ten however bad the results look, then trips immediately once the rate is evaluated. Afterwards, skip the wait and try recovering - both with passing probes and with a failing one.
Play the scripted sequence a call at a time. After it trips, skip the wait and drive the probes yourself.
next transition · 10 more calls before the failure rate is evaluated at all.
The neighbouring patterns
Timeout, retry, bulkhead, and rate limiter each solve a different failure.
These are routinely confused in interviews, and the distinction is simple: a timeout bounds one call, a retry repeats one call, a breaker stops a series of calls, and a bulkhead limits what any one dependency can consume.
Ordering matters when you stack them. The timeout sits innermost so slow calls become failures the breaker can count; retry sits inside the breaker so its attempts feed the same statistics; the breaker sits outermost so that once it is open, the retries do not run at all.
Retry: assume the failure was a blip
A retry policy repeats the same call, on the theory that the failure was transient. For a dropped packet or a brief restart, it is exactly right.
Both are reasonable. They do opposite things to the load on a service that is already struggling.
Implementations by language
The same pattern, in the runtimes it comes up in most often.
Note how differently they express the open condition: a percentage in Resilience4j and Polly, a consecutive-failure count in pybreaker, and a plain function in gobreaker. The Node example is the same opossum breaker used in the POC above.
Wraps any promise-returning function. Its own timeout turns slow calls into failures, and it emits events for every state change.
const CircuitBreaker = require('opossum');
const breaker = new CircuitBreaker(callPaymentService, {
timeout: 3000, // slow call = failed call
volumeThreshold: 5, // judge only after 5 calls - without it, ONE failure opens it
errorThresholdPercentage: 50,
resetTimeout: 10000, // OPEN duration before HALF-OPEN
});
breaker.fallback(() => ({ queued: true }));
breaker.on('open', () => logger.warn('payment breaker OPEN'));
await breaker.fire(orderId);The reference implementation for the JVM, and the one whose vocabulary the rest of this page borrows. Ships Retry, RateLimiter, Bulkhead, and TimeLimiter alongside the breaker.
CircuitBreakerConfig config = CircuitBreakerConfig.custom()
.failureRateThreshold(50)
.slowCallRateThreshold(50)
.slowCallDurationThreshold(Duration.ofSeconds(2))
.minimumNumberOfCalls(10)
.waitDurationInOpenState(Duration.ofSeconds(30))
.permittedNumberOfCallsInHalfOpenState(3)
.build();
CircuitBreaker breaker = CircuitBreaker.of("paymentService", config);Resilience strategies composed into a pipeline, so the breaker, retry, and timeout are declared in the order they wrap each other.
var pipeline = new ResiliencePipelineBuilder()
.AddCircuitBreaker(new CircuitBreakerStrategyOptions
{
FailureRatio = 0.5,
SamplingDuration = TimeSpan.FromSeconds(10),
MinimumThroughput = 8,
BreakDuration = TimeSpan.FromSeconds(30)
})
.Build();A decorator-based breaker. Count-based rather than rate-based: it opens after a fixed number of consecutive failures.
import pybreaker
breaker = pybreaker.CircuitBreaker(
fail_max=5, # consecutive failures before OPEN
reset_timeout=30, # seconds before HALF-OPEN
)
@breaker
def call_payment_service(order_id):
return requests.post(PAYMENT_URL, json={"order": order_id}, timeout=2)Takes a ReadyToTrip function instead of a threshold number, so the open condition is ordinary Go rather than configuration.
cb := gobreaker.NewCircuitBreaker(gobreaker.Settings{
Name: "payment-service",
MaxRequests: 3, // probes allowed in HALF-OPEN
Timeout: 30 * time.Second, // OPEN duration
ReadyToTrip: func(c gobreaker.Counts) bool {
return c.Requests >= 10 &&
float64(c.TotalFailures)/float64(c.Requests) >= 0.5
},
})
result, err := cb.Execute(func() (interface{}, error) {
return callPaymentService()
})Breaking at the infrastructure layer
The breaker does not have to live in your code.
A service mesh applies the same behaviour at the proxy in front of each service, which means it covers every language in the fleet and is configured by operations rather than shipped in a release.
The trade-off is fallback. A sidecar can reject the call, but only application code knows what to serve instead - so meshes usually handle the breaking while the application still owns the degraded response.
Outlier detection ejects unhealthy instances from the pool; connection and pending-request limits cap what can queue up. No application code is involved.
Tap Payment Service to see everything whose health now depends on the breaker in front of it.
A real outage, in aggregate
A thousand orders during a Payment outage, counted rather than animated.
A food delivery app takes a thousand orders while Payment is unavailable. Drawing a thousand requests would teach nothing; the counters are where the story is, and they are also what an on-call engineer would actually be looking at.
Watch how few calls reach the dependency before the breaker has seen enough to stop them, and how everything after that costs the caller nothing.
Only one of these is unavailable. Without a breaker, it is enough to take the whole order flow down.
These are the numbers an operator watches, not a thousand animated arrows. The values are simulation inputs, not a measurement of a production system.
Where it goes wrong
Four mistakes that turn a breaker into an outage of its own.
- Counting the wrong errors. A 400 or a 404 means the request was wrong, not that the service is unhealthy. Counting them opens the circuit on healthy traffic and blocks everyone. Count connection failures, timeouts, 5xx, and slow calls.
- Retry storms. Aggressive retries against a struggling service add load exactly when it can least afford it, which causes more failures, which causes more retries. Cap attempts, use exponential backoff with jitter, and put the breaker outside the retry.
- Thresholds set without data. A threshold below the dependency’s normal error rate opens constantly; one far above it never opens at all. Set both the threshold and the slow-call limit from observed latency and error percentiles.
- No visibility. A breaker that opens silently looks identical to a dependency that got fast. State transitions belong in your metrics and your alerts.
- Current state per breaker, and every transition between states
- Failure rate and slow-call rate inside the window
- Calls rejected while OPEN - the traffic the breaker absorbed
- Fallback invocations, split by which fallback ran
- Time spent in OPEN, and how many probe cycles recovery took
Practice
Configure it yourself, then predict what the sequence does.
Ten calls, six of them failures, against a breaker you can retune. Answer first, then run the sequence and see whether you were right. Then change the minimum call count or the threshold and run it again - the same ten calls can produce a different answer.
Configuration: 10 minimum calls, 50% threshold, 10 second open duration, 2 probes. The sequence below is fixed until you change it.
next transition · 10 more calls before the failure rate is evaluated at all.
After all ten calls, what state is the circuit in?
Interview questions
Twelve questions, from the definition to the failure modes of recovery.
BeginnerWhat is a circuit breaker?
A resilience pattern that stops a caller from repeatedly invoking a dependency that is failing or degraded, so the caller fails fast instead of tying up resources on calls that will not succeed.
BeginnerWhat are the three states?
CLOSED, where calls pass through and failures are counted. OPEN, where calls are rejected immediately without reaching the dependency. HALF-OPEN, where a limited number of probes test whether the dependency has recovered.
BeginnerWhat happens in the OPEN state?
Calls are rejected at the breaker and never reach the dependency. The caller gets an immediate error or a fallback response, and the dependency gets a period with no traffic from this caller.
BeginnerWhat is HALF-OPEN for?
To test recovery without committing full traffic to it. A small number of probes go through; if they succeed the circuit closes, and if any fails it returns to OPEN for another wait period.
IntermediateHow is a circuit breaker different from a retry?
A retry assumes the failure was transient and repeats the same call. A breaker assumes the dependency is unhealthy and stops calling it. Retry adds load; the breaker removes it. They address opposite situations and are usually used together.
IntermediateCan a successful HTTP response still open the circuit?
Yes, in implementations with a slow-call threshold. A call that returns 200 after twelve seconds is counted as a failure for breaker purposes, because the latency is what damages the caller regardless of status code.
IntermediateWhy should some exceptions be excluded from the failure count?
Because not every error says something about the dependency’s health. A 400 or a 404 means the request was invalid. Counting those opens the circuit on the basis of client mistakes, blocking healthy traffic for everyone.
IntermediateWhy does the minimum call count exist?
To keep the failure rate statistically meaningful. Without it, the first failed call on a quiet service is a 100% failure rate and the circuit opens on a sample size of one.
AdvancedWhat is a cascading failure, and how does a breaker prevent one?
One service fails, its callers block on it, their resources fill with pending work, and they begin failing for their own callers. The breaker cuts the chain at its first link: the caller stops waiting on the unhealthy dependency, so its resources stay available for everything else it does.
AdvancedWhere does the breaker belong relative to retry and timeout?
Timeout innermost, so a slow call becomes a countable failure. Retry next, so its attempts are recorded in the same statistics. Breaker outermost, so once it is open the retries do not run at all. Putting the breaker inside the retry lets the retry hammer an open circuit.
AdvancedIs breaker state per instance or shared?
In libraries like opossum, Resilience4j, and gobreaker it is per process, so each instance learns independently and a fleet of twenty opens at twenty slightly different times. Shared state through Redis makes the decision global but adds a dependency on the path it is meant to protect. Service meshes sidestep this by deciding per proxy against a shared health view.
AdvancedWhat can go wrong when the circuit closes again?
Every caller that was failing fast resumes at once, and the full load lands on a service that has only just recovered, which can knock it straight back over. Permitting a small number of HALF-OPEN probes, and ramping rather than switching, is what keeps recovery from becoming the next outage.
Key takeaways
- The breaker exists to protect the caller first and the dependency second. Both benefits are real, but the caller is the one that would otherwise fall over.
- CLOSED counts, OPEN rejects, HALF-OPEN probes. Every configuration knob is about when to move between those three.
- Slow is a failure. If the breaker only counts errors it will miss the most common real degradation.
- An open circuit needs an answer, not just a rejection - cached data, a safe default, a queued job, or an honest error.
- Pair it with a timeout inside and a bulkhead beside it; a breaker alone leaves the other failure modes open.
- Count only errors that indicate dependency health, and never let retries run inside an open circuit.
- Thresholds come from observed latency and error rates, not from the example in the README.
- State transitions belong in metrics and alerts. A silently open breaker looks exactly like a healthy, very fast service.
Build it: Order Service, Payment Service, and a breaker between them
Everything above explained the pattern. Now you build it and measure it. Two small Node.js services run on your own computer - no Docker. You will break Payment Service on purpose, watch what happens to Order Service without protection, then add a timeout, then a circuit breaker, and measure each step. Every output on this page is from a real run (Node 22, Express 5.2.1, opossum 10.0.0); your times will differ a little.
Set up the project
Two packages, five small files, no Docker.
You need Node.js 18 or newer - it has fetch and AbortSignal.timeout built in. Make a new folder and install two packages: Express, to write the two web services, and opossum, a circuit breaker library for Node.js. Setting "type": "module" lets the files use import.
You will run each service in its own terminal window. Payment Service listens on port 3002, Order Service on port 3001. A third terminal sends test orders.
mkdir breaker-lab && cd breaker-lab
npm init -y
npm pkg set type=module
npm install express opossumbreaker-lab/
payment-service.js step 2 - the service that fails
order-service-naive.js step 3 - calls Payment with no protection
send-orders.js step 3 - sends test orders and measures them
order-service-timeout.js step 4 - adds a timeout
order-service.js step 5 - adds the circuit breaker
payment-service-safe.js step 8 - charges each order only once
check.js step 9 - checks your breaker automaticallyPayment Service: a dependency you can break
A pretend payment provider with a switch: ok, slow, or down.
This is the service that will fail. POST /payments/:orderId charges an order. A small admin API lets you change its behaviour while it runs: "ok" answers at once, "slow" waits 5 seconds and then charges, "down" returns HTTP 500 immediately.
It also counts how many requests it received and how many charges really happened. Those two numbers are how you will measure every step.
// payment-service.js - a pretend payment provider on port 3002
import express from "express";
const app = express();
let mode = "ok"; // "ok" | "slow" | "down"
let received = 0; // how many payment requests arrived
const charges = []; // every charge that really happened
const wait = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
app.post("/payments/:orderId", async (req, res) => {
received++;
if (mode === "down") return res.status(500).json({ error: "payment provider down" });
if (mode === "slow") await wait(5000); // very slow, but it still works
charges.push(req.params.orderId); // the customer is charged here
res.json({ orderId: req.params.orderId, status: "PAID" });
});
// controls for the lab
app.post("/admin/mode/:mode", (req, res) => {
mode = req.params.mode;
res.json({ mode });
});
app.get("/admin/stats", (req, res) => res.json({ mode, received, charges: charges.length }));
app.post("/admin/reset", (req, res) => {
received = 0;
charges.length = 0;
res.json({ ok: true });
});
app.listen(3002, () => console.log("payment-service on http://localhost:3002"));node payment-service.js
# change its behaviour from another terminal:
curl -X POST localhost:3002/admin/mode/slow
curl -X POST localhost:3002/admin/mode/down
curl -X POST localhost:3002/admin/mode/okOrder Service with no protection
Measure what happens when Payment is slow, and when it is down.
The first Order Service calls Payment with a plain fetch: no timeout, no breaker. It also counts how many requests are waiting for Payment at the same moment. send-orders.js sends a number of orders - all at once, or one every few milliseconds - and prints the results, the response times, and Payment's counters.
When Payment works, 10 orders took about 60 ms in total. When Payment is slow, every single order waited about 5 seconds, and 20 requests were stuck waiting at the same time. In a real service each waiting request holds memory and a connection; with enough traffic, Order Service runs out and stops answering everything - including requests that have nothing to do with payments. That is the cascading failure from section 03.
When Payment is down, the answers are fast (errors come back in about 20 ms), but notice the last line: Payment received every one of the 20 requests. Order Service keeps hitting a service that is already failing, which makes it harder for that service to recover.
// order-service-naive.js - calls Payment with NO protection, on port 3001
import express from "express";
const app = express();
let waiting = 0; // requests currently stuck waiting for Payment
let maxWaiting = 0;
app.post("/orders/:orderId", async (req, res) => {
waiting++;
maxWaiting = Math.max(maxWaiting, waiting);
try {
const response = await fetch(`http://localhost:3002/payments/${req.params.orderId}`, {
method: "POST",
});
if (!response.ok) throw new Error(`payment returned ${response.status}`);
res.json(await response.json());
} catch (error) {
res.status(502).json({ error: error.message });
} finally {
waiting--;
}
});
app.get("/admin/stats", (req, res) => res.json({ waiting, maxWaiting }));
app.listen(3001, () => console.log("order-service (naive) on http://localhost:3001"));// send-orders.js - sends N orders at the same time and reports what happened
const count = Number(process.argv[2] ?? 10);
const gapMs = Number(process.argv[3] ?? 0); // wait between orders (0 = all at once)
const wait = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
const started = Date.now();
const results = await Promise.all(
Array.from({ length: count }, async (_, i) => {
await wait(i * gapMs);
const t = Date.now();
const response = await fetch(`http://localhost:3001/orders/order-${i + 1}`, { method: "POST" });
const body = await response.json();
return { status: body.status ?? body.error, ms: Date.now() - t };
}),
);
const summary = {};
for (const r of results) summary[r.status] = (summary[r.status] ?? 0) + 1;
const times = results.map((r) => r.ms);
console.log("orders sent:", count, gapMs ? `(one every ${gapMs} ms)` : "(all at once)");
console.log("results: ", summary);
console.log("fastest: ", Math.min(...times), "ms | slowest:", Math.max(...times), "ms");
console.log("total time: ", Date.now() - started, "ms");
const payment = await (await fetch("http://localhost:3002/admin/stats")).json();
console.log("payment received", payment.received, "requests, charged", payment.charges, "times");node order-service-naive.js # terminal 2
node send-orders.js 10 # terminal 3: payment ok
curl -X POST localhost:3002/admin/mode/slow
node send-orders.js 20
curl localhost:3001/admin/stats
curl -X POST localhost:3002/admin/mode/down
node send-orders.js 20payment ok
orders sent: 10 (all at once)
results: { PAID: 10 }
fastest: 36 ms | slowest: 57 ms
total time: 62 ms
payment slow
orders sent: 20 (all at once)
results: { PAID: 20 }
fastest: 5026 ms | slowest: 5040 ms <- every order waited 5 seconds
order-service stats: {"maxWaiting":20} <- 20 requests stuck at once
payment down
orders sent: 20 (all at once)
results: { 'payment returned 500': 20 }
fastest: 12 ms | slowest: 26 ms
payment received 20 requests, charged 0 times <- still hitting a failing serviceAdd a timeout first
Never wait forever - but a timeout alone still sends every request.
The first fix is a timeout: AbortSignal.timeout(1000) gives up on Payment after 1 second. Now a slow Payment costs each order 1 second instead of 5, and requests stop piling up for long.
But look at what Payment received: all 20 requests. A timeout protects the caller from waiting; it does nothing to stop the caller from sending. Every new order still tries Payment, waits a full second, and fails. A breaker adds the missing piece: after enough failures, stop trying for a while.
// order-service-timeout.js - step 3: a timeout, but no breaker
import express from "express";
const app = express();
app.post("/orders/:orderId", async (req, res) => {
try {
const response = await fetch(`http://localhost:3002/payments/${req.params.orderId}`, {
method: "POST",
signal: AbortSignal.timeout(1000), // give up after 1 second
});
if (!response.ok) throw new Error(`payment returned ${response.status}`);
res.json(await response.json());
} catch (error) {
res.status(502).json({ error: error.name === "TimeoutError" ? "payment timed out" : error.message });
}
});
app.listen(3001, () => console.log("order-service (timeout only) on http://localhost:3001"));results: { 'payment timed out': 20 }
fastest: 1003 ms | slowest: 1040 ms <- every order still pays 1 second
total time: 4809 ms
payment received 20 requests <- every order still reaches PaymentAdd the circuit breaker
Wrap the call; after enough failures, fail fast.
Now the real thing. The Payment call goes into a function, callPayment, which keeps its 1-second timeout. opossum wraps it in a breaker. Instead of calling callPayment directly, the route calls breaker.fire(orderId). The breaker decides whether to call Payment at all.
The settings: judge only after at least 5 calls (volumeThreshold), open when half of them fail (errorThresholdPercentage), stay open for 5 seconds (resetTimeout). While the breaker is open, fallback() decides what the customer gets. Three event listeners print every state change, and /admin/breaker shows the breaker's state and counters.
We measured three situations. With orders arriving one every 200 ms and Payment slow, the breaker opened after the first failures: Payment received 11 of the 20 requests, and the other 9 orders were answered in about 3 ms instead of 1 second. With Payment down and orders every 100 ms, only 5 requests reached Payment; 15 failed fast.
The third situation shows a real limit. When all 20 orders arrived in the same moment, Payment received all 20. The breaker can only judge calls that have finished, and none had finished yet - all 20 were already on their way. A breaker protects you from a stream of traffic over time, not from one big burst. For bursts you need a bulkhead or a rate limiter (section 11).
// order-service.js - step 4: timeout + circuit breaker
import express from "express";
import CircuitBreaker from "opossum";
const app = express();
// 1. The risky call, with its own timeout
async function callPayment(orderId) {
const response = await fetch(`http://localhost:3002/payments/${orderId}`, {
method: "POST",
signal: AbortSignal.timeout(1000),
});
if (!response.ok) throw new Error(`payment returned ${response.status}`);
return response.json();
}
// 2. Wrap it in a breaker
const breaker = new CircuitBreaker(callPayment, {
timeout: 1500, // opossum's own limit (a backup for the fetch timeout)
volumeThreshold: 5, // judge only after at least 5 calls
errorThresholdPercentage: 50, // open when half of them fail
resetTimeout: 5000, // stay OPEN for 5 s, then try one call (HALF-OPEN)
});
// 3. What the user gets while the breaker is open
breaker.fallback((orderId) => ({ orderId, status: "PAYMENT_UNAVAILABLE" }));
// 4. Log every state change
breaker.on("open", () => console.log(new Date().toISOString().slice(11, 23), "breaker OPEN"));
breaker.on("halfOpen", () => console.log(new Date().toISOString().slice(11, 23), "breaker HALF-OPEN"));
breaker.on("close", () => console.log(new Date().toISOString().slice(11, 23), "breaker CLOSED"));
app.post("/orders/:orderId", async (req, res) => {
const result = await breaker.fire(req.params.orderId);
res.json(result);
});
app.get("/admin/breaker", (req, res) => {
const state = breaker.opened ? "OPEN" : breaker.halfOpen ? "HALF-OPEN" : "CLOSED";
res.json({ state, ...breaker.stats });
});
app.listen(3001, () => console.log("order-service (breaker) on http://localhost:3001"));node order-service.js
# terminal 3
curl -X POST localhost:3002/admin/mode/slow
node send-orders.js 20 200 # 20 orders, one every 200 ms
curl localhost:3001/admin/breakerpayment slow, one order every 200 ms
results: { PAYMENT_UNAVAILABLE: 20 }
fastest: 3 ms | slowest: 1034 ms
payment received 11 requests <- 9 orders never reached Payment
order-service log: breaker OPEN
payment down, one order every 100 ms
results: { PAYMENT_UNAVAILABLE: 20 }
fastest: 2 ms | slowest: 42 ms
payment received 5 requests <- 15 failed fast
payment slow, all 20 orders at the same moment
fastest: 1030 ms | slowest: 1046 ms
payment received 20 requests <- a burst: the breaker had nothing to judge yetNo protectionEvery order waited ~5 s; 20 requests stuck at once; Payment got 20.Timeout onlyEvery order failed after ~1 s; Payment still got 20.Timeout + breakerPayment got 11; the other 9 orders failed in ~3 ms.Breaker, all at oncePayment got 20 - a burst arrives before any call has failed.Watch it recover
OPEN, then HALF-OPEN after 5 seconds, then CLOSED.
Break Payment, open the breaker, then fix Payment. An order sent immediately still gets the fallback - the breaker is open and does not even try. After resetTimeout (5 seconds) the breaker moves to HALF-OPEN and lets the next call through as a test. It succeeded, so the breaker closed, and orders flow normally again.
Look at the times in the log: OPEN at :33.115, HALF-OPEN at :38.116 - exactly 5 seconds later - and CLOSED as soon as the test call succeeded.
curl -X POST localhost:3002/admin/mode/down
node send-orders.js 10 100
curl -X POST localhost:3002/admin/mode/ok
curl -X POST localhost:3001/orders/early # right away
sleep 5
curl -X POST localhost:3001/orders/probe # after the reset timeout
curl -X POST localhost:3001/orders/nextbreaker: {"state":"OPEN","failures":5,"fallbacks":10,"rejects":5, ...}
early -> {"orderId":"early","status":"PAYMENT_UNAVAILABLE"} still OPEN
probe -> {"orderId":"probe","status":"PAID"} the test call
next -> {"orderId":"next","status":"PAID"}
order-service log:
05:27:33.115 breaker OPEN
05:27:38.116 breaker HALF-OPEN
05:27:38.902 breaker CLOSEDThe setting that opens on one failure
Without volumeThreshold, a single bad call shuts the door.
Many examples - including the opossum example most people copy - set errorThresholdPercentage: 50 and nothing else. We removed volumeThreshold from order-service.js to see what that does.
One order failed while Payment was briefly down. That is 1 failure out of 1 call: 100%, which is more than 50%. The breaker opened at once. Payment was fine again a moment later, but the next 5 orders all got PAYMENT_UNAVAILABLE without even trying - Payment received only that 1 request. A single unlucky call blocked every customer for 5 seconds.
volumeThreshold tells the breaker: do not judge until you have seen at least this many calls. Choose it from your real traffic: high enough that one or two failures cannot trip it, low enough that a real outage trips it quickly.
payment down for one order:
unlucky -> PAYMENT_UNAVAILABLE
payment ok again, 5 more orders:
results: { PAYMENT_UNAVAILABLE: 5 } <- refused, although Payment works
payment received 1 requests
order-service log: breaker OPEN <- opened after ONE failureWatch out: Always set a minimum number of calls before the breaker may open: volumeThreshold in opossum, minimumNumberOfCalls in Resilience4j. Without it, the first failure after a restart can open the circuit.
The payment trap: a timeout is not a failure
Customers were told payment failed - and 11 of them were charged.
Go back to step 5's slow-Payment run. All 20 customers got PAYMENT_UNAVAILABLE. Payment received 11 requests. Our script printed "charged 0 times" - but it printed that too early. We waited 6 seconds and asked Payment again: 11 charges. Payment was slow, not broken. It finished those 11 charges after Order Service had already given up and told the customers it failed.
Then it gets worse. A customer who sees "payment unavailable" presses Pay again. We simulated all 20 customers retrying once Payment was healthy: Payment charged every retry. 31 charges for 20 orders - 11 customers paid twice.
Two fixes work together. First, Payment Service must be idempotent: the order id is an idempotency key, and a second request for the same order returns the first result instead of charging again (payment-service-safe.js). With it, the same retries produced exactly 20 charges. Second, the fallback must not say "failed" when it does not know. Change it to PAYMENT_PENDING, and give the shop a way to ask Payment what really happened. We checked order-7 right after the timeout: NOT_PAID. Five seconds later: PAID.
1 - The order goes to Payment
20 customers place orders, one every 200 ms. Payment is slow today: it will answer each request after 5 seconds.
Step 8, a real run: Payment is slow, the breaker gives up after 1 second - and Payment finishes the charge anyway.
// payment-service-safe.js - the same provider, but each order is charged at most once
import express from "express";
const app = express();
let mode = "ok";
let received = 0;
const charges = new Map(); // orderId -> charge. The order id is the idempotency key.
const wait = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
app.post("/payments/:orderId", async (req, res) => {
received++;
const { orderId } = req.params;
if (charges.has(orderId)) { // already charged: do NOT charge again
return res.json({ ...charges.get(orderId), repeated: true });
}
if (mode === "down") return res.status(500).json({ error: "payment provider down" });
if (mode === "slow") await wait(5000);
if (!charges.has(orderId)) charges.set(orderId, { orderId, status: "PAID" });
res.json(charges.get(orderId));
});
// "Did this order get paid?" - lets the shop check instead of guessing
app.get("/payments/:orderId", (req, res) => {
res.json(charges.get(req.params.orderId) ?? { orderId: req.params.orderId, status: "NOT_PAID" });
});
app.post("/admin/mode/:mode", (req, res) => {
mode = req.params.mode;
res.json({ mode });
});
app.get("/admin/stats", (req, res) => res.json({ mode, received, charges: charges.size }));
app.listen(3002, () => console.log("payment-service (safe) on http://localhost:3002"));// retry-orders.js - every customer who saw an error presses "Pay again"
const count = Number(process.argv[2] ?? 20);
const results = {};
for (let i = 1; i <= count; i++) {
const response = await fetch(`http://localhost:3002/payments/order-${i}`, { method: "POST" });
const body = await response.json();
const key = body.repeated ? "already paid - not charged again" : body.status;
results[key] = (results[key] ?? 0) + 1;
}
console.log("retries:", results);
const stats = await (await fetch("http://localhost:3002/admin/stats")).json();
console.log("payment: total charges =", stats.charges, "for", count, "orders");Payment slow, 20 orders, breaker on - then wait 6 s:
customers saw: { PAYMENT_UNAVAILABLE: 20 }
payment stats: {"received":11,"charges":11} <- 11 customers WERE charged
Every customer presses "Pay again":
payment-service.js total charges = 31 for 20 orders <- 11 charged twice
payment-service-safe.js total charges = 20 for 20 orders
retries: { 'already paid - not charged again': 11, PAID: 9 }breaker.fallback((orderId) => ({
orderId,
status: "PAYMENT_PENDING", // not "failed": we do not know yet
message: "We could not confirm your payment yet. We will check and update your order.",
}));
// Ask the payment service what really happened
app.get("/orders/:orderId/payment", async (req, res) => {
const response = await fetch(`http://localhost:3002/payments/${req.params.orderId}`);
res.json(await response.json());
});
// measured, with payment-service-safe.js in slow mode:
// POST /orders/order-7 -> {"status":"PAYMENT_PENDING", ...}
// GET /orders/order-7/payment right away -> {"status":"NOT_PAID"}
// GET /orders/order-7/payment after 5 s -> {"status":"PAID"}Watch out: Never retry a payment, or any call that changes money or data, without an idempotency key - and never tell a customer a payment failed when the call only timed out. "We are not sure yet" is the honest answer.
Check your work
An automated check, instead of judging by eye.
check.js tests the behaviour this lab is about. With Payment down and orders arriving every 100 ms, Payment must receive 6 requests or fewer, the last orders must be answered in under 50 ms, and the breaker must be OPEN. Then it fixes Payment, waits for the reset timeout, and checks that an order succeeds and the breaker is CLOSED.
Run it with payment-service.js and your order-service.js running. Our breaker passed all 5 checks. As a test of the test, we also ran it against the unprotected order-service-naive.js: it failed 3 checks - Payment received 10 of 10 requests, and there is no breaker to open or close.
// check.js - run while payment-service and order-service are running: node check.js
const wait = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
const post = (url) => fetch(url, { method: "POST" }).then((r) => r.json());
const get = (url) => fetch(url).then((r) => r.json()).catch(() => ({}));
let passed = 0;
let failed = 0;
const check = (ok, text) => {
console.log(ok ? " PASS" : " FAIL", text);
ok ? passed++ : failed++;
};
console.log("1. Payment goes down; 10 orders arrive, one every 100 ms");
await post("http://localhost:3002/admin/mode/down");
const before = (await get("http://localhost:3002/admin/stats")).received;
const times = [];
for (let i = 1; i <= 10; i++) {
const t = Date.now();
await post(`http://localhost:3001/orders/check-${i}`);
times.push(Date.now() - t);
await wait(100);
}
const reached = (await get("http://localhost:3002/admin/stats")).received - before;
check(reached <= 6, `payment received ${reached} of 10 requests (should be 6 or fewer)`);
check(times.slice(-4).every((ms) => ms < 50), `last 4 orders answered in ${times.slice(-4).join(", ")} ms (should be under 50 ms)`);
check((await get("http://localhost:3001/admin/breaker")).state === "OPEN", "breaker is OPEN");
console.log("2. Payment recovers; wait for the reset timeout");
await post("http://localhost:3002/admin/mode/ok");
await wait(5500);
const probe = await post("http://localhost:3001/orders/check-probe");
check(probe.status === "PAID", `first order after the wait was ${probe.status} (should be PAID)`);
check((await get("http://localhost:3001/admin/breaker")).state === "CLOSED", "breaker is CLOSED again");
console.log(failed === 0 ? `All ${passed} checks passed.` : `${failed} check(s) failed.`);
process.exit(failed === 0 ? 0 : 1);1. Payment goes down; 10 orders arrive, one every 100 ms
PASS payment received 5 of 10 requests (should be 6 or fewer)
PASS last 4 orders answered in 7, 3, 6, 4 ms (should be under 50 ms)
PASS breaker is OPEN
2. Payment recovers; wait for the reset timeout
PASS first order after the wait was PAID (should be PAID)
PASS breaker is CLOSED again
All 5 checks passed. FAIL payment received 10 of 10 requests (should be 6 or fewer)
FAIL breaker is OPEN
FAIL breaker is CLOSED again
3 check(s) failed.Practice on your own
- 1.
Change the fallback to PAYMENT_PENDING and add GET /orders/:orderId/payment, as in step 8. Put Payment in slow mode, place one order, and check its status right away and after 5 seconds.
Hint
The fallback receives the same arguments as fire(). The status route only needs to forward the question to Payment.
- 2.
Set resetTimeout to 2000 and run step 6 again. How long does the breaker stay OPEN now? What could go wrong if it is much too short?
Hint
Each HALF-OPEN test call reaches Payment. How often would a struggling Payment be tested?
- 3.
Make Payment "slow" mean 1.5 seconds instead of 5, and lower the fetch timeout to 500 ms. Does the breaker still open? Why - what does opossum count as a failure?
Hint
A call that times out throws, and a thrown error is a failure.
- 4.
Send 20 orders all at once with the breaker on. Payment receives all 20. Which pattern from section 11 would limit how many calls can be in flight at the same time?
Hint
Think of ship compartments.
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...