← Back to microservices patterns map
senior engineering interviews

Deployment Interview Questions

The kind of questions asked in senior system design interviews at big tech, fintech, and large companies. Some ask what a thing is, some ask how you would build it, some give you a situation and ask what you would do.

5 Easy9 Medium11 Hard
Q01What is the difference between Blue-Green and Canary deployments? When would you choose one over the other?Easy
Blue-Green keeps two identical setups and moves everyone across in one step. It is built for undoing a bad release instantly, with no downtime. Canary sends a small share of traffic to the new version and widens it bit by bit.

Pick Blue-Green when: undoing a release has to be near-instant (payments, login), you can afford to run twice the servers, and the release is all-or-nothing. Pick Canary when: you want proof from real traffic before everyone gets it, your monitoring is good, and you can live with a small number of users hitting a problem while you check.
Q02How do you handle database migrations in a Blue-Green deployment?Hard
This is the hardest part of Blue-Green. During the switch, the old version and the new one are both talking to the same database. Change it in three steps - a pattern called Expand-Contract, or Parallel Change:

  • Expand: add the new columns or tables. The old code ignores them; the new code reads and writes both old and new.
  • Migrate: users are now on Green. Copy the existing data across into the new shape.
  • Contract: once Blue is gone for good, remove the old columns or tables in a later release.

Never drop or rename a column in the middle of a Blue-Green switch - the old version is still using it.
Q03What is a Rolling deployment and what are maxSurge and maxUnavailable?Easy
A Rolling deployment swaps old servers for new ones a few at a time, so there is always enough capacity left to serve users. In Kubernetes a running copy of your app is called a pod, and the number you want running is the replica count.

maxSurge: how many extra pods you allow above that count while the swap is happening. With 4 replicas and maxSurge: 1, you may briefly have 5.
maxUnavailable: how many pods are allowed to be missing during the swap. maxUnavailable: 1 means at most one can be down.

Setting both to 0 is not allowed - nothing could ever move. A common production setting is maxSurge: 25%, maxUnavailable: 25%.
Q04Explain Shadow deployment. What are the risks and how do you isolate side effects?Hard
Shadow deployment mirrors real production traffic to a new version asynchronously. The shadow response is discarded — users never see it.

Key risks:
  • Write side effects: Shadow requests that hit the same DB will cause double-writes. Solution: use a separate shadow DB, or implement a shadow-only write path with a x-shadow: true header that services intercept.
  • Third-party calls: Shadow requests may trigger real emails, payments, or API calls. Solution: stub external calls in shadow mode using if (req.headers['x-shadow']) return mockResponse().
  • Infrastructure load: Doubles compute. Set async timeouts so shadow failures never impact the live path.
Q05How do you handle user sessions and authentication during a Blue-Green switch?Medium
If you keep sessions in the memory of the Blue servers, they are lost the moment you switch - the Green servers cannot see them, so everyone is logged out.

Solutions:
  • JWT tokens (best): the token itself carries the proof of who the user is, signed by your server, so nothing has to be stored server-side. Green can check a token on its own.
  • Redis session store: Both Blue and Green connect to the same Redis cluster. Sessions survive the switch.
  • Sticky sessions: the load balancer keeps each user on whichever side they started on while the switch is in progress. It works, but it is fiddly and it makes undoing the release messier.
The habit to build: design services that remember nothing between requests. When something really must be remembered, put it in a token or in Redis, not in one server's memory.
Q06What metrics do you monitor during a Canary release? What triggers an automatic rollback?Medium
Primary metrics to watch:
  • Failures: the share of requests that come back as a server error. Act if it is more than 0.1% above normal.
  • Speed: the typical request (p50), the slow ones (p95) and the slowest 1% (p99). Raise an alert if p99 climbs more than 10%.
  • How hard the machines are working: CPU, memory, and how many database connections are in use.
  • Business numbers: how many people complete a purchase, a checkout, a login.
Automatic rollback triggers: Use tools like Flagger (Kubernetes), AWS CodeDeploy hooks, or custom health-check scripts. A common rule: if more than 1% of requests fail over five minutes, send all traffic back to the old version automatically and wake up whoever is on call. Implement with Prometheus + AlertManager + a webhook that triggers a deployment rollback.
Q07What is the difference between A/B testing deployment and Canary deployment?Easy
Canary is about releasing safely. The question is: is this version solid enough to replace the old one? Users are split at random, the point is to spot problems early, and it finishes when the new version has everyone.

A/B Testing is about trying out a product idea. The question is: does this version get better results from a particular group? Users are split by who they are - account, plan, country - the point is to measure things like sign-ups and usage, and it can run for weeks or stay on forever behind a feature flag.

They can run simultaneously: a canary rollout of a feature that is also being A/B tested by user segment.
Q08You are deploying a Node.js microservice that calls a third-party payment API. How would you approach the deployment strategy?Hard
This one is testing your judgement, not your memory. Use Blue-Green, with a few extra precautions:

  • Send a unique key with each payment: if the same request somehow reaches both sides during the overlap, the payment provider sees the same key twice and charges the customer once. This is called an idempotency key.
  • A circuit breaker on each side: if the payment API starts failing, stop calling it for a while instead of piling on more requests. Blue and Green each need their own. In Node.js, opossum does this.
  • Request tracing: Add a x-deployment-version header to outgoing payment calls so you can correlate issues in payment provider logs with your deployment version.
  • Canary first, then Blue-Green: Route 1% of traffic to Green → verify zero payment failures → flip fully to Green.
Q09How does Kubernetes implement Rolling updates internally? What happens if a pod's readiness probe fails?Medium
Kubernetes rolling update process:

1. It makes a new group of pods for the new version, and keeps the old group around
2. It starts new pods one at a time, staying within maxSurge
3. For each new pod it waits until the pod says it is ready to take traffic
4. Only then does it shut down one old pod, staying within maxUnavailable

If a new pod never reports ready: no traffic is sent to it, the rollout stops where it is, and the old pods keep serving users. The deployment sits there, stuck but still working. You can set progressDeadlineSeconds (e.g., 600s) after which Kubernetes marks the deployment as failed and you can configure automated rollback via kubectl rollout undo or CI/CD hooks.
Q10What is the "Expand-Contract" (Parallel Change) pattern and when is it required?Hard
Expand-Contract is how you change a database when the old and new versions of your code both have to use it at the same time - which is the case with Rolling, Blue-Green, and Canary.

Phases:
  • Expand: add the new columns or tables without touching the old ones. The old code ignores them; the new code uses them. Both can run.
  • Migrate: copy the existing data into the new shape. Both versions still work throughout.
  • Contract: in a later release, once the old version is gone for good, delete the old columns or tables.
You need this any time you rename a column, change its type, split a table, or merge two. People often skip the contract step, and the database slowly fills with things nothing reads.
Q11How would you implement Blue-Green deployment in AWS without Kubernetes?Medium
AWS-native Blue-Green options:
  • AWS CodeDeploy: Native Blue-Green support for EC2 Auto Scaling Groups and Lambda. Automatically shifts traffic, runs health checks, and rolls back. Supports AllAtOnce, HalfAtATime, OneAtATime traffic shifts.
  • ALB Weighted Target Groups: Create two Target Groups (Blue TG, Green TG) behind an Application Load Balancer. Use aws elbv2 modify-rule to adjust weights from 100/0 to 0/100.
  • ECS with CodeDeploy: AWS ECS + CodeDeploy integration provides built-in Blue-Green with automatic task set switching.
  • Lambda Traffic Shifting: Lambda Aliases support weighted traffic splitting natively — perfect for serverless canary/blue-green.
Q12Design a deployment strategy for a high-traffic Node.js e-commerce checkout API (50k RPM). Walk me through your full approach.Hard
This is a full system design answer. Recommended approach: Canary → Blue-Green hybrid
  • Infrastructure: Kubernetes on EKS, 3 AZs, ALB with weighted target groups
  • Phase 1 – Canary (1%): Deploy v2 pods, route 1% of checkout traffic. Monitor: error rate, p99 latency, payment success rate, cart abandonment rate. Duration: 30 minutes.
  • Phase 2 – Expand (5% → 25% → 50%): Automated promotion if error rate < 0.05% for 10 consecutive minutes per stage.
  • Phase 3 – Blue-Green flip: At 50% validation, do a full flip to 100% v2. Blue remains warm for 30 minutes.
  • Database: Expand-Contract pattern, separate read replica for v2 testing. Redis sessions shared.
  • Rollback SLA: <60 seconds via automated Flagger rollback + PagerDuty alert.
Q13What is a feature flag and how does it differ from a deployment strategy?Easy
Feature flags (feature toggles) decouple code deployment from feature activation. You deploy code with the feature disabled, then turn it on in production without a new deployment.

Deployment strategy: Controls how the binary/container is replaced on servers (Blue-Green, Rolling, etc.)
Feature flag: Controls which users see which features at runtime, independent of deployment.

They complement each other: you can do a Blue-Green deployment of code that has a feature flag — all instances (Blue and Green) get the new code, but the feature is disabled until you flip the flag in LaunchDarkly or your own flag service. This gives you "dark launches" — ship code safely, activate features on your schedule.
Q14How do you prevent cache poisoning during a Blue-Green switch when both environments share a Redis cache?Hard
The problem: the new version writes data into Redis in a new shape, and the old version cannot read it - or the other way round, if you undo the release.

Strategies:
  • Put the version in the key: v1:user:123 and v2:user:123. Neither version can see the other's entries. The cost is that the new version starts with an empty cache.
  • Make the new version read both shapes: it writes the new one but understands the old one too, so nothing breaks if you have to go back.
  • Cache flush on switch: Acceptable for non-critical caches. Add a warm-up phase (send synthetic traffic to pre-warm Green before switching).
  • Blue-Green Redis instances: Separate Redis per environment. More expensive but eliminates all cache conflict risks.
Q15What is a "readiness probe" vs "liveness probe" in Kubernetes and why are both critical for zero-downtime deployments?Medium
Liveness probe answers: is this thing still alive? If it fails, Kubernetes restarts the container. It catches an app that has frozen, run out of memory, or got itself into a bad state. Node.js example: /health/live returns 200 if the event loop is responsive.

Readiness probe answers: is this thing ready for users yet? If it fails, Kubernetes stops sending it traffic - but it does not restart it. Node.js example: /health/ready checks DB connection, Redis connection, and any warm-up completion.

Why both matter for zero-downtime: During a Rolling update, a new pod is only added to the load balancer pool when its readiness probe passes. This ensures no traffic is sent to a pod that hasn't finished initializing (DB pool warming, cache loading, etc.). Without a readiness probe, users hit a half-started server and get errors every time you deploy.
Q16How do you handle in-flight requests during a Blue-Green switch to avoid dropped connections?Hard
When flipping from Blue to Green, in-flight requests on Blue must complete before Blue shuts down.

Solution: Graceful shutdown in Node.js
  • Listen for SIGTERM signal (sent by K8s/Docker on shutdown)
  • Stop accepting new connections: server.close()
  • Wait for existing connections to drain (track active requests with a counter)
  • Set a hard timeout (30s) in case connections never close
  • In Kubernetes, set terminationGracePeriodSeconds: 30 and add a preStop: sleep 5 hook to account for iptables propagation delay before Blue stops getting traffic
At the load balancer level, configure connection draining (AWS ALB: "Deregistration delay" = 30s) so the LB waits for in-flight requests to complete before removing Blue from the target group.
Q17When would you NOT use Canary deployment despite its benefits?Medium
Canary is not always the right choice. Avoid it when:

  • Breaking API changes: If v2's API is incompatible with v1, routing some users to v2 and others to v1 will break cross-service calls. Use Blue-Green instead.
  • Stateful workloads: Databases, message queues, or services with sticky state that can't run two versions simultaneously.
  • Rules you have to follow: in finance and healthcare, auditors may require every user to be on the same version at the same time.
  • Weak monitoring: a canary you are not watching is worse than none at all - you get all the extra complexity and still only find out when everyone is affected.
  • Quiet services: at under 100 requests a minute, 5% of traffic is a handful of requests - far too few to tell a real problem from bad luck.
Q18Explain how Flagger automates Canary deployments on Kubernetes.Hard
Flagger is a piece of software you install into Kubernetes that runs the whole canary process for you. You describe what you want in a file of type Canary and it does the rest.

How it works:
  • You write down your limits - how many failures and how much slowdown you will accept - and the steps to climb (5% → 10% → 25% → 50% → 100%)
  • Flagger creates a primary deployment (stable) and a canary deployment (new) automatically
  • It queries Prometheus/Datadog metrics to evaluate health after each step
  • If metrics pass: automatically advances to next weight increment
  • If metrics fail: automatically rolls back to 0% canary and alerts
  • Supports Istio, Linkerd, NGINX, and AWS ALB for traffic shifting
Flagger takes the manual steps out of a canary - deciding when to widen and when to go back - which are exactly the steps people get wrong under pressure.
Q19You deployed a new version and users report bugs. Walk me through your complete rollback procedure for each strategy.Medium
  • Blue-Green: Update load balancer config to point back to Blue. Run nginx -s reload or update ALB target group weights. Time: ~5 seconds. Zero impact.
  • Canary: Set canary weight to 0% in your traffic controller (Flagger, NGINX, Istio). All traffic returns to stable v1. Time: ~10 seconds.
  • Rolling (K8s): Run kubectl rollout undo deployment/name. Kubernetes starts a new rolling update back to the previous version. Time: minutes (depends on replica count).
  • Shadow: No rollback needed — users were never exposed to the shadow version. Simply stop mirroring traffic.
  • Recreate: Re-deploy the old version. Has downtime. This is why Recreate is risky in production.
  • A/B Testing: Update routing rules to send all segments to the stable version. Feature flag off.
Q20How do microservice dependencies affect deployment strategy choice?Hard
Microservices add complexity because Service A's deployment may break Service B if their contracts change.

Key principles:
  • Consumer-driven contracts (Pact): Service consumers define what they expect. Producers must satisfy contracts before deploying. Prevents breaking changes from reaching production.
  • Backward compatibility window: When changing an API, maintain the old endpoint for at least one full deployment cycle. Rolling or Canary deployments may have the caller on v1 while the callee is on v2.
  • Deploy dependencies first: If Service B depends on Service A, deploy A first (with backward-compatible additions), then deploy B, then clean up A's old endpoints.
  • Prefer Canary over Rolling for breaking changes: Canary lets you validate the full service graph before committing. Rolling exposes all instances to the change gradually.
  • Service mesh (Istio/Linkerd): Enables fine-grained traffic control per service — run canary and rolling strategies independently across your service mesh.
Q21What is a "dark launch" and how does it relate to Shadow deployment?Easy
A dark launch means deploying code to production but keeping the feature invisible to users. The code is live but inactive.

Relationship to Shadow: They're related but different. Shadow deployment mirrors traffic to test new infrastructure. A dark launch deploys code that's controlled by feature flags.

Example: You dark-launch a new recommendation engine by deploying it with a feature flag set to false. Then you enable Shadow mode — the engine processes real traffic in parallel, and you compare its recommendations to the current engine's. Once confident, you enable the flag for 1% of users (canary). This is the full progression: dark launch → shadow → canary → full release.
Q22How do you implement zero-downtime deployment for a WebSocket-heavy Node.js service?Hard
A WebSocket is a connection that stays open, sometimes for hours. If a deploy cuts it, the user sees it: a disconnect, or a message that never arrives.

Strategy: Blue-Green with graceful WebSocket drain
  • Reconnect from the browser: when the connection drops, try again after 1 second, then 2, then 4, and so on. This is not optional - any server restart cuts every open connection.
  • Graceful shutdown: On SIGTERM, stop accepting new WS connections, send a { type: "server_restart", reconnect_in: 3000 } message to all connected clients, then wait for clients to acknowledge and reconnect to Green.
  • Session handoff via Redis: Store active subscription state in Redis. Green picks up the subscription state when the client reconnects.
  • Keep old connections where they are: tell the load balancer to leave existing connections on Blue while new ones go to Green. The old ones finish and disappear on their own.
  • No Rolling for WebSocket services: Rolling causes partial reconnects. Use Blue-Green where the switch is clean and the drain window is explicit.
Q23Compare CI/CD pipeline design for Blue-Green vs Rolling deployments.Medium
Rolling CI/CD pipeline (simpler):
Build → Test → Push image → kubectl set image → Monitor rollout → Auto-rollback on probe failure

Blue-Green CI/CD pipeline (more complex):
Build → Test → Push image → Deploy to inactive environment (Green) → Run smoke tests against Green → Validate health checks → Switch LB traffic → Monitor → Decommission Blue (after TTL) OR keep warm for rollback

Key difference: Blue-Green requires a validation gate between deploy and traffic switch. Rolling is continuous. Tools: GitHub Actions + Helm for Rolling; AWS CodeDeploy or custom scripts with ALB API for Blue-Green. For Kubernetes Blue-Green, tools like Argo Rollouts provide native support without manual scripting.
Q24What is "traffic shadowing" in Istio and how does it compare to application-level mirroring?Hard
Istio traffic mirroring (also called shadowing) mirrors a percentage of live traffic to a shadow service at the infrastructure level — no application code changes needed.

Istio VirtualService config example: mirror: host: v2-shadow with mirrorPercentage: 10 — sends 10% of requests to shadow, 100% to production. The shadow response is dropped by Istio.

vs. Application-level mirroring (custom middleware):
  • Istio: No code changes, infrastructure-managed, percentage-based, cleaner isolation
  • App-level: More flexible (can modify shadow request, log diffs inline), works without a service mesh, but adds latency risk to production path if not properly async
For production at scale: Istio/Envoy mirroring is preferred. For quick validation or ML testing: app-level middleware is faster to implement.
Q25You're a tech lead. A junior dev asks: "Why not just always use Blue-Green? It seems safest." How do you respond?Medium
Good question - Blue-Green is excellent. It is just not the right answer every time:

  • Cost: twice the servers, always, not just during a release. Across 200 services that is double the cloud bill forever. At that size Canary or Rolling is far cheaper.
  • Database complexity: Blue-Green's hardest problem is DB compatibility. Rolling/Canary with Expand-Contract migrations are often easier to manage in practice.
  • Things that hold state: Blue-Green assumes your service remembers nothing between requests. Databases, queue consumers, and background workers cannot simply be run twice over.
  • Warm-up: a freshly started server is slow for its first minutes - the runtime has not optimised the hot code yet and the database connections are not open. Switch a busy service onto a cold Green and users feel it.
  • Not always necessary: For dev/staging — Recreate is perfectly fine. For an internal tool with 10 users — Rolling is more than enough.
The answer is always the same: match the strategy to how critical the service is, how much traffic it gets, and how much operational experience the team has.