# Zero-Downtime Deployment Checklist for Production Releases

> A production-ready zero-downtime deployment checklist covering readiness probes, graceful shutdown, database migrations, traffic management, and rollback verification.

- **Published:** 2025-06-09
- **Author:** Clixo
- **Reading time:** 5 min read
- **Tags:** zero-downtime, deployment, checklist, production
- **Canonical URL:** https://clixo.sh/blog/zero-downtime-deployment-checklist-production

Most teams claim their deployments are zero-downtime. Most are wrong. Requests get dropped during container restarts, database migrations lock tables mid-deploy, and load balancers route traffic to instances that are not ready yet. The downtime is real — it is just brief enough that no one notices until it affects a peak-traffic window.

This checklist covers what actually needs to be verified before calling a deployment zero-downtime.

```mermaid
flowchart LR
  A["Pre-deploy checks"] --> B["Run backwards-compatible DB migrations"]
  B --> C["Start new instances"]
  C --> D{"Readiness probe passes?"}
  D -- No --> WAIT["Wait and retry"]
  WAIT --> D
  D -- Yes --> E["Drain old connections"]
  E --> F["Add to load balancer rotation"]
  F --> G{"Smoke tests pass?"}
  G -- No --> ROLL["Rollback to previous version"]
  G -- Yes --> DONE["Deploy complete"]
```

## Zero-Downtime Deployment Checklist

### Pre-Deploy: Application Readiness

**Readiness probe configuration**

- [ ] The `/health` or `/ready` endpoint returns a non-200 status until the application is fully initialized — not just started
- [ ] The readiness probe waits for database connections, cache warmup, and any async initialization before returning healthy
- [ ] The probe has a realistic `initialDelaySeconds` value based on observed startup time, not a guess
- [ ] Traffic is not routed to a new instance until the readiness probe passes

**Graceful shutdown**

- [ ] The application catches `SIGTERM` and stops accepting new requests immediately
- [ ] In-flight requests are given time to complete before the process exits (typically 15-30 seconds)
- [ ] The `terminationGracePeriodSeconds` value is longer than your longest expected request duration
- [ ] Long-running background jobs are paused or handed off, not killed mid-execution

### Pre-Deploy: Database Migrations

Database migrations are the most common cause of silent downtime in otherwise well-designed deploys.

- [ ] Migrations are backwards-compatible — the current version of the code can run against both the old and new schema
- [ ] No `ALTER TABLE ... LOCK` operations on large tables during peak traffic (run them with `pt-online-schema-change` or equivalent)
- [ ] Migrations are run before the new application version starts receiving traffic, not after
- [ ] A rollback migration exists and has been tested
- [ ] Column renames happen in two deploys: add the new column, then remove the old one — never rename in one step

### Deploy: Traffic Management

- [ ] The load balancer removes an instance from rotation before the new version starts deploying to it
- [ ] Connection draining is enabled — existing connections complete before the load balancer stops sending to the old instance
- [ ] The new instance is added to rotation only after passing the readiness probe
- [ ] The deploy is staged — not all instances are updated simultaneously

### Deploy: Kubernetes-Specific

- [ ] `rollingUpdate.maxUnavailable` is set to 0 — no pod is terminated until a replacement is ready
- [ ] `rollingUpdate.maxSurge` allows at least one extra pod to spin up during the rollout
- [ ] Pod disruption budgets prevent the cluster from evicting too many pods simultaneously during node maintenance
- [ ] The `preStop` hook introduces a short sleep (typically 5 seconds) to account for endpoint propagation delay from the service registry

### Deploy: Feature Flags

- [ ] New or risky functionality is gated behind a feature flag, not shipped fully enabled
- [ ] The flag is tested in the disabled state before the deploy goes out
- [ ] The flag can be toggled without a new deploy

### Post-Deploy: Verification

- [ ] Automated smoke tests run immediately after the deployment completes
- [ ] Key HTTP endpoints return expected status codes
- [ ] Error rate (HTTP 5xx) is at baseline for at least 5 minutes post-deploy
- [ ] P95 latency has not regressed
- [ ] Any critical business flows (checkout, authentication, primary API calls) work end-to-end

### Post-Deploy: Rollback Readiness

- [ ] The previous artifact version is still available and can be re-deployed in under 5 minutes
- [ ] The rollback procedure has been documented and is accessible to any on-call engineer
- [ ] If a migration was run, the rollback plan accounts for the schema state
- [ ] The on-call engineer knows the rollback command or process before the deploy starts — not after something goes wrong

## The Items Teams Most Commonly Skip

**Readiness probes that lie.** A health endpoint that returns 200 immediately on startup — before the database connection pool is established — will route traffic to an unready instance. This drops requests during every deploy.

**No draining on the load balancer.** Removing an instance from rotation without draining existing connections terminates in-flight requests. Enable connection draining and set a drain timeout that matches your request duration.

**Single-step column renames.** Renaming a database column in one migration breaks the running application version that still references the old column name. This is a preventable outage.

**Untested rollback paths.** The rollback plan is only useful if it works. Test it in a staging environment before you need it in production at 2 AM.

---

Zero-downtime deployment is not a single setting you enable. It is the result of getting a handful of things right simultaneously — application lifecycle, database hygiene, traffic management, and observability. This checklist is the starting point.

[Clixo helps engineering teams design deployment infrastructure that ships safely — let's review your setup](https://clixo.sh/#contact)

---

Clixo · 1141 W Bryn Mawr Ave, Itasca, IL 60143, US · [hello@clixo.sh](mailto:hello@clixo.sh)
[Start a build](https://clixo.sh/#contact) · [All services](https://clixo.sh/services) · [Agent guide (llms.txt)](https://clixo.sh/llms.txt)
