Home Backend

Production Readiness Checklist for Node.js APIs

Backend

September 24, 2026

Production Readiness Checklist for Node.js APIs

The first deploy is rarely the problem. It is the Tuesday night six weeks later, when a container restarts in a loop and nobody can say why. Production readiness is less about clever code and more about the boring guarantees you put in place before real traffic arrives.

Node.js makes it easy to get an API listening on a port. It does not make that API a service. The difference is configuration, observability, failure behaviour and shutdown — the things that decide whether an incident lasts five minutes or five hours.

Work through the sections below before you point users at the thing. None of it requires a large team; most of it is an afternoon of careful work.

Configuration and secrets

Everything environment-specific belongs in environment variables, and every one of them should be validated at boot. A schema — zod, envalid, joi, whatever your team already uses — turns a missing DATABASE_URL into a clear failure at startup instead of a confusing 500 on the four hundredth request.

  • Fail fast. Parse and validate config once, export a typed object, and refuse to start if anything is missing or malformed.
  • Keep secrets out of the repository. Use your platform's secret store or a managed vault, and have a plan for rotation that does not require a code change.
  • Set the proxy correctly. Behind a load balancer or ingress, enable the trust proxy setting, or every request will appear to come from the same internal address. That quietly breaks rate limiting and audit logging.
  • Pin the runtime. Commit your lockfile, declare the Node version in the engines field, and use the same major version in CI and production.

Keep NODE_ENV set to production in production, but resist the temptation to branch application behaviour on it. Use explicit flags for anything that changes how the API behaves.

Logging that earns its keep

Write structured JSON to stdout and let the platform collect it. Do not write log files inside a container; they vanish with the container and fill the disk while they last. A library such as pino handles this well and keeps the per-request overhead low.

For every request, log the method, the route pattern rather than the raw URL, the status code, the duration, and a request ID. Return that same ID in a response header. When a support ticket says "it failed at about half past three", the ID is what turns a guess into a lookup.

Redact before you ship

Authorisation headers, cookies, session tokens, passwords and payment fields must never reach your log pipeline. Configure redaction paths rather than relying on reviewers to notice. Request bodies should be off by default; if you need them for a specific route, mask the fields you do not need.

If you cannot answer "what happened to request 8f3c…?" in under a minute, the logging is not finished.

Keep noise under control too. Log lifecycle events — server started, connection pool created, shutdown begun — at info level, and leave per-request chatter at debug. A log stream nobody reads is just a bill.

Health checks that tell the truth

Most teams need two endpoints, not one, because they answer different questions.

  • Liveness (/healthz) asks whether the process is still running. It should be cheap and should not touch the database. If it fails, the orchestrator restarts the container.
  • Readiness (/readyz) asks whether this instance can serve traffic right now: database reachable, migrations applied, critical downstream dependencies responding. Return 503 when it cannot.

Check dependencies with a short timeout — a second or two at most — and cache the result for a few seconds so a slow database does not cause every probe to hang. Be selective: making readiness depend on everything turns a minor third-party wobble into a full outage. Keep health routes out of your authentication layer and out of public documentation, and add a startup probe if your app takes a while to boot.

Security headers and the perimeter

A JSON API still benefits from sensible headers. Helmet sets most of them in one line; if you prefer to be explicit, the essentials are Strict-Transport-Security, X-Content-Type-Options: nosniff, a restrictive Referrer-Policy, and a content security policy that defaults to none for any HTML error page you render. Remove X-Powered-By. Set HSTS with a long max-age only once you are certain every subdomain is served over HTTPS.

Then cover the predictable attack surface:

  • Lock CORS to an explicit allowlist. Never combine a wildcard origin with credentials.
  • Cap request body size and set server timeouts, including headersTimeout and requestTimeout, so a slow client cannot hold a socket indefinitely.
  • Rate limit per IP and per authenticated token, with limits that reflect real usage rather than a round number.
  • Validate input at the edge of every route and use parameterised queries throughout.
  • Run a dependency audit in CI and keep the lockfile updated on a schedule.

Finally, make sure errors do not leak internals. Return a generic message with the request ID, log the stack trace server-side, and let the ID connect the two.

Graceful shutdown

Deploys cause most avoidable 502s. The fix is a shutdown sequence that gives in-flight requests time to finish.

  1. Listen for SIGTERM and SIGINT, and guard against receiving the signal twice.
  2. Flip readiness to unhealthy so the load balancer stops sending new traffic.
  3. Stop accepting new connections, then allow existing requests to complete within a deadline of, say, ten to fifteen seconds.
  4. Close database pools, cache clients and message consumers.
  5. Flush logs and exit with code 0. If the deadline expires, exit anyway — a stuck process is worse than a dropped request.

The subtle part is keep-alive. Calling server.close() waits for idle sockets that may never close on their own, so close idle connections explicitly and track active ones. Make sure the orchestrator's termination grace period is longer than your shutdown deadline, and treat uncaughtException and unhandledRejection as reasons to log and shut down rather than carry on in an unknown state.

Test it properly: send SIGTERM to an instance under load and watch your error rate. If you see 502s, the sequence needs work.

The final pass before you ship

Rehearse the whole thing in a staging environment that resembles production, then walk this list one last time:

  1. Config validated at boot, secrets in a managed store, proxy trust configured.
  2. Structured logs with request IDs and redaction in place.
  3. Separate liveness and readiness endpoints, both excluded from auth and rate limits.
  4. Security headers set, CORS restricted, body limits and timeouts tuned.
  5. Shutdown handled, tested under load, with alerts on error rate, latency and restart loops.
  6. A rollback plan you have actually tried, not one you have only written down.

An afternoon spent here buys you quiet nights. Skip it, and you will rebuild the same confidence at two in the morning, with worse tooling and an audience.

Photo: StockSnap / Pixabay

Related Posts

Developer Laptop Setup Checklist for New UK Hires
Tools

October 10, 2026

Developer Laptop Setup Checklist for New UK Hires

A practical checklist for setting up a secure, comfortable development laptop as a new UK hire, from disk encryption and access requests...

read more
A Beginner's Guide to Database Normalisation for Small Business Apps
Databases

October 09, 2026

A Beginner's Guide to Database Normalisation for Small Business Apps

A practical introduction to first, second and third normal forms, with clear examples showing how to structure small business data...

read more
How to Run Zero-Downtime Database Migrations
Databases

October 07, 2026

How to Run Zero-Downtime Database Migrations

Practical steps for changing production schemas without downtime: the expand-and-contract pattern, lock-aware statements, deploy...

read more