The Production Readiness Review: 42 Checks Before a Service Takes Traffic

A 42-point production readiness checklist for a service about to take real traffic: ownership and on-call, probes and graceful shutdown, instrumentation, SLOs and alerts, dependency timeouts, migrations and restore, rollback, capacity, and secrets. Every line is something you can look at and mark done.

A production readiness review is the conversation that keeps a new service from becoming somebody else's 3 a.m. problem. This is the checklist form of it: 42 checks, grouped into the nine areas where launches actually come apart, each one a thing you can look at and mark done.

Use it as a gate rather than a survey. Walk it with the team that will carry the pager, in one sitting, and give every line one of three answers: done, deliberately accepted with a name against it, or blocking. The middle answer is the one that earns the exercise — almost every check here is fine to skip on purpose and expensive to skip by accident.

Nothing here requires a particular vendor. Where a check names a tool, it is because the mechanism is easiest to point at in that tool; the check is about the mechanism.

This page is the whole deliverable. There is no course behind it, no part two, and no follow-up series — if you arrived here from a form, this is what you were sent.


1. Ownership and On-Call

Four checks. If the service has no owner who can be reached and can act, nothing else on this list gets used.

  1. A named owning team, in a machine-readable place. Not a person — a team, recorded somewhere a tool can read: the service catalogue entry, CODEOWNERS, repo metadata. Test it by asking who gets paged while the one person everybody names is on holiday. If the answer needs that person, ownership is not recorded anywhere.
  2. A rotation that is populated for the next 30 days. Look at the schedule in the paging tool, not at the intent in a document. A single-person rotation means the launch has one point of failure, and that person sleeps.
  3. Escalation that terminates. Every path ends at someone who can make a decision, with a stated timeout at each hop. Fire a test page and watch where it lands when nobody acknowledges it.
  4. The on-call can reach production. Access, VPN, break-glass credentials, and the ability to execute the rollback — proven by that person doing it once, in working hours, before launch. Permission that has never been exercised is not access.

2. Health, Readiness, and Graceful Shutdown

Five checks. Most launch-day instability is an orchestrator and a process disagreeing about whether the process is ready.

  1. Liveness and readiness are different endpoints. Liveness answers “restart me”; readiness answers “send me traffic”. Wire them to the same handler and a slow dependency will get a healthy process killed in a loop. The probe semantics are worth reading properly in the Kubernetes reference.
  2. Readiness fails when a hard dependency is unavailable — and checks nothing else. It should test what the service genuinely cannot work without. A readiness probe that opens a database connection on every call becomes its own load problem: cache the result for a few seconds.
  3. Start-up is distinguished from failure. A slow-booting process needs a startup probe or a generous initial delay, or the orchestrator kills it mid-warm-up, forever, and the symptom looks like a crash loop rather than a timing bug.
  4. SIGTERM is handled. On signal: stop accepting new work, finish in-flight requests, close pools, exit. Delete a pod deliberately and read the logs — you should see the drain, not an abrupt end mid-request.
  5. Drain ordering is correct. The load balancer must stop sending traffic before the process stops accepting it. That means a short pre-stop delay, and a termination grace period longer than your slowest legitimate request.

3. Instrumentation You Can Debug From

Five checks. The question is not whether you have dashboards; it is whether one person can answer “which request failed, and where” without redeploying.

  1. Structured logs, with a request id on every line. JSON or logfmt, one event per line, and the id propagated to downstream calls so that one user's bad minute is one query rather than a manual join across four services.
  2. Traces cross the process boundary. Context must travel on outbound calls. The interoperable default is the W3C Trace Context traceparent header, which OpenTelemetry SDKs emit by default. Verify by opening one real trace and confirming it contains more than one service.
  3. Latency is recorded as a distribution, not an average. You need p50, p95 and p99 to be derivable from the stored data. An average cannot be re-cut after the fact, and it hides precisely the tail that will page you. The performance reference covers why percentiles do not average.
  4. Rate, errors, duration and saturation exist per endpoint. Request rate, error rate, duration distribution, and the saturation of whatever is scarce here — connection pool, queue depth, worker slots. Anything missing is a blind spot you will discover during the incident rather than before it. See observability for what each signal actually answers.
  5. No secrets and no personal data in the logs. Before launch, grep a day of staging logs for tokens, cookies, Authorization values and email addresses. Logged data is copied to places you do not control within minutes and cannot be recalled.

4. SLOs and Alerts That Mean Something

Five checks. An alert that nobody can act on trains the rotation to ignore alerts, which is worse than having none.

  1. One written objective, with a number and a window. “99.9% of requests under 400 ms, measured over 28 days” is an objective. “Fast and reliable” is a preference. Get the number agreed by whoever will complain when it is missed.
  2. The indicator is measured where the user is. Server-side latency excludes the network, the TLS handshake and the client. If the commitment is about experience, measure at the edge or in the client — and state which, so the number is not quietly re-interpreted later.
  3. Pages fire on symptoms, not causes. Page on user-visible failure and on the objective burning down. High CPU on its own is a dashboard line; it becomes a page only when it predicts harm you cannot otherwise see.
  4. Every page carries a runbook link and an owner in its payload. Open each link and confirm the document exists and names the first three actions. A runbook that says “investigate the issue” is a missing runbook. The incident management reference covers the roles this assumes.
  5. Non-urgent alerts have a ticket path. Anything that does not need hands within the hour goes to a queue with an owner, not to the phone. Mixed-urgency paging is the mechanism by which a rotation stops reading pages.

5. Dependencies and Failure Containment

Five checks. Your availability is bounded by how your service behaves when something it calls stops answering.

  1. A written dependency list, each entry marked hard or soft. Hard means the service cannot serve without it; soft means it degrades. If every entry is hard, you have no degradation story and your availability is the product of everyone else's.
  2. Every outbound call has a timeout you chose. Check the client construction, not the documentation: a Go http.Client left at its zero value has no timeout, and Python's requests without an explicit timeout argument waits indefinitely. An unbounded call is how one slow dependency consumes all your workers.
  3. Retries are bounded, backed off, jittered, and safe to repeat. Retrying a non-idempotent write duplicates it; retrying in lockstep converts a blip into a stampede against a service that is already struggling. Send an idempotency key, or do not retry.
  4. Soft-dependency failure has been exercised, not assumed. Block the dependency in staging — firewall rule, wrong port, null route — and watch the request path. “It degrades gracefully” is a hypothesis until you have watched it happen. Chaos engineering is this check, done continuously and with a blast radius.
  5. Overload is shed deliberately, and says so. Past capacity, reject early with 429 or 503 and a Retry-After header (defined in RFC 9110) instead of queueing work until every request times out. Silent queueing turns a capacity problem into a total outage.

6. Data, Migrations, and Restore

Five checks. This is the section whose failures are not recoverable by redeploying, which is why it is worth an hour rather than a glance.

  1. Schema changes are compatible with both versions of the code. During any rolling deploy, old and new code run at once. Expand first — add the nullable column, dual-write — ship the code, and contract in a later release. Never in one step.
  2. No blocking DDL on a hot table without knowing the lock it takes. Know which statements take an exclusive lock on your engine and version, set a lock timeout so the migration fails instead of the application, and know the abort plan before you start.
  3. A restore has been performed, not merely a backup taken. Restore into a scratch environment, then run a query that proves row counts and the presence of a recent write. A backup nobody has restored is an assumption with a cron entry.
  4. Recovery point and recovery time objectives are written as numbers. And the restore you just timed fits inside the recovery time objective. If it does not, either the number is wrong or the plan is — decide which now, in daylight.
  5. Connection pool arithmetic done against the server limit. Instances multiplied by pool size must stay under the database's connection limit with headroom for migrations and for a human holding a shell. Set a statement timeout so one pathological query cannot hold a connection indefinitely.

7. Deploy and Rollback

Four checks. The value of a rollback is entirely in whether it has been run recently on this service.

  1. Rollback executed this week, in production, on this service. The previous artefact still deploys and someone has actually deployed it. A rollback path that has never run in production is a hypothesis with a runbook attached. Artefact promotion is what makes this cheap.
  2. The forward-only cases are named. Some migrations and every external side effect — mail sent, payment captured, webhook delivered — cannot be undone by redeploying the old build. Write down which ones, and what the recovery is instead of rollback.
  3. Deploy is one command or one click, and its output states what changed. Any procedure that requires remembering a sequence will be performed wrong at 3 a.m. by someone who did not write it.
  4. New behaviour sits behind a flag that defaults off. Then the launch is a config change reversible in seconds without a build, and deploy and release become separate events with separate blast radii. See feature flags at scale for the lifecycle cost of doing this badly.

8. Capacity, Load, and Cost

Four checks. The useful output of a load test is not a pass — it is a number and a bottleneck.

  1. Load tested to expected peak, then past it, until something breaks. Record the point at which latency turns the corner and which resource ran out first. That pair is your capacity model; “the test passed” is not. Capacity planning covers why high utilisation is fragile rather than efficient.
  2. Resource requests and limits set from measurement. In Kubernetes, requests drive scheduling and limits drive enforcement: a memory limit below real usage gets the container OOM-killed under load, a CPU limit below real usage throttles it instead, and a missing request lets it be scheduled onto a node that cannot feed it.
  3. Headroom for the loss of one failure domain. If losing one zone or one replica puts you over capacity, you are already inside the incident. Add a disruption budget so that a routine node drain cannot take the quorum with it.
  4. Resources tagged for cost attribution at creation time. Retro-tagging competes with whatever else is on the roadmap, and loses. Tag owning team and environment now, so the first surprising invoice is answerable from the tags themselves rather than by reconstruction — the mechanics are in cost allocation tagging.

9. Secrets, Access, and Data Handling

Five checks. These are the ones that are cheap before launch and expensive to retrofit afterwards, because by then the credential is everywhere.

  1. No secret in the image, the repository, or a plain environment dump. Delivered at runtime from a secret manager, and the process must not print its own environment on an error page or in a crash handler. Secrets management covers the injection patterns.
  2. Rotation has been executed once. Rotate a credential before launch and watch the service pick up the new value without a restart you did not plan. A rotation procedure nobody has run is a procedure that does not work yet.
  3. Least privilege demonstrated by removal. Take the permission away in staging; if nothing breaks, it was not needed. Roles granted wide “for now” are not narrowed after launch, because after launch narrowing them is a risk nobody wants to own.
  4. The data this service stores is written down. Which fields, where they live, how long they are kept, and who can read them. If any of it is personal data it belongs in your organisation's processing records before launch — under the GDPR that is the Article 30 record — and whether that obligation falls on your organisation at all, and in what form, is a question for whoever owns your compliance programme, not for this page.
  5. Transport encrypted end to end, including inside the cluster. “Internal traffic” crosses hardware you do not own wherever a managed service, a peering link or another region sits in the path. Monitor certificate expiry with enough lead time that renewal is a task rather than an incident.

If you only have an hour, do these four and defer the rest: run the rollback for real (30), restore a backup and time it (27), block a soft dependency and watch what happens (23), and confirm the person who will be paged can actually reach production (4). Those four are the checks whose absence turns an ordinary incident into a long one — each of them replaces a belief with an observation, which is the only thing this list is really for.

Jakub Dimitri Rezayev
Jakub Dimitri Rezayev
Founder & Chief Architect • Garnet Grid Consulting

Jakub holds an M.S. in Customer Intelligence & Analytics and a B.S. in Finance & Computer Science from Pace University. With deep expertise spanning D365 F&O, Azure, Power BI, and AI/ML systems, he architects enterprise solutions that bridge legacy systems and modern technology — and has led multi-million dollar ERP implementations for Fortune 500 supply chains.

View Full Profile →