Cloud architecture & reliabilityPractical guides / September 2026
Web Dev QA DB Fra

Home / Reliability

Reliability notebook

Readiness, liveness and startup checks do different jobs

Choose health signals that route traffic and recover failed workloads without creating restart loops.

· 2 min read

Cloud architecture & reliability
The useful takeawayA dependency outage should not automatically become a restart storm.

Give each check one responsibility

In Kubernetes, readiness indicates whether a container is ready to accept traffic. Liveness can trigger a restart when its failure conditions are met. A startup probe lets slow-starting applications initialise before liveness and readiness checks begin. These are distinct mechanisms, not three names for the same endpoint.

Write the intended decision beside every check. “Can this instance serve a useful request?” is different from “is this process stuck in a way that restarting can repair?” If the decision is unclear, the check is probably collecting the wrong evidence.

Consider dependency failures

Suppose the database is temporarily unavailable. Restarting every application instance may not repair the database and may add startup load. Decide which failure belongs in readiness and which genuinely justifies restarting the process.

Test normal startup, a slow startup, a dependency interruption and recovery. Record the sequence of events instead of checking only the final healthy state. A service may eventually recover while still causing unnecessary disruption along the way.

Tune from observed behaviour

Choose timeouts and failure thresholds using measured startup and response behaviour. A check that is too impatient can reject healthy instances; an overly tolerant check can leave a broken instance in service for too long.

Keep checks lightweight and protect diagnostic details from unnecessary public exposure. Review the settings after meaningful application changes. The goal is a predictable response to a failure, with enough time and information for operators to understand what happened.

Before you finish

  • Each probe has a distinct purpose
  • Slow startup tested
  • Dependency outage tested
  • Recovery sequence observed

Technical reference
Kubernetes: container probes