Map the path of a real request
Choose one user action and list the services needed to complete it. Include name resolution, identity, the application and its data stores. Then add the control systems needed to operate the service, such as deployment and secret retrieval.
Draw arrows with a meaning: “must read from”, “must authenticate with” or “must receive traffic through”. A collection of cloud logos is not a dependency map. Name the data that crosses each boundary and the consequence when a dependency is unavailable.
Look for shared assumptions
Several instances may share a database, region, account or configuration source. Consider failures that affect all replicas at once. A configuration mistake can propagate across a redundant fleet just as efficiently as a correct change.
Ask a practical question for every dependency: if this stopped responding, what would the user see and what could the operator still do? Record the answer as an assumption until it has been tested. Avoid labelling a service resilient simply because it uses several machines.
Turn the map into a rehearsal
Choose a safe test environment and one failure condition. Predict the user-visible effect, run an authorised controlled test, and compare the observation with the prediction. Stop if the test affects systems outside the agreed scope.
Keep the map beside the runbook and update it after architecture changes. Its purpose is to support decisions, not to produce a perfect diagram. The Google SRE monitoring chapter provides useful context for connecting service symptoms to the underlying system.
Before you finish
- User and operator dependencies included
- Shared failure points identified
- Untested assumptions marked
- One failure scenario rehearsed
Technical reference
Google SRE: monitoring distributed systems
