Ask what must work first
Start with a workflow people depend on: reading published content, accepting an order or accessing a document. Ask how long that workflow can be interrupted and how much recent work could be lost. Different workflows on the same platform may have different needs.
Recovery time and recovery point objectives express those tolerances. They are targets to design and test against, not promises created by writing a number in a document. Record who agreed to them and which assumptions they depend on.
List every recovery step
Include detection, decision-making, access retrieval, provisioning, data restoration and verification. A technically fast restore can still miss the objective if nobody has the credentials or authority to start it.
Check what the backup method can actually recover. PostgreSQL’s documentation distinguishes backup approaches with different recovery requirements. Use the documentation for your database and version, and do not assume a daily copy can meet a much shorter data-loss tolerance.
Run a timed exercise
Use an isolated environment and a clearly defined starting condition. Have someone follow the written instructions and record where they need help. Measure the full elapsed time and identify the restored data point.
Compare the outcome with the agreed objectives. If there is a gap, decide whether to improve the recovery process, change the architecture or revise the requirement with the responsible people. A rehearsal that reveals a gap has produced useful information; hiding the gap leaves the service exposed to a surprise.
Before you finish
- Workflow owner agrees the targets
- Access and decisions included in timing
- Restored data point verified
- Gaps assigned to owners
Technical reference
PostgreSQL: backup and restore
