Observability and incident response
Status: Outline. Expand the operational material around diagnosis and decision-making.
To cover
Section titled “To cover”- Logs, metrics, traces, profiles, and correlation
- SLIs, SLOs, SLAs, error budgets, and alert ownership
- Symptoms versus diagnostic signals
- Triage, containment, mitigation, communication, and recovery
- Blameless review, contributing conditions, and durable follow-up
- Backup verification, restore exercises, and disaster recovery objectives
Interview prompts
Section titled “Interview prompts”- Which signals distinguish PHP saturation, database lock wait, and downstream latency?
- What makes an incident action item more useful than “be more careful”?