Part 14 · Observability & Production Debugging
14.1 / Postmortem

Postmortem, action quality và telemetry governance

Blameless không nghĩa thiếu accountability; tập trung conditions/defenses/system changes với owner và verification.

Postmortem structure

Impact/duration/SLIs, timeline, detection/response, trigger, contributing conditions, why defenses failed, what went well, lessons, actions. Distinguish root-cause label from causal graph; human action occurs within system incentives/tools.

Action items

Specific owner/deadline/priority/verification. Prefer eliminate/reduce blast/detect/automate/runbook over “be careful/training” alone. Track closure and effectiveness; recurring incidents show actions weak.

Telemetry quality

Schema/semantic convention ownership, instrumentation tests, dashboards-as-code, alert reviews, cardinality budgets, sampling/retention policies and pipeline SLO. Detect missing labels, clock skew, exporter drops and stale dashboards.

Cost

Volume × retention × index/cardinality. Tier logs, sample traces, aggregate metrics/profiles, filter low-value data—but preserve compliance/security and incident-critical signals. Cost allocation per team/service.

Runbooks/game days

Runbook trigger/preconditions/diagnostics/mitigations/rollback/escalation. Game days validate detection, access, backup/failover and communication. Update dashboards/runbooks from real use.

Postmortem Culture · Observability Primer
← IncidentCâu hỏi →