Postmortem, action quality và telemetry governance
Blameless không nghĩa thiếu accountability; tập trung conditions/defenses/system changes với owner và verification.
Postmortem structure
Impact/duration/SLIs, timeline, detection/response, trigger, contributing conditions, why defenses failed, what went well, lessons, actions. Distinguish root-cause label from causal graph; human action occurs within system incentives/tools.
Action items
Specific owner/deadline/priority/verification. Prefer eliminate/reduce blast/detect/automate/runbook over “be careful/training” alone. Track closure and effectiveness; recurring incidents show actions weak.
Telemetry quality
Schema/semantic convention ownership, instrumentation tests, dashboards-as-code, alert reviews, cardinality budgets, sampling/retention policies and pipeline SLO. Detect missing labels, clock skew, exporter drops and stale dashboards.
Cost
Volume × retention × index/cardinality. Tier logs, sample traces, aggregate metrics/profiles, filter low-value data—but preserve compliance/security and incident-critical signals. Cost allocation per team/service.
Runbooks/game days
Runbook trigger/preconditions/diagnostics/mitigations/rollback/escalation. Game days validate detection, access, backup/failover and communication. Update dashboards/runbooks from real use.