Part 14 · Observability & Production Debugging
Part 14 / 14.3

8 lab observability và production debugging

Mỗi lab phải lưu hypothesis, evidence, mitigation, verification và learning; dashboard đẹp nhưng không hỗ trợ quyết định chưa phải kết quả.

Lab 01 · SLI/SLO và burn-rate alert

Định nghĩa availability/latency SLI từ user journey, xử lý valid events, đặt SLO/error budget; mô phỏng fast/slow burn, viết alert có owner, severity, runbook và test false positive.

Lab 02 · Prometheus và Grafana

Instrument RED/USE bằng counter, gauge và histogram; cấu hình scrape, viết PromQL/recording rule và dựng Grafana dashboard theo SLO → RED → USE. Cố ý tạo unbounded label và query/panel nặng; đo series, memory và latency rồi sửa cardinality, buckets, variables và refresh interval. Tạo một alert có NoData/Error policy, routing, runbook và kiểm thử notification.

Lab 03 · Structured logging và audit

Thiết kế event schema, correlation, error taxonomy, masking và retention. Chứng minh không log token/PII; tách audit trail, kiểm tra access/integrity và query một failure xuyên services.

Lab 04 · OpenTelemetry distributed trace

Instrument HTTP, DB và async message; propagate context, baggage có kiểm soát, resource attributes. So head/tail sampling, làm đứt context rồi tìm và sửa đoạn trace mất.

Lab 05 · JVM CPU/memory/lock

Tạo ba faults: CPU hot loop, retained memory, lock contention. Dùng OS metrics, JFR, GC logs, nhiều thread dumps và heap histogram/dump khi an toàn để phân biệt; lưu before/after evidence.

Lab 06 · Tomcat và dependency saturation

Load test connector với giới hạn threads/accept queue và DB pool nhỏ; quan sát active/queued/rejected, acquire time, latency. Thử tăng từng pool, chứng minh bottleneck dịch chuyển và chọn bounded capacity.

Lab 07 · Retry storm và queue lag

Inject downstream latency/error và slow consumer; quan sát retry amplification, deadlines, lag/partition skew. Áp dụng timeout budget, backoff+jitter, circuit/load shedding và verify recovery không thundering herd.

Lab 08 · Game day và postmortem

Chạy incident có commander/ops/comms/scribe: alert → triage → mitigation → recovery. Ghi timeline/decision log, status update, postmortem blameless và actions có owner/deadline/test.

Rubric

MứcTiêu chí
1Thu thập được telemetry và tìm symptom.
2Liên kết SLO, hypotheses và cross-signal evidence.
3 · SeniorMitigation an toàn, capacity/privacy/cost guardrails và learning được kiểm chứng.
← Câu hỏiTự đánh giá →