Part 14 · Observability & Production Debugging
Part 14 / 14.2
36 câu hỏi observability và production debugging
Trả lời theo user impact → hypothesis → evidence → mitigation → verification, luôn nêu giới hạn của tín hiệu.
Signals, SLO và telemetry
1. Monitoring khác observability?
Monitoring kiểm tra known conditions; observability cho phép suy luận internal state từ outputs để điều tra cả unknown failure. Hệ thống cần cả hai.
2. Metrics, logs, traces và profiles bổ sung nhau thế nào?
Metrics phát hiện xu hướng, logs giải thích events, traces nối request path, profiles chỉ ra nơi CPU/allocation/time được tiêu; correlation context liên kết chúng.
3. SLI, SLO và SLA khác nhau?
SLI là phép đo service level, SLO là target nội bộ, SLA là cam kết có hậu quả kinh doanh/pháp lý.
4. Error budget dùng để làm gì?
Khoảng unreliability còn được phép giúp cân bằng release velocity và reliability; burn rate cho biết budget đang bị tiêu nhanh đến đâu.
5. Vì sao alert theo symptom?
Latency/error/availability phản ánh user impact và actionable hơn alert mọi cause; cause metrics vẫn hữu ích khi chẩn đoán.
6. Multi-window burn-rate alert?
Cửa sổ ngắn phát hiện burn nhanh, cửa sổ dài lọc spike; kết hợp severity với mức budget bị đe dọa.
7. RED và USE?
RED: rate/errors/duration cho request services. USE: utilization/saturation/errors cho resources; cả hai cần gắn với SLO.
8. Counter khác gauge và histogram?
Counter tăng tích lũy rồi rate; gauge lên/xuống; histogram phân phối theo buckets để aggregate quantile phía server.
9. Cardinality explosion là gì?
Label values không giới hạn như user/request ID tạo quá nhiều time series, tăng memory/cost và làm query chậm.
10. Client-side percentile có thể aggregate?
Percentile đã tính thường không merge đúng giữa instances; histogram buckets có thể sum rồi tính quantile với sai số bucket.
Prometheus và Grafana chuyên sâu
11. Prometheus và Grafana khác vai trò thế nào?
Prometheus scrape, lưu time series, chạy PromQL và rules. Grafana query Prometheus cùng data sources khác để visualize và có thể evaluate Grafana-managed alerts. Grafana không tự thu thập metrics ứng dụng.
12. Vì sao rate trước rồi mới sum counter?
Counter reset xảy ra riêng trên từng instance. rate xử lý reset theo series; sum trước có thể che reset và tạo spike hoặc kết quả sai.
13. Recording rule dùng khi nào?
Khi PromQL đắt hoặc dùng lặp lại cho dashboard/SLO/alert. Rule giảm query latency nhưng thêm độ trễ evaluation và cần versioning, naming, dependency review.
14. Panel No data bắt đầu debug ở đâu?
Kiểm tra time range, variables và query inspector; chạy selector đơn giản trong Explore; sau đó kiểm tra target scrape, labels, data-source permission và transformations. No data không đồng nghĩa value bằng zero.
15. Grafana variables có rủi ro gì?
Regex All/multi-value, label cardinality cao và repeat panels có thể tạo query explosion. Dashboard critical cần default/scope rõ và bounded dimensions.
16. Grafana-managed alert và Prometheus alert chọn thế nào?
Chọn một ownership model rõ theo data source, portability và operations. Prometheus rules gần metrics và Alertmanager; Grafana alerting thuận tiện cho multi-data-source. Không duplicate cùng condition nếu không có dedupe và silence strategy.
Logs, traces và profiling
17. Structured log cần trường gì?
Timestamp chuẩn, level, service/version/environment, event, trace/span hoặc correlation ID, outcome/error taxonomy; tránh secrets và PII.
18. Log level và sampling chọn thế nào?
Giữ errors/security/audit cần thiết, sample high-volume success theo policy; dynamic verbosity phải time-bound, audited và có cost/privacy guardrail.
19. Audit log khác application log?
Audit log ghi actor/action/resource/outcome để truy vết và có yêu cầu integrity/retention/access riêng; không nên trộn tùy tiện với debug log.
20. Trace context truyền qua đâu?
Qua standardized headers và message metadata; async fan-out có links/parent semantics, không nhét request ID vào metric label.
21. Head sampling vs tail sampling?
Head quyết định sớm, rẻ nhưng có thể bỏ rare errors; tail xem completed trace để giữ slow/error nhưng cần buffering và collector capacity.
22. Trace bị đứt do đâu?
Không propagate context qua HTTP/message/executor, instrumentation version lệch, sampling không thống nhất hoặc collector/export failure.
23. Continuous profiling mang lại gì?
So sánh CPU/allocation/lock theo thời gian và version; profile cần sampling overhead control và symbol/context đủ để giải thích.
24. CPU cao điều tra thế nào?
Xác nhận scope/user impact, correlate deploy/load, lấy process/thread CPU và JFR/profile, phân biệt compute, spin, GC, lock rồi thử mitigation có kiểm soát.
25. Memory leak khác traffic-driven memory?
Leak có retained set tăng qua các chu kỳ GC; load tăng có thể giảm khi traffic/queue/cache giảm. Dùng GC trend, class histogram và heap dump/dominator evidence.
26. Heap dump có rủi ro gì?
Có thể pause/IO lớn, chứa secrets/PII và làm đầy disk; cần capacity check, access/retention/encryption và ưu tiên histogram/JFR trước khi phù hợp.
Runtime, dependencies và incidents
27. Thread dump đọc dấu hiệu nào?
RUNNABLE hotspots, BLOCKED owner/contended monitor, WAITING/TIMED_WAITING theo pool/queue, deadlock và nhiều stacks giống nhau; lấy nhiều snapshots.
28. GC pause và allocation pressure phân biệt?
Xem allocation rate, heap-after-GC, pause/concurrent phase, promotion và CPU. Pause là symptom; root cause có thể object churn, leak, sizing hoặc collector behavior.
29. Tomcat maxThreads tăng có luôn tốt?
Không; tăng concurrency có thể đẩy saturation sang DB/downstream, tăng memory/context switch. Phải xét accept queue, active threads, latency và Little's Law.
30. Tomcat acceptCount và connection backlog?
Khi worker capacity bận, accepted requests chờ trong connector/OS queues đến giới hạn; đầy queue dẫn timeout/refusal. Các lớp backlog và timeout cần đo riêng.
31. Connection pool saturation biểu hiện gì?
Acquire wait/timeouts tăng, active gần max, downstream latency/transactions kéo dài; chỉ tăng pool có thể overload database.
32. Retry storm hình thành thế nào?
Timeout/retry không budget, nhiều layers retry đồng bộ làm traffic khuếch đại lúc dependency yếu; dùng deadline, bounded retry, backoff+jitter và load shedding.
33. Queue lag tăng nhưng consumer CPU thấp?
Có thể blocked I/O, partition imbalance, poison message, lock, downstream throttle hoặc consumer rebalance; cần per-partition/processing-stage evidence.
34. Bước đầu của incident commander?
Xác định impact/severity, thiết lập roles/channel/timeline, ưu tiên mitigation an toàn và giữ investigation hypotheses/evidence tách khỏi phỏng đoán.
35. Rollback không giảm lỗi nói lên gì?
Có thể config/data/dependency/capacity issue, rollback chưa hoàn tất hoặc telemetry lag. Verify version/traffic và mở rộng hypotheses thay vì rollback lặp.
36. Postmortem tốt gồm gì?
Blameless timeline, impact/detection/response, contributing system conditions, what worked/failed và actions có owner/deadline/verification.