Part 14 · Observability & Production Debugging
14.1 / Grafana

Grafana: dashboard hỗ trợ quyết định

Grafana query data sources, biến dữ liệu thành panels và giúp người vận hành đi từ user impact tới nguyên nhân có thể kiểm chứng.

1. Grafana nằm ở đâu?

Prometheus ─┐
Loki ──────┼──> Grafana data sources ─> dashboards/alerts
Tempo ─────┤
SQL/Cloud ─┘

Grafana thường không lưu metrics gốc. Panel “No data” có thể do query, variable, time range, permission, data source hoặc upstream ingestion.

2. Dashboard hierarchy

  1. Service overview: SLO, traffic, errors, latency, saturation.
  2. Breakdown: route, status, region/AZ, version, dependency.
  3. Resource detail: JVM, CPU, memory, pools, queues, DB/cache.
  4. Cross-signal links: logs/traces/profile cùng time range và labels.
Dashboard tốt trả lời: Có impact không? Bắt đầu khi nào? Scope nào? Dimension nào khác biệt? Điều tra tiếp ở đâu?

3. Chọn visualization đúng

PanelDùng khiBẫy
Time seriesTrend, rate, latency, saturation.Quá nhiều series, unit không nhất quán.
Stat/GaugeMột giá trị hiện tại.Mất context thời gian.
TableTop dimensions, exact values, links.Không ưu tiên thông tin.
HeatmapHistogram/latency theo thời gian.Query bucket sai.
State timelineState/deployment/incident periods.Không hợp continuous value.

Đặt đúng unit. Threshold màu chỉ có nghĩa khi gắn SLO hoặc capacity action.

4. Variables

Variables chọn environment, cluster, namespace, service hoặc instance. Multi-value/all có thể tạo regex rất lớn.

label_values(up{job=~"$job"}, service)

sum by (service) (
  rate(http_server_requests_seconds_count{service=~"$service"}[5m])
)

Dashboard critical cần default an toàn, scope rõ và URL state chia sẻ được.

5. Transformations và expressions

Transformations join, rename, filter hoặc reduce dữ liệu sau query. Logic SLI/alert quan trọng nên nằm trong PromQL/recording rules có version control, tránh panel và alert tính khác nhau.

6. Performance

7. Grafana Alerting

Query/expression → evaluation
→ Normal / Pending / Alerting / NoData / Error
→ notification policy → contact point

Rule cần rõ behavior khi NoDataError. Pending period giảm noise; recovery threshold tránh flapping. Labels dùng routing/grouping; annotations chứa impact, dashboard và runbook.

Tránh page trùng: cùng condition không nên vô thức tồn tại cả Prometheus rule + Alertmanager và Grafana-managed alert. Chọn ownership rõ.

8. Provisioning và dashboard as code

Data sources, dashboards, folders và alerts production nên được provision/version control. Dùng stable UID, render và verify sau thay đổi.

9. Review checklist

10. Debug panel sai

  1. Kiểm tra time range, timezone, refresh và variables.
  2. Mở query inspector xem query/response/error.
  3. Chạy selector đơn giản tại Explore trước khi aggregate.
  4. Kiểm tra labels, unit, transformations và overrides.
  5. Kiểm tra counter reset, missing samples và cách dùng rate.
  6. So dashboard version với deploy annotation.
Dashboards · Visualizations · Alerting · Provisioning
English interview answer:

Prometheus collects and stores time-series metrics and evaluates PromQL rules. Grafana queries Prometheus and other data sources to build dashboards and can also evaluate managed alerts. I keep alert ownership explicit, version dashboards and rules, and design the overview around SLOs before drilling into RED, USE, logs, and traces.

← PrometheusLogging →