Grafana: dashboard hỗ trợ quyết định
Grafana query data sources, biến dữ liệu thành panels và giúp người vận hành đi từ user impact tới nguyên nhân có thể kiểm chứng.
1. Grafana nằm ở đâu?
Prometheus ─┐
Loki ──────┼──> Grafana data sources ─> dashboards/alerts
Tempo ─────┤
SQL/Cloud ─┘Grafana thường không lưu metrics gốc. Panel “No data” có thể do query, variable, time range, permission, data source hoặc upstream ingestion.
2. Dashboard hierarchy
- Service overview: SLO, traffic, errors, latency, saturation.
- Breakdown: route, status, region/AZ, version, dependency.
- Resource detail: JVM, CPU, memory, pools, queues, DB/cache.
- Cross-signal links: logs/traces/profile cùng time range và labels.
3. Chọn visualization đúng
| Panel | Dùng khi | Bẫy |
|---|---|---|
| Time series | Trend, rate, latency, saturation. | Quá nhiều series, unit không nhất quán. |
| Stat/Gauge | Một giá trị hiện tại. | Mất context thời gian. |
| Table | Top dimensions, exact values, links. | Không ưu tiên thông tin. |
| Heatmap | Histogram/latency theo thời gian. | Query bucket sai. |
| State timeline | State/deployment/incident periods. | Không hợp continuous value. |
Đặt đúng unit. Threshold màu chỉ có nghĩa khi gắn SLO hoặc capacity action.
4. Variables
Variables chọn environment, cluster, namespace, service hoặc instance. Multi-value/all có thể tạo regex rất lớn.
label_values(up{job=~"$job"}, service)
sum by (service) (
rate(http_server_requests_seconds_count{service=~"$service"}[5m])
)Dashboard critical cần default an toàn, scope rõ và URL state chia sẻ được.
5. Transformations và expressions
Transformations join, rename, filter hoặc reduce dữ liệu sau query. Logic SLI/alert quan trọng nên nằm trong PromQL/recording rules có version control, tránh panel và alert tính khác nhau.
6. Performance
- Giới hạn time range, series và panels; không refresh nhanh hơn scrape interval vô ích.
- Dùng recording rules cho query nặng; xem query inspector.
- Giảm regex rộng, high-cardinality grouping và repeat panels.
- Chọn min interval phù hợp resolution màn hình.
- Đo Grafana và data sources: latency, errors, alert evaluation.
7. Grafana Alerting
Query/expression → evaluation
→ Normal / Pending / Alerting / NoData / Error
→ notification policy → contact pointRule cần rõ behavior khi NoData và Error. Pending period giảm noise; recovery threshold tránh flapping. Labels dùng routing/grouping; annotations chứa impact, dashboard và runbook.
8. Provisioning và dashboard as code
Data sources, dashboards, folders và alerts production nên được provision/version control. Dùng stable UID, render và verify sau thay đổi.
- Folder/RBAC và least privilege.
- Không nhúng secret vào dashboard JSON/query.
- Backup database/config/plugins và test restore/upgrade.
- Annotate deploy, feature flag và incident.
9. Review checklist
- Overview bắt đầu bằng user-centric SLI/SLO, không phải CPU.
- Mỗi panel có title, unit, legend và scope rõ.
- Numerator/denominator và grouping đúng semantics.
- Variables không tạo query explosion.
- No data, null, zero và stale data được phân biệt.
- Có drill-down tới logs, traces và runbook.
- Dashboard/rules có owner và version control.
- Load time đủ thấp để dùng trong incident.
10. Debug panel sai
- Kiểm tra time range, timezone, refresh và variables.
- Mở query inspector xem query/response/error.
- Chạy selector đơn giản tại Explore trước khi aggregate.
- Kiểm tra labels, unit, transformations và overrides.
- Kiểm tra counter reset, missing samples và cách dùng
rate. - So dashboard version với deploy annotation.
Prometheus collects and stores time-series metrics and evaluates PromQL rules. Grafana queries Prometheus and other data sources to build dashboards and can also evaluate managed alerts. I keep alert ownership explicit, version dashboards and rules, and design the overview around SLOs before drilling into RED, USE, logs, and traces.