Failure, security, observability và rollout
Một sơ đồ HA chỉ trở thành reliability design khi ta mô tả được cách phát hiện lỗi, hành vi degraded, cách phục hồi và cách xác minh rằng recovery thực sự hoạt động.
1. Failure matrix
Source BK yêu cầu xét tối thiểu các failure mode: timeout, overload, data corruption, node loss, AZ loss và region loss. Với từng dependency, cần ghi rõ blast radius, detection, fallback, retry/idempotency, RPO/RTO và test. Backup không được xem là giải pháp hoàn chỉnh nếu restore chưa từng được kiểm thử.
| Failure mode | Điều cần trả lời | Ví dụ degraded / recovery behavior |
|---|---|---|
| Timeout / network uncertainty | Client có biết request thất bại hay chỉ không biết kết quả? Timeout budget ở từng hop là bao nhiêu? | Retry có backoff + jitter; operation có idempotency key; tránh retry storm. |
| Overload / saturation | Queue, thread pool, connection pool, CPU, memory hay downstream nào bão hòa trước? | Admission control, shed load, bounded queue, degrade tính năng phụ. |
| Data corruption | Lỗi logical hay physical? Có checksum/version/audit để phát hiện không? Replica có thể nhân bản corruption không? | Quarantine dữ liệu, restore từ bản sạch, reconciliation theo source of truth. |
| Node loss | Service còn đủ capacity sau khi mất một node không? | Health check loại node; traffic chuyển sang healthy instances; auto-healing. |
| AZ loss | Compute, DB, cache, broker có thật sự độc lập AZ không? | Failover trong region; đảm bảo remaining AZ có headroom. |
| Region loss | Traffic steering, data replication, DNS/control plane và credentials phụ thuộc region nào? | Failover theo runbook; chấp nhận RPO/RTO đã định nghĩa; verify trước khi failback. |
RPO, RTO và blast radius
- RPO trả lời: tối đa có thể mất bao nhiêu dữ liệu khi phục hồi?
- RTO trả lời: tối đa mất bao lâu để khôi phục dịch vụ về mức chấp nhận được?
- Blast radius trả lời: lỗi này ảnh hưởng một request, một tenant, một shard, một AZ hay toàn bộ hệ thống?
2. Multi-region
Source BK phân biệt rõ trade-off: active-passive đơn giản hơn về conflict nhưng phải trả giá bằng failover path và RTO; active-active có thể cải thiện latency/availability nhưng làm routing, consistency và conflict phức tạp hơn. Failback thường nguy hiểm không kém failover. Ngoài technical goals, data residency có thể quyết định topology.
| Topology | Điểm mạnh | Đổi lại | Câu hỏi review |
|---|---|---|---|
| Active-passive | Conflict model đơn giản; source of truth rõ hơn. | Standby có thể lạnh; failover cần automation/runbook; RTO thường lớn hơn. | Standby có đủ capacity? Data lag bao nhiêu? DNS/routing chuyển trong bao lâu? |
| Active-active | Phục vụ gần user hơn; giảm phụ thuộc một region. | Conflict resolution, write ownership, consistency, duplicate processing và split-brain khó hơn. | Write conflict giải bằng rule nào? Ordering scope? Global invariant nào không được phá? |
Failover và failback là hai bài toán riêng
Failover thường chuyển traffic khỏi vùng lỗi; failback đưa traffic và quyền ghi trở lại topology bình thường. Sau một incident, hai region có thể có trạng thái khác nhau, replication backlog hoặc dữ liệu cần reconciliation. Vì vậy “đã failover được” không đồng nghĩa “có thể failback ngay”.
- Define source of truth trong từng phase của incident.
- Đo replication lag và divergence trước khi đổi write ownership.
- Throttling catch-up để recovery không làm sập region đang healthy.
- Chạy consistency/reconciliation checks trước khi đóng incident.
3. Security và abuse
Source BK yêu cầu bắt đầu bằng threat model: assets, actors và trust boundaries. Sau đó mới kiểm tra controls như authentication, authorization, encryption, secret/key rotation, least privilege, rate/quota, audit, PII minimization và deletion.
| Area | Review question | Failure nếu bỏ qua |
|---|---|---|
| Authentication | Ai/điều gì đang gọi? Credential lifecycle và rotation thế nào? | Stolen credential trở thành access lâu dài. |
| Authorization | Actor này được phép làm gì trên resource cụ thể? | Authenticated nhưng vẫn đọc/sửa dữ liệu không thuộc quyền. |
| Least privilege | Service account có quyền tối thiểu cho workload không? | Một compromise lan thành blast radius lớn. |
| Rate / quota | Giới hạn theo IP, account, tenant, API key hay resource? | Abuse hoặc noisy neighbor làm cạn capacity. |
| Audit | Ai đã đổi quyền, xóa dữ liệu, rotate key, approve action? | Không điều tra được incident hoặc compliance gap. |
| PII lifecycle | Thật sự cần lưu gì? retention/deletion/export được thực thi ra sao? | Tăng hậu quả breach và chi phí compliance. |
Abuse case phải theo domain
- URL shortener: malicious redirect, phishing, domain reputation, automated scanning/rate limit.
- Upload service: malware, content type spoofing, size bombs, archive bombs, unsafe preview/parser.
- Chat: spam, harassment, report/block, account farming, rate limit và moderation workflow.
4. Observability
Source BK yêu cầu đo từ SLI của user journey, kết hợp RED/USE metrics, structured logs và traces. Mỗi queue/cache/replica cần signal về saturation hoặc freshness. Alert phải actionable, ưu tiên theo burn rate; với workflow async cần correlation ID và audit trạng thái để truy vết xuyên nhiều bước.
Đo symptom trước, cause sau
| Layer | Ví dụ signal | Ý nghĩa |
|---|---|---|
| User journey / SLI | Success rate, end-to-end latency, freshness | User có nhận đúng kết quả trong thời gian hứa hẹn không? |
| RED | Rate, Errors, Duration | Phù hợp request/service-facing workloads. |
| USE | Utilization, Saturation, Errors | Phù hợp resource như CPU, disk, network, pool. |
| Queue | Oldest message age, backlog, consume rate | Backlog lớn chưa chắc nguy hiểm bằng message age vượt SLO. |
| Cache | Hit ratio, eviction, memory pressure, stale age | Phân biệt hiệu quả cache với saturation/freshness. |
| Replica | Replication lag, apply error, last successful sync | Cho biết read freshness và recovery readiness. |
Google SRE mô tả bốn “golden signals” cho user-facing systems là latency, traffic, errors và saturation. Đây là một baseline tốt, nhưng source BK nhấn mạnh rằng distributed workflow cần thêm domain signal như queue age, replica freshness và workflow state.
Async workflow: correlation và state audit
Với flow như API → queue → worker → downstream, chỉ nhìn request log ở API là không đủ. Correlation ID / trace context nên đi qua message metadata; workflow state cần có timestamp và stable business/event ID để phân biệt “đang chờ”, “retry”, “đã xử lý”, “duplicate” và “unknown outcome”.
5. Rollout và rollback
Source BK yêu cầu thiết kế rollout bằng canary/blue-green, feature flag, mixed-version contracts, expand-contract schema và rollback. Điểm quan trọng nhất: rollback binary không tự hoàn tác data/schema; nhiều incident cần forward fix hoặc reconciliation thay vì “deploy bản cũ”.
| Technique | Dùng để giảm rủi ro gì? | Điểm cần nhớ |
|---|---|---|
| Canary | Giới hạn blast radius khi release. | Canary cohort phải đại diện workload; success criteria gắn SLI. |
| Blue-green | Chuẩn bị environment mới và chuyển traffic có kiểm soát. | Data store/shared state có thể khiến “switch back” không sạch. |
| Feature flag | Tách code deployment khỏi feature activation. | Flag cũng là configuration state; cần ownership, cleanup và audit. |
| Mixed-version contract | Cho phép rolling deploy khi old/new instances cùng tồn tại. | API/event/schema phải backward/forward compatible trong rollout window. |
| Expand-contract | Thay schema/data model mà không lockstep deploy. | Expand trước, migrate/read-both/write-both nếu cần, contract chỉ sau khi old path biến mất. |
Rollback decision
- Nếu chỉ code/stateless behavior thay đổi và data contract chưa đổi, rollback binary có thể đủ.
- Nếu đã write dữ liệu theo schema mới, gửi side effect ra ngoài hoặc thay state machine, rollback có thể tạo inconsistency.
- Khi đó cần forward fix, compensating action hoặc reconciliation theo stable source of truth.
6. Review checklist tổng hợp
- Đã có failure matrix cho timeout, overload, corruption, node/AZ/region loss chưa?
- Mỗi failure có blast radius, detection, degraded behavior, retry/idempotency, RPO/RTO và recovery test chưa?
- Multi-region topology đã giải thích active-passive/active-active, conflict, residency, failover và failback chưa?
- Threat model đã xác định assets, actors, trust boundaries và abuse cases đặc thù domain chưa?
- Authn/authz, encryption, secret/key rotation, least privilege, quotas, audit và PII lifecycle có owner không?
- SLI có bám user journey? Queue/cache/replica có saturation/freshness signal?
- Alert có actionable và dựa trên user impact / error-budget burn thay vì metric noise không?
- Async flow có correlation ID, stable event/business ID và state audit?
- Rollout có canary/blue-green/flag, mixed-version contract và expand-contract schema?
- Rollback plan có tính đến data/schema/side effects và reconciliation?
Version note: OWASP ASVS stable 5.0.0 (released 2025). AWS/Google references are living documentation; verify organization-specific controls and cloud/service behavior against the version actually deployed.