Part 10 · System Design · 10.1.09

Failure, security, observability và rollout

Một sơ đồ HA chỉ trở thành reliability design khi ta mô tả được cách phát hiện lỗi, hành vi degraded, cách phục hồi và cách xác minh rằng recovery thực sự hoạt động.

Mental model từ source BK: với mỗi dependency, đừng chỉ hỏi “có replica không?”. Hãy hỏi failure xảy ra thế nào → blast radius là gì → detect bằng gì → hệ thống degrade/fallback ra sao → retry có an toàn không → phục hồi đến mức nào → đã test recovery chưa?

1. Failure matrix

Source BK yêu cầu xét tối thiểu các failure mode: timeout, overload, data corruption, node loss, AZ loss và region loss. Với từng dependency, cần ghi rõ blast radius, detection, fallback, retry/idempotency, RPO/RTO và test. Backup không được xem là giải pháp hoàn chỉnh nếu restore chưa từng được kiểm thử.

Failure modeĐiều cần trả lờiVí dụ degraded / recovery behavior
Timeout / network uncertaintyClient có biết request thất bại hay chỉ không biết kết quả? Timeout budget ở từng hop là bao nhiêu?Retry có backoff + jitter; operation có idempotency key; tránh retry storm.
Overload / saturationQueue, thread pool, connection pool, CPU, memory hay downstream nào bão hòa trước?Admission control, shed load, bounded queue, degrade tính năng phụ.
Data corruptionLỗi logical hay physical? Có checksum/version/audit để phát hiện không? Replica có thể nhân bản corruption không?Quarantine dữ liệu, restore từ bản sạch, reconciliation theo source of truth.
Node lossService còn đủ capacity sau khi mất một node không?Health check loại node; traffic chuyển sang healthy instances; auto-healing.
AZ lossCompute, DB, cache, broker có thật sự độc lập AZ không?Failover trong region; đảm bảo remaining AZ có headroom.
Region lossTraffic steering, data replication, DNS/control plane và credentials phụ thuộc region nào?Failover theo runbook; chấp nhận RPO/RTO đã định nghĩa; verify trước khi failback.

RPO, RTO và blast radius

Backup ≠ recovery. Một backup “tồn tại” chưa chứng minh rằng nó giải được incident. Cần restore drill định kỳ, đo thời gian phục hồi, kiểm tra tính toàn vẹn dữ liệu và xác nhận ứng dụng thật sự khởi động/đọc được dữ liệu sau restore.
Bổ sung production detail: AWS Reliability Pillar nhấn mạnh automatic recovery và test recovery procedures. Vì vậy failure matrix nên nối trực tiếp sang runbook/game day thay vì chỉ nằm trong tài liệu thiết kế.

2. Multi-region

Source BK phân biệt rõ trade-off: active-passive đơn giản hơn về conflict nhưng phải trả giá bằng failover path và RTO; active-active có thể cải thiện latency/availability nhưng làm routing, consistency và conflict phức tạp hơn. Failback thường nguy hiểm không kém failover. Ngoài technical goals, data residency có thể quyết định topology.

TopologyĐiểm mạnhĐổi lạiCâu hỏi review
Active-passiveConflict model đơn giản; source of truth rõ hơn.Standby có thể lạnh; failover cần automation/runbook; RTO thường lớn hơn.Standby có đủ capacity? Data lag bao nhiêu? DNS/routing chuyển trong bao lâu?
Active-activePhục vụ gần user hơn; giảm phụ thuộc một region.Conflict resolution, write ownership, consistency, duplicate processing và split-brain khó hơn.Write conflict giải bằng rule nào? Ordering scope? Global invariant nào không được phá?

Failover và failback là hai bài toán riêng

Failover thường chuyển traffic khỏi vùng lỗi; failback đưa traffic và quyền ghi trở lại topology bình thường. Sau một incident, hai region có thể có trạng thái khác nhau, replication backlog hoặc dữ liệu cần reconciliation. Vì vậy “đã failover được” không đồng nghĩa “có thể failback ngay”.

3. Security và abuse

Source BK yêu cầu bắt đầu bằng threat model: assets, actors và trust boundaries. Sau đó mới kiểm tra controls như authentication, authorization, encryption, secret/key rotation, least privilege, rate/quota, audit, PII minimization và deletion.

AreaReview questionFailure nếu bỏ qua
AuthenticationAi/điều gì đang gọi? Credential lifecycle và rotation thế nào?Stolen credential trở thành access lâu dài.
AuthorizationActor này được phép làm gì trên resource cụ thể?Authenticated nhưng vẫn đọc/sửa dữ liệu không thuộc quyền.
Least privilegeService account có quyền tối thiểu cho workload không?Một compromise lan thành blast radius lớn.
Rate / quotaGiới hạn theo IP, account, tenant, API key hay resource?Abuse hoặc noisy neighbor làm cạn capacity.
AuditAi đã đổi quyền, xóa dữ liệu, rotate key, approve action?Không điều tra được incident hoặc compliance gap.
PII lifecycleThật sự cần lưu gì? retention/deletion/export được thực thi ra sao?Tăng hậu quả breach và chi phí compliance.

Abuse case phải theo domain

Bổ sung production detail: OWASP ASVS là checklist verification cho technical security controls. Stable ASVS 5.0.0 được phát hành năm 2025; khi ghi requirement vào design/review nên pin version để tránh identifier thay đổi giữa các phiên bản.

4. Observability

Source BK yêu cầu đo từ SLI của user journey, kết hợp RED/USE metrics, structured logs và traces. Mỗi queue/cache/replica cần signal về saturation hoặc freshness. Alert phải actionable, ưu tiên theo burn rate; với workflow async cần correlation ID và audit trạng thái để truy vết xuyên nhiều bước.

Đo symptom trước, cause sau

LayerVí dụ signalÝ nghĩa
User journey / SLISuccess rate, end-to-end latency, freshnessUser có nhận đúng kết quả trong thời gian hứa hẹn không?
REDRate, Errors, DurationPhù hợp request/service-facing workloads.
USEUtilization, Saturation, ErrorsPhù hợp resource như CPU, disk, network, pool.
QueueOldest message age, backlog, consume rateBacklog lớn chưa chắc nguy hiểm bằng message age vượt SLO.
CacheHit ratio, eviction, memory pressure, stale agePhân biệt hiệu quả cache với saturation/freshness.
ReplicaReplication lag, apply error, last successful syncCho biết read freshness và recovery readiness.

Google SRE mô tả bốn “golden signals” cho user-facing systems là latency, traffic, errors và saturation. Đây là một baseline tốt, nhưng source BK nhấn mạnh rằng distributed workflow cần thêm domain signal như queue age, replica freshness và workflow state.

Async workflow: correlation và state audit

Với flow như API → queue → worker → downstream, chỉ nhìn request log ở API là không đủ. Correlation ID / trace context nên đi qua message metadata; workflow state cần có timestamp và stable business/event ID để phân biệt “đang chờ”, “retry”, “đã xử lý”, “duplicate” và “unknown outcome”.

Alert actionable, không alert vì metric “trông xấu”. Alert tốt phải chỉ ra điều khẩn cấp, có user impact hoặc imminent impact và có hành động rõ. Burn-rate alert giúp ưu tiên khi error budget đang bị tiêu quá nhanh thay vì dùng một threshold tĩnh cho mọi tình huống.

5. Rollout và rollback

Source BK yêu cầu thiết kế rollout bằng canary/blue-green, feature flag, mixed-version contracts, expand-contract schema và rollback. Điểm quan trọng nhất: rollback binary không tự hoàn tác data/schema; nhiều incident cần forward fix hoặc reconciliation thay vì “deploy bản cũ”.

TechniqueDùng để giảm rủi ro gì?Điểm cần nhớ
CanaryGiới hạn blast radius khi release.Canary cohort phải đại diện workload; success criteria gắn SLI.
Blue-greenChuẩn bị environment mới và chuyển traffic có kiểm soát.Data store/shared state có thể khiến “switch back” không sạch.
Feature flagTách code deployment khỏi feature activation.Flag cũng là configuration state; cần ownership, cleanup và audit.
Mixed-version contractCho phép rolling deploy khi old/new instances cùng tồn tại.API/event/schema phải backward/forward compatible trong rollout window.
Expand-contractThay schema/data model mà không lockstep deploy.Expand trước, migrate/read-both/write-both nếu cần, contract chỉ sau khi old path biến mất.

Rollback decision

6. Review checklist tổng hợp

Official references từ source BK: AWS Reliability Principles · OWASP ASVS · Google SRE — Monitoring Distributed Systems

Version note: OWASP ASVS stable 5.0.0 (released 2025). AWS/Google references are living documentation; verify organization-specific controls and cloud/service behavior against the version actually deployed.