Sentinel, Cluster và failover
Sentinel thêm high availability cho một primary dataset; Redis Cluster thêm sharding cùng failover. Cả hai đều dựa asynchronous replication, nên availability không đồng nghĩa zero data loss hoặc strong consistency.
Sentinel: monitoring, discovery và failover
Một Sentinel deployment thực hiện bốn vai trò: health monitoring, notification, service discovery cho primary hiện tại và automatic failover. Nhiều Sentinel processes phải nằm trên failure domains độc lập; chạy ba processes trên cùng một host không tạo quorum chịu được host failure.
| Trạng thái/quyết định | Ý nghĩa |
|---|---|
| SDOWN — subjective down | Một Sentinel tự đánh giá instance không phản hồi trong down-after-milliseconds. |
| ODOWN — objective down | Đủ số Sentinel theo configured quorum cùng báo primary down. |
| Leader election / authorization | Một Sentinel cần quyền từ majority của Sentinel population đã biết để tiến hành failover; quorum và majority phục vụ hai bước khác nhau. |
| Discovery | Client hỏi Sentinels để tìm primary mới và phải reconnect/retry khi topology đổi. |
Quorum quá thấp làm false failover dễ hơn; timeout quá ngắn nhạy với pause hoặc transient network. Quorum cao nhưng không đủ majority Sentinel đang sống có thể detect ODOWN mà không hoàn tất failover.
Sentinel failover semantics
detect ODOWN
=> elect Sentinel leader
=> select replica candidate
=> promote candidate
=> reconfigure remaining replicas
=> publish new primary to clients
=> reconfigure old primary as replica when it returns
Replica selection cân nhắc replica priority, connectivity và replication progress. Vì replication asynchronous, acknowledged writes chưa đến candidate có thể mất khi promotion. Old primary ở network partition có thể tiếp tục nhận writes từ clients trỏ trực tiếp vào nó; các writes đó không tự merge vào primary mới.
Redis Cluster và 16,384 hash slots
Redis Cluster chia keyspace thành 16,384 hash slots. Mỗi primary node sở hữu một tập slots; replicas theo primary để failover. Nodes trao đổi topology và failure information qua cluster bus/gossip, còn client gửi commands trực tiếp tới node sở hữu slot.
key --CRC16 / hash tag--> slot 0..16383 --slot map--> primary node
wrong stable owner => MOVED slot host:port
slot đang migration => ASK slot host:port + ASKING
MOVED cho biết mapping ổn định đã đổi và client nên refresh slot cache. ASK là redirection tạm trong migration; client gửi ASKING tới target cho command kế tiếp nhưng không thay slot map vĩnh viễn. Cluster-aware client cần xử lý redirection, topology refresh, reconnect và bounded retry.
Multi-key operations và hash tags
Multi-key commands, transactions và scripts/functions chỉ có thể thao tác keys cùng hash slot khi command yêu cầu co-location. Phần đầu tiên nằm trong {...} là hash tag dùng để tính slot, ví dụ cart:{user-42} và cart-lock:{user-42}.
- Hash tag phải phản ánh atomicity boundary thật, không gom toàn tenant vào một slot nếu tenant có traffic lớn.
- Key schema là một phần của sharding contract; đổi schema cần compatibility và migration plan.
- Cross-slot aggregate thường cần fan-out tại application, precomputed structure hoặc data-model redesign.
- Một hot key vẫn nằm trên một primary; thêm shards không chia tải của chính key đó.
Resharding và capacity
Resharding chuyển slots online giữa primaries. Trong lúc migrate, source đánh dấu slot migrating và target đánh dấu importing; keys được chuyển dần, còn client phải xử lý ASK/MOVED. Theo dõi migration progress, latency, network, source/target memory và retries; tránh chạy cùng lúc với persistence rewrite hoặc node recovery nếu headroom thấp.
Capacity-plan theo memory gồm overhead, peak operations/second, bandwidth, hot-key distribution, per-slot skew và replica/failover headroom. Rebalancing slots chỉ hiệu quả khi load tương đối phân bố theo keys; một sorted set hoặc stream khổng lồ không tự được chia nhỏ.
Cluster availability và network partitions
Primary failure cần được majority của primary nodes đánh dấu để replica được authorize promotion. Phía minority của network partition cuối cùng ngừng phục vụ writes sau timeout thay vì tiếp tục tạo hai writable histories. cluster-node-timeout cân bằng detection speed với false positives khi network jitter hoặc event-loop stalls.
Topology nên phân bố primary và replica across hosts/zones để một failure domain không lấy cả hai. Replica migration có thể cải thiện coverage, nhưng không thay explicit placement validation. Nếu mỗi primary chỉ có một replica và chúng cùng zone, cluster có nhiều nodes nhưng vẫn mất slot availability khi zone hỏng.
Sentinel hay Cluster?
| Tiêu chí | Sentinel | Redis Cluster |
|---|---|---|
| Mục tiêu | HA cho một primary dataset. | Sharding/scale-out cùng HA. |
| Keyspace | Toàn dataset trên primary và replicas. | Chia 16,384 slots qua nhiều primaries. |
| Client | Sentinel discovery và reconnect. | Slot-aware routing, MOVED/ASK và topology refresh. |
| Multi-key | Không có cross-slot constraint. | Related keys phải cùng slot cho atomic multi-key operations. |
| Scaling writes/memory | Giới hạn bởi một primary. | Mở rộng bằng thêm primaries và resharding, nếu workload shard được. |
| Operational cost | Thấp hơn nhưng vẫn cần quorum topology. | Cao hơn: slot balance, reshard, cluster-aware clients và failure coverage. |
Chọn Sentinel khi dataset và write throughput vừa một node nhưng cần automated failover. Chọn Cluster khi một primary không đủ memory/throughput hoặc cần horizontal shard ownership. Không chọn Cluster chỉ vì “production phải có cluster”; complexity phải giải quyết một capacity hoặc isolation requirement thật.
Failure drill và observability
- Đo replication lag và xác định acknowledged-write loss có thể chấp nhận.
- Kill primary, partition network và pause node; quan sát detection, election, promotion và client recovery.
- Xác minh old primary không còn nhận production traffic và rejoin đúng role.
- Với Cluster, test MOVED/ASK, reshard under load, replica promotion và mất toàn một zone.
- Đo error budget: failed requests, duplicate retries, stale reads, lost writes và time-to-recovery.
- Kiểm tra backup/restore riêng; Sentinel/Cluster replicas không phải backup.