Part 04 · PostgreSQL & Redis · 4.1.09

Sentinel, Cluster và failover

Sentinel thêm high availability cho một primary dataset; Redis Cluster thêm sharding cùng failover. Cả hai đều dựa asynchronous replication, nên availability không đồng nghĩa zero data loss hoặc strong consistency.


Sentinel: monitoring, discovery và failover

Một Sentinel deployment thực hiện bốn vai trò: health monitoring, notification, service discovery cho primary hiện tại và automatic failover. Nhiều Sentinel processes phải nằm trên failure domains độc lập; chạy ba processes trên cùng một host không tạo quorum chịu được host failure.

Trạng thái/quyết địnhÝ nghĩa
SDOWN — subjective downMột Sentinel tự đánh giá instance không phản hồi trong down-after-milliseconds.
ODOWN — objective downĐủ số Sentinel theo configured quorum cùng báo primary down.
Leader election / authorizationMột Sentinel cần quyền từ majority của Sentinel population đã biết để tiến hành failover; quorum và majority phục vụ hai bước khác nhau.
DiscoveryClient hỏi Sentinels để tìm primary mới và phải reconnect/retry khi topology đổi.

Quorum quá thấp làm false failover dễ hơn; timeout quá ngắn nhạy với pause hoặc transient network. Quorum cao nhưng không đủ majority Sentinel đang sống có thể detect ODOWN mà không hoàn tất failover.

Sentinel failover semantics

detect ODOWN
    => elect Sentinel leader
    => select replica candidate
    => promote candidate
    => reconfigure remaining replicas
    => publish new primary to clients
    => reconfigure old primary as replica when it returns

Replica selection cân nhắc replica priority, connectivity và replication progress. Vì replication asynchronous, acknowledged writes chưa đến candidate có thể mất khi promotion. Old primary ở network partition có thể tiếp tục nhận writes từ clients trỏ trực tiếp vào nó; các writes đó không tự merge vào primary mới.

Client là một phần của HA: hard-code primary address làm Sentinel gần như vô nghĩa. Client phải hỗ trợ Sentinel discovery, bounded reconnect, command retry theo idempotency và đóng stale pooled connections sau role change.

Redis Cluster và 16,384 hash slots

Redis Cluster chia keyspace thành 16,384 hash slots. Mỗi primary node sở hữu một tập slots; replicas theo primary để failover. Nodes trao đổi topology và failure information qua cluster bus/gossip, còn client gửi commands trực tiếp tới node sở hữu slot.

key --CRC16 / hash tag--> slot 0..16383 --slot map--> primary node

wrong stable owner  => MOVED slot host:port
slot đang migration => ASK slot host:port + ASKING

MOVED cho biết mapping ổn định đã đổi và client nên refresh slot cache. ASK là redirection tạm trong migration; client gửi ASKING tới target cho command kế tiếp nhưng không thay slot map vĩnh viễn. Cluster-aware client cần xử lý redirection, topology refresh, reconnect và bounded retry.

Multi-key operations và hash tags

Multi-key commands, transactions và scripts/functions chỉ có thể thao tác keys cùng hash slot khi command yêu cầu co-location. Phần đầu tiên nằm trong {...} là hash tag dùng để tính slot, ví dụ cart:{user-42}cart-lock:{user-42}.

Resharding và capacity

Resharding chuyển slots online giữa primaries. Trong lúc migrate, source đánh dấu slot migrating và target đánh dấu importing; keys được chuyển dần, còn client phải xử lý ASK/MOVED. Theo dõi migration progress, latency, network, source/target memory và retries; tránh chạy cùng lúc với persistence rewrite hoặc node recovery nếu headroom thấp.

Capacity-plan theo memory gồm overhead, peak operations/second, bandwidth, hot-key distribution, per-slot skew và replica/failover headroom. Rebalancing slots chỉ hiệu quả khi load tương đối phân bố theo keys; một sorted set hoặc stream khổng lồ không tự được chia nhỏ.

Cluster availability và network partitions

Primary failure cần được majority của primary nodes đánh dấu để replica được authorize promotion. Phía minority của network partition cuối cùng ngừng phục vụ writes sau timeout thay vì tiếp tục tạo hai writable histories. cluster-node-timeout cân bằng detection speed với false positives khi network jitter hoặc event-loop stalls.

Topology nên phân bố primary và replica across hosts/zones để một failure domain không lấy cả hai. Replica migration có thể cải thiện coverage, nhưng không thay explicit placement validation. Nếu mỗi primary chỉ có một replica và chúng cùng zone, cluster có nhiều nodes nhưng vẫn mất slot availability khi zone hỏng.

Sentinel hay Cluster?

Tiêu chíSentinelRedis Cluster
Mục tiêuHA cho một primary dataset.Sharding/scale-out cùng HA.
KeyspaceToàn dataset trên primary và replicas.Chia 16,384 slots qua nhiều primaries.
ClientSentinel discovery và reconnect.Slot-aware routing, MOVED/ASK và topology refresh.
Multi-keyKhông có cross-slot constraint.Related keys phải cùng slot cho atomic multi-key operations.
Scaling writes/memoryGiới hạn bởi một primary.Mở rộng bằng thêm primaries và resharding, nếu workload shard được.
Operational costThấp hơn nhưng vẫn cần quorum topology.Cao hơn: slot balance, reshard, cluster-aware clients và failure coverage.

Chọn Sentinel khi dataset và write throughput vừa một node nhưng cần automated failover. Chọn Cluster khi một primary không đủ memory/throughput hoặc cần horizontal shard ownership. Không chọn Cluster chỉ vì “production phải có cluster”; complexity phải giải quyết một capacity hoặc isolation requirement thật.

Failure drill và observability

  1. Đo replication lag và xác định acknowledged-write loss có thể chấp nhận.
  2. Kill primary, partition network và pause node; quan sát detection, election, promotion và client recovery.
  3. Xác minh old primary không còn nhận production traffic và rejoin đúng role.
  4. Với Cluster, test MOVED/ASK, reshard under load, replica promotion và mất toàn một zone.
  5. Đo error budget: failed requests, duplicate retries, stale reads, lost writes và time-to-recovery.
  6. Kiểm tra backup/restore riêng; Sentinel/Cluster replicas không phải backup.
Senior answer: mô tả topology chưa đủ. Câu trả lời tốt luôn nối failure detection → quorum/majority → promotion → client routing → consistency window → old-primary handling → drill evidence.
Nguồn tham khảo