Part 08 · RabbitMQ · 8.1.03

Durability, quorum queues và recovery

Durable topology, persistent message, publisher confirms và replicated queue bảo vệ các lớp failure khác nhau. Không có một flag duy nhất biến toàn bộ pipeline thành “không mất message”.

Guarantee stack: topology sống qua restart + delivery được lưu + replica majority còn hoạt động + producer nhận confirm + consumer idempotent. Bỏ một tầng sẽ mở một failure window tương ứng.

1. Durability stack và scope

Cơ chếBảo vệKhông bảo vệ
Durable exchange/queueTopology qua broker restartTransient messages, node loss nếu không replica
Persistent delivery modeYêu cầu broker lưu messageUnroutable publish, consumer side effect
Publisher confirmProducer quan sát broker acceptanceConsumer DB commit
Quorum replicationQueue log qua một số node failuresMajority loss, operator delete, logical corruption
Consumer ack + dedupeRedelivery và duplicate-safe processingExternal side effect không idempotent

Disk flush, filesystem/hardware behavior, queue type và confirm timing quyết định guarantee thực tế. Đọc semantics của version đang vận hành và test bằng failure injection; tên “persistent” không tự chứng minh RPO bằng 0 trong mọi crash.

2. Quorum queue replication model

Quorum queue dùng replicated log theo Raft với một leader và members trên nhiều cluster nodes. Writes đi qua leader và cần quorum/majority để commit. Minority partition không tự nhận writes, ưu tiên consistency/safety hơn availability để tránh split brain.

Replica count thường là số lẻ và phải trải qua failure domains thật. Ba members chịu được một member unavailable; năm chịu được hai nhưng tăng disk/network/write amplification. Thêm nhiều replicas không làm throughput tăng tuyến tính và có thể làm latency cao hơn.

Replication placement: ba replicas trên ba nodes nhưng cùng host, rack hoặc availability zone vẫn có correlated failure. Broker topology phải map vào infrastructure failure domains, không chỉ đếm process.

3. Leader placement, capacity và queue choice

Mỗi queue có leader; client traffic và replication tập trung qua leader. Nếu nhiều hot queues có leader trên một node, CPU, disk và network lệch dù replica count cân bằng. Theo dõi leader distribution và rebalance bằng procedure được hỗ trợ.

Quorum queues ưu tiên replicated safety và có workload assumptions riêng. Queue ngắn throughput rất cao, payload lớn, backlog dài hoặc số queue lớn phải benchmark với chính version/storage. Không suy diễn từ classic queue hoặc synthetic test không có confirms/consumers.

4. Network partitions và availability

Cluster membership, metadata và individual queue consensus là các lớp liên quan nhưng không đồng nhất. Khi partition làm quorum queue mất majority, queue dừng tiến triển cho tới khi majority được phục hồi; đây là expected safety behavior, không phải retry vô hạn ở client.

Client retry phải có backoff/jitter và deadline để không tạo thundering herd khi leader election hoặc node recovery. Load balancer/DNS chỉ giúp tìm node reachable; chúng không tạo queue majority hoặc làm unknown publish outcome chắc chắn.

5. Client recovery và operation uncertainty

Automatic recovery có thể tái tạo connection, channel, topology và consumer. Nó không biết chắc operation đúng lúc disconnect đã commit hay chưa. Producer có thể duplicate khi retry; delivery đang xử lý có thể redeliver; confirms/acks scoped theo channel cũ không thể tiếp tục trên channel mới.

Declarations phải idempotent và tương thích. Exclusive/auto-delete resources dễ race giữa cleanup server và redeclare client; server-generated names và recovery ordering giảm collision. Application cần stable message ID, outbox/deduplication và explicit readiness trong thời gian topology chưa phục hồi.

6. Backup, upgrades và recovery evidence

Cluster replication không thay backup: delete, bad policy, credential compromise hoặc logical error có thể lan tới replicas. Backup definitions, policies, users/permissions và dữ liệu theo khả năng hỗ trợ; quan trọng hơn là restore drill với RPO/RTO đo được.

Rolling upgrade phải theo official compatibility matrix và supported version path. Theo dõi unavailable queues, quorum health, leader changes, node alarms, disk fsync/latency, confirm latency/nacks, replica catch-up và recovery duration.

Review checklist: từng durability layer được nêu rõ; replica trải failure domains; majority-loss behavior được chấp nhận; client retry bounded; unknown outcome duplicate-safe; backups restore được; rolling procedure và rollback đã diễn tập.
Tài liệu: Quorum Queues · Reliability Guide · Network Partitions · Clustering Guide · Backup and Restore