Durability, quorum queues và recovery
Durable topology, persistent message, publisher confirms và replicated queue bảo vệ các lớp failure khác nhau. Không có một flag duy nhất biến toàn bộ pipeline thành “không mất message”.
1. Durability stack và scope
| Cơ chế | Bảo vệ | Không bảo vệ |
|---|---|---|
| Durable exchange/queue | Topology qua broker restart | Transient messages, node loss nếu không replica |
| Persistent delivery mode | Yêu cầu broker lưu message | Unroutable publish, consumer side effect |
| Publisher confirm | Producer quan sát broker acceptance | Consumer DB commit |
| Quorum replication | Queue log qua một số node failures | Majority loss, operator delete, logical corruption |
| Consumer ack + dedupe | Redelivery và duplicate-safe processing | External side effect không idempotent |
Disk flush, filesystem/hardware behavior, queue type và confirm timing quyết định guarantee thực tế. Đọc semantics của version đang vận hành và test bằng failure injection; tên “persistent” không tự chứng minh RPO bằng 0 trong mọi crash.
2. Quorum queue replication model
Quorum queue dùng replicated log theo Raft với một leader và members trên nhiều cluster nodes. Writes đi qua leader và cần quorum/majority để commit. Minority partition không tự nhận writes, ưu tiên consistency/safety hơn availability để tránh split brain.
Replica count thường là số lẻ và phải trải qua failure domains thật. Ba members chịu được một member unavailable; năm chịu được hai nhưng tăng disk/network/write amplification. Thêm nhiều replicas không làm throughput tăng tuyến tính và có thể làm latency cao hơn.
3. Leader placement, capacity và queue choice
Mỗi queue có leader; client traffic và replication tập trung qua leader. Nếu nhiều hot queues có leader trên một node, CPU, disk và network lệch dù replica count cân bằng. Theo dõi leader distribution và rebalance bằng procedure được hỗ trợ.
Quorum queues ưu tiên replicated safety và có workload assumptions riêng. Queue ngắn throughput rất cao, payload lớn, backlog dài hoặc số queue lớn phải benchmark với chính version/storage. Không suy diễn từ classic queue hoặc synthetic test không có confirms/consumers.
- Capacity tính cả replicas, retention/backlog và compaction/maintenance overhead.
- Disk free threshold cần headroom cho recovery và re-replication.
- Consumer locality không bỏ được leader/replication cost.
- Queue type là architectural choice; migration cần compatibility và drain/cutover plan.
4. Network partitions và availability
Cluster membership, metadata và individual queue consensus là các lớp liên quan nhưng không đồng nhất. Khi partition làm quorum queue mất majority, queue dừng tiến triển cho tới khi majority được phục hồi; đây là expected safety behavior, không phải retry vô hạn ở client.
Client retry phải có backoff/jitter và deadline để không tạo thundering herd khi leader election hoặc node recovery. Load balancer/DNS chỉ giúp tìm node reachable; chúng không tạo queue majority hoặc làm unknown publish outcome chắc chắn.
5. Client recovery và operation uncertainty
Automatic recovery có thể tái tạo connection, channel, topology và consumer. Nó không biết chắc operation đúng lúc disconnect đã commit hay chưa. Producer có thể duplicate khi retry; delivery đang xử lý có thể redeliver; confirms/acks scoped theo channel cũ không thể tiếp tục trên channel mới.
Declarations phải idempotent và tương thích. Exclusive/auto-delete resources dễ race giữa cleanup server và redeclare client; server-generated names và recovery ordering giảm collision. Application cần stable message ID, outbox/deduplication và explicit readiness trong thời gian topology chưa phục hồi.
6. Backup, upgrades và recovery evidence
Cluster replication không thay backup: delete, bad policy, credential compromise hoặc logical error có thể lan tới replicas. Backup definitions, policies, users/permissions và dữ liệu theo khả năng hỗ trợ; quan trọng hơn là restore drill với RPO/RTO đo được.
Rolling upgrade phải theo official compatibility matrix và supported version path. Theo dõi unavailable queues, quorum health, leader changes, node alarms, disk fsync/latency, confirm latency/nacks, replica catch-up và recovery duration.