Part 08 · RabbitMQ · 8.1.05

Backpressure, capacity và production operations

Queue hấp thụ burst nhưng không tạo processing capacity. Backlog cuối cùng biến thành business latency, disk consumption, redelivery volume và recovery time.

Operating invariant: nếu arrival rate vượt sustainable service rate trong thời gian dài, hệ thống phải giảm ingress, tăng capacity hoặc chấp nhận loss/degradation. Tăng queue limit chỉ dời thời điểm failure.

1. Flow control từ consumer đến publisher

Consumer-side backpressure đến từ prefetch, worker concurrency và processing rate. Broker-side flow control giảm tốc connection đang publish khi queues/disk không theo kịp. Memory hoặc disk alarms có thể block publishers để bảo vệ node; client phải quan sát blocked/unblocked events, confirm latency và publish timeout.

LayerControlFailure khi cấu hình sai
ConsumerPrefetch, concurrency, deadlineUnacked cao, unfairness, overload downstream
QueueMax length/bytes, TTL, overflow policySilent business loss hoặc disk growth
BrokerFlow control, memory/disk alarmsPublisher stall, cluster instability
ProducerRate limit, bounded buffer, retry budgetMemory blow-up, retry storm
BusinessAdmission/load sheddingKhông phân biệt critical và optional work
Client buffer không phải queue durable. Khi broker block, unbounded in-memory buffering ở producer có thể làm process OOM. Publish path cần bounded concurrency/buffer và failure response rõ cho upstream.

2. Backlog math và latency

Gọi arrival rate là λ, service rate mỗi worker là μ và số workers hữu hiệu là c. Khi λ ≥ cμ lâu dài, backlog tăng không giới hạn. Ngay cả utilization gần 100%, processing variance và downstream pauses làm tail latency tăng mạnh; cần capacity headroom cho burst và failures.

Queue depth phải đọc cùng publish/ack rate và message age. 100.000 messages có thể bình thường nếu drain vài giây, hoặc là incident nếu oldest age vượt SLA. Payload size quyết định byte backlog, memory/disk I/O và network recovery, không chỉ message count.

3. Metrics và cardinality

Per-queue metrics có cardinality lớn; giữ detailed telemetry cho critical queues và aggregate/archive hợp lý. Dashboard average có thể che một hot leader hoặc tenant; alert theo rate-of-change, age và saturation.

4. Capacity test và load shedding

Benchmark với queue type, replica count, persistence, confirms, payload distribution, routing fan-out, consumer work và storage thật. Warm cache test ngắn không mô phỏng backlog, compaction, node loss hay recovery. Chạy steady state, burst, soak và degraded-node scenarios.

Queue max length/bytes và overflow mode phải gắn business policy: reject publish, dead-letter hay drop-oldest có hậu quả khác. Admission control nên ưu tiên critical commands, giảm optional events hoặc trả overload rõ thay vì để disk đầy rồi toàn cluster dừng.

5. Security và change operations

Dùng TLS theo threat model, least-privilege vhost permissions tách configure/write/read, credential rotation, network allowlist và hạn chế management/metrics endpoints. Không đặt password trong URI/log; audit user, policy và topology changes.

Rolling upgrade, policy change, queue migration và certificate rotation cần compatibility matrix, canary node/workload, health gate và rollback. Không restart hàng loạt khi backlog cao; mỗi node loss có thể kích hoạt leader election, replication và client reconnect cùng lúc.

6. Incident runbook

Đầu tiên xác định producer spike, consumer slowdown, downstream dependency, poison message, hot leader hay resource alarm. Bảo vệ disk/quorum trước; giảm ingress hoặc shed optional work; scale consumers có giới hạn theo downstream capacity; cô lập poison; rồi drain/replay có kiểm soát.

  1. Chụp timeline và baseline rates/age/resources trước thay đổi.
  2. Stop amplification: unbounded retry, reconnect storm, requeue loop.
  3. Ổn định broker và downstream, không chỉ giảm queue depth.
  4. Drain theo rate budget; verify idempotency và business reconciliation.
  5. Sau incident, cập nhật capacity model, alerts, runbook và failure test.
Review checklist: publisher backpressure observable; client buffers bounded; age/SLA metrics có; capacity test dùng durability thật; load shedding có business policy; security endpoints bị giới hạn; incident runbook chống amplification và có reconciliation.
Tài liệu: Flow Control · Monitoring · Resource Alarms · Production Checklist · Networking and TLS