RDB, AOF và replication
Redis durability không phải một công tắc. Persistence policy, filesystem flush, replica lag, failover selection và thời điểm client nhận acknowledgement tạo ra những cửa sổ mất dữ liệu khác nhau.
RDB snapshots
RDB lưu point-in-time snapshot compact của dataset. Redis thường fork background child để tạo snapshot trong khi parent tiếp tục phục vụ commands. RDB thuận tiện cho backup/transfer và thường khởi động nhanh, nhưng crash có thể mất mọi write kể từ snapshot thành công gần nhất.
Fork chia sẻ pages qua copy-on-write (COW). Khi write rate cao trong lúc snapshot, changed pages được copy và peak RSS có thể tăng mạnh; fork latency cũng đáng chú ý với dataset lớn và memory pressure. Theo dõi thời gian save, last-save status, COW bytes, free memory headroom và disk throughput.
AOF và fsync policy
Append-only file ghi lại mutating commands để rebuild dataset. Durability phụ thuộc lúc bytes được ghi và lúc operating system flush xuống storage.
appendfsync | Cách hoạt động | Failure window / cost |
|---|---|---|
always | Yêu cầu fsync cho mỗi batch/event-loop iteration có writes. | Cửa sổ nhỏ nhất nhưng write latency và throughput phụ thuộc storage. |
everysec | Fsync khoảng mỗi giây ở background. | Thường có thể mất khoảng một giây writes khi host crash; cân bằng phổ biến. |
no | Để operating system quyết định flush. | Latency thấp hơn nhưng cửa sổ mất dữ liệu lớn và khó dự đoán hơn. |
AOF rewrite tạo biểu diễn tối giản tương đương dataset hiện tại thay vì giữ toàn lịch sử commands. Redis hiện đại quản lý multi-part AOF bằng base file, incremental files và manifest; backup/restore phải giữ bộ artifacts nhất quán. Disk full, fsync stalls, failed rewrite và AOF growth đều cần alert và tested response.
Kết hợp persistence và restore
| Chế độ | Ưu điểm | Điều phải chấp nhận |
|---|---|---|
| Không persistence | Đơn giản cho disposable cache. | Restart làm mất dataset; phải rebuild được và chống stampede. |
| RDB | Artifact compact, backup và restart thuận tiện. | Mất dữ liệu từ snapshot cuối; fork/COW cost. |
| AOF | Durability window nhỏ hơn và replay rõ. | File/storage work lớn hơn; rewrite và corruption procedure. |
| RDB + AOF | Kết hợp backup convenience với AOF recovery. | Thêm disk, memory và operational complexity; khi AOF enabled, Redis dùng AOF để reconstruct dataset lúc restart. |
Đừng xem persistence “bật thành công” là restore đã sẵn sàng. Phải kiểm tra file ownership/permissions, encryption, version compatibility, checksum/corruption handling, available disk và thời gian load vào memory. Copy artifact phải theo quy trình nhất quán; sau restore cần validate key counts, sampled values, TTL behavior và application invariants.
Replication: full sync và partial resync
replica connects
|
+-- history still in replication backlog
| => partial resynchronization
|
+-- history unavailable / identity changed
=> full synchronization
=> transfer snapshot + buffered new writes
Replica theo dõi replication ID và offset. Nếu disconnect ngắn và backlog còn đủ history, partial resync chỉ gửi phần bị thiếu; nếu không, full sync phải chuyển snapshot rồi apply commands được buffer trong thời gian đó. Full sync tiêu tốn CPU, memory, bandwidth và disk tùy cấu hình, nên reconnect storm hoặc backlog quá nhỏ có thể lặp expensive resync.
Redis replication mặc định asynchronous: primary có thể acknowledge write trước khi replica nhận nó. Khi primary mất đột ngột, node được promote có thể thiếu acknowledged writes. Replication backlog hỗ trợ resync, không phải durable transaction log thay persistence.
Read replicas và consistency
Replica thường read-only nhưng trả dữ liệu theo replay progress của chính nó. Route read sang replica tạo eventual consistency: read-after-write, monotonic read và session guarantees phải do application/router giải quyết, ví dụ stick read về primary trong một window hoặc gắn version/token để kiểm tra freshness.
- Giám sát byte/second lag và time lag; time lag một mình dễ gây hiểu nhầm khi traffic thấp.
- Slow hoặc disconnected replica có thể làm output buffer tăng; đặt limits và cảnh báo trước khi memory cạn.
- Replica restart/full sync có thể tạo burst load lên primary và network.
- Read scaling không miễn phí nếu stale data làm business decision sai.
WAIT và acknowledgement semantics
WAIT chờ một số replicas xác nhận đã nhận/processed writes trước đó của client trong timeout. Nó cải thiện xác suất write tồn tại trên nhiều nodes nhưng không biến Redis thành strongly consistent store: replica acknowledgement không nhất thiết nghĩa dữ liệu đã fsync, failover manager vẫn có thể chọn node khác và network partition vẫn tạo các giới hạn consistency.
Client cần định nghĩa rõ acknowledgement mong muốn: accepted in primary memory, appended/fsynced locally, received by replicas hay committed vào durable source of truth bên ngoài. Không gắn nhãn chung “write thành công” cho các mức khác nhau.
Failure-window matrix
| Failure | RDB | AOF | Replication | Kiểm soát chính |
|---|---|---|---|---|
| Redis process crash | Khôi phục snapshot cuối. | Replay tới flush boundary. | Replica có thể tiếp tục nếu được promote. | Persistence status, restart/failover runbook. |
| Host/power loss | Phụ thuộc snapshot đã durable. | Phụ thuộc fsync/storage guarantee. | Async replica có thể lag. | RPO, fsync policy, multi-node placement. |
| Logical delete/corruption | Snapshot mới có thể chứa lỗi. | Command lỗi được ghi lại. | Lỗi replicate sang replicas. | Offline backup, retention và restore point. |
| Failover | Không quyết định node mới nhất. | Chỉ bảo vệ local node chứa AOF. | Promoted replica có thể thiếu writes. | Topology, lag policy, acknowledgement và fencing. |
Checklist vận hành
- Chọn persistence và fsync policy từ RPO/latency budget, không từ default.
- Capacity-plan peak RSS trong fork/COW, AOF rewrite và full replication sync.
- Alert last persistence error, rewrite duration, disk usage/latency, replica lag, backlog và output buffers.
- Backup đủ bộ RDB/AOF artifacts theo quy trình nhất quán và tách blast radius.
- Diễn tập process crash, host loss, corrupt artifact, full resync và failover dưới write load.
- Đo actual data loss, restart time và application readiness sau recovery.