Part 04 · PostgreSQL & Redis · 4.1.08

RDB, AOF và replication

Redis durability không phải một công tắc. Persistence policy, filesystem flush, replica lag, failover selection và thời điểm client nhận acknowledgement tạo ra những cửa sổ mất dữ liệu khác nhau.


RDB snapshots

RDB lưu point-in-time snapshot compact của dataset. Redis thường fork background child để tạo snapshot trong khi parent tiếp tục phục vụ commands. RDB thuận tiện cho backup/transfer và thường khởi động nhanh, nhưng crash có thể mất mọi write kể từ snapshot thành công gần nhất.

Fork chia sẻ pages qua copy-on-write (COW). Khi write rate cao trong lúc snapshot, changed pages được copy và peak RSS có thể tăng mạnh; fork latency cũng đáng chú ý với dataset lớn và memory pressure. Theo dõi thời gian save, last-save status, COW bytes, free memory headroom và disk throughput.

AOF và fsync policy

Append-only file ghi lại mutating commands để rebuild dataset. Durability phụ thuộc lúc bytes được ghi và lúc operating system flush xuống storage.

appendfsyncCách hoạt độngFailure window / cost
alwaysYêu cầu fsync cho mỗi batch/event-loop iteration có writes.Cửa sổ nhỏ nhất nhưng write latency và throughput phụ thuộc storage.
everysecFsync khoảng mỗi giây ở background.Thường có thể mất khoảng một giây writes khi host crash; cân bằng phổ biến.
noĐể operating system quyết định flush.Latency thấp hơn nhưng cửa sổ mất dữ liệu lớn và khó dự đoán hơn.

AOF rewrite tạo biểu diễn tối giản tương đương dataset hiện tại thay vì giữ toàn lịch sử commands. Redis hiện đại quản lý multi-part AOF bằng base file, incremental files và manifest; backup/restore phải giữ bộ artifacts nhất quán. Disk full, fsync stalls, failed rewrite và AOF growth đều cần alert và tested response.

Storage latency đi vào request latency: với policy fsync chặt, disk stall có thể trực tiếp làm write chậm. Không đánh giá durability riêng khỏi p99 latency, filesystem behavior và cloud-volume failure modes.

Kết hợp persistence và restore

Chế độƯu điểmĐiều phải chấp nhận
Không persistenceĐơn giản cho disposable cache.Restart làm mất dataset; phải rebuild được và chống stampede.
RDBArtifact compact, backup và restart thuận tiện.Mất dữ liệu từ snapshot cuối; fork/COW cost.
AOFDurability window nhỏ hơn và replay rõ.File/storage work lớn hơn; rewrite và corruption procedure.
RDB + AOFKết hợp backup convenience với AOF recovery.Thêm disk, memory và operational complexity; khi AOF enabled, Redis dùng AOF để reconstruct dataset lúc restart.

Đừng xem persistence “bật thành công” là restore đã sẵn sàng. Phải kiểm tra file ownership/permissions, encryption, version compatibility, checksum/corruption handling, available disk và thời gian load vào memory. Copy artifact phải theo quy trình nhất quán; sau restore cần validate key counts, sampled values, TTL behavior và application invariants.

Replication: full sync và partial resync

replica connects
    |
    +-- history still in replication backlog
    |      => partial resynchronization
    |
    +-- history unavailable / identity changed
           => full synchronization
           => transfer snapshot + buffered new writes

Replica theo dõi replication ID và offset. Nếu disconnect ngắn và backlog còn đủ history, partial resync chỉ gửi phần bị thiếu; nếu không, full sync phải chuyển snapshot rồi apply commands được buffer trong thời gian đó. Full sync tiêu tốn CPU, memory, bandwidth và disk tùy cấu hình, nên reconnect storm hoặc backlog quá nhỏ có thể lặp expensive resync.

Redis replication mặc định asynchronous: primary có thể acknowledge write trước khi replica nhận nó. Khi primary mất đột ngột, node được promote có thể thiếu acknowledged writes. Replication backlog hỗ trợ resync, không phải durable transaction log thay persistence.

Read replicas và consistency

Replica thường read-only nhưng trả dữ liệu theo replay progress của chính nó. Route read sang replica tạo eventual consistency: read-after-write, monotonic read và session guarantees phải do application/router giải quyết, ví dụ stick read về primary trong một window hoặc gắn version/token để kiểm tra freshness.

WAIT và acknowledgement semantics

WAIT chờ một số replicas xác nhận đã nhận/processed writes trước đó của client trong timeout. Nó cải thiện xác suất write tồn tại trên nhiều nodes nhưng không biến Redis thành strongly consistent store: replica acknowledgement không nhất thiết nghĩa dữ liệu đã fsync, failover manager vẫn có thể chọn node khác và network partition vẫn tạo các giới hạn consistency.

Client cần định nghĩa rõ acknowledgement mong muốn: accepted in primary memory, appended/fsynced locally, received by replicas hay committed vào durable source of truth bên ngoài. Không gắn nhãn chung “write thành công” cho các mức khác nhau.

Failure-window matrix

FailureRDBAOFReplicationKiểm soát chính
Redis process crashKhôi phục snapshot cuối.Replay tới flush boundary.Replica có thể tiếp tục nếu được promote.Persistence status, restart/failover runbook.
Host/power lossPhụ thuộc snapshot đã durable.Phụ thuộc fsync/storage guarantee.Async replica có thể lag.RPO, fsync policy, multi-node placement.
Logical delete/corruptionSnapshot mới có thể chứa lỗi.Command lỗi được ghi lại.Lỗi replicate sang replicas.Offline backup, retention và restore point.
FailoverKhông quyết định node mới nhất.Chỉ bảo vệ local node chứa AOF.Promoted replica có thể thiếu writes.Topology, lag policy, acknowledgement và fencing.
Quyết định source of truth: chỉ dùng Redis làm authoritative store khi business đã định lượng được data-loss window, topology và restore path đáp ứng nó. Với cache hoặc derived data, ưu tiên khả năng rebuild và stampede protection hơn durability giả tạo.

Checklist vận hành

  1. Chọn persistence và fsync policy từ RPO/latency budget, không từ default.
  2. Capacity-plan peak RSS trong fork/COW, AOF rewrite và full replication sync.
  3. Alert last persistence error, rewrite duration, disk usage/latency, replica lag, backlog và output buffers.
  4. Backup đủ bộ RDB/AOF artifacts theo quy trình nhất quán và tách blast radius.
  5. Diễn tập process crash, host loss, corrupt artifact, full resync và failover dưới write load.
  6. Đo actual data loss, restart time và application readiness sau recovery.
Nguồn tham khảo