Part 09 · Distributed Systems

Deadline budgets, retry và cancellation

Retry tiêu thêm capacity đúng lúc hệ thống đang yếu; phải có owner, budget và idempotency.

Deadline propagation

End-to-end deadline chia cho connection, request và downstream hops; mỗi hop dùng remaining budget, không reset full timeout. Khi caller hủy/hết deadline, propagate cancellation nếu safe để tránh zombie work.

Retry policy

Retry transient failure, bounded attempts/time, exponential backoff + jitter. Chỉ retry operation idempotent hoặc có idempotency key. Nhiều tầng mỗi tầng retry 3 lần tạo multiplicative attempts; chọn một layer ownership và retry budget.

Retry amplification: nếu năm tầng đều có tối đa ba attempts, downstream có thể nhận 3^5 = 243 calls cho một request gốc. Nếu “retry ba lần” nghĩa là ba retries ngoài initial attempt, con số là 4^5 = 1024. Luôn ghi rõ attempts hay retries, chọn một retry owner và áp end-to-end budget.

Timeout selection

Dựa latency percentile, network variance, connection/TLS cost và false-timeout tolerance. Timeout quá ngắn tạo retry load; quá dài giữ threads/connections. Warm connection hoặc prewarm có thể tránh deploy spike.

Failure handling

Phân biệt connect/read/write/pool-acquire timeout và response status. Retry-After/rate limit cần tôn trọng. Nếu outcome unknown, query operation status hoặc reconcile bằng business key thay vì gửi lại mù.

Timeouts, Retries and Backoff with Jitter · gRPC Deadlines