Part 09 · Distributed Systems

Partial failure, unknown outcome và clocks

Timeout là quan sát của caller, không phải bằng chứng server không thực hiện operation.

Remote-call timeline

Request có thể fail trước server, đang chạy, commit rồi mất response, hoặc response đến sau deadline. Caller chỉ biết không nhận kết quả đúng hạn. Retry non-idempotent có thể nhân side effect; idempotency key và reconciliation biến unknown outcome thành state có thể xử lý.

Failure detector

Timeout/heartbeat chỉ nghi ngờ node, không chứng minh node chết: network partition, GC pause và overload trông giống nhau. Chọn timeout cân bằng false suspicion với recovery speed.

Time và ordering

Wall clock có thể skew/jump; monotonic clock phù hợp đo duration. Timestamp không luôn tạo causal order. Sequence/version, logical clock hoặc broker partition order có scope rõ hơn. “Last write wins” theo clock lệch có thể mất update hợp lệ.

Failure modes

English interview answer:

A timeout is an observation from the caller, not proof that the server did nothing. The request may have failed before reaching the server, may still be running, or may have committed while the response was lost. I use idempotency and reconciliation to handle that uncertainty.

Google SRE: Handling Overload · AWS Builders' Library: Timeouts and Retries