Part 09 · Distributed Systems

38 câu hỏi Distributed Systems

Bắt đầu bằng timeline và invariant; chỉ gọi tên pattern sau khi đã xác định failure cần giải quyết.

Failure và resilience

1. Vì sao timeout là unknown outcome?
Request có thể chưa tới, đang chạy, đã commit nhưng mất response hoặc phản hồi trễ; caller chỉ biết deadline hết.
2. Idempotency key giải quyết gì?
Cho server nhận diện cùng logical operation khi retry và trả cùng outcome/không lặp side effect; cần atomic persistence.
3. Retry storm hình thành thế nào?
Latency/failure kích retry, retry tăng load, làm latency/failure tăng; nhiều tầng retry nhân attempts.
4. Chọn timeout dựa trên gì?
Latency percentile, network/TLS/pool cost, deadline tổng và false-timeout tolerance; đo production thay vì số mặc định.
5. Backoff cần jitter vì sao?
Không jitter, clients retry cùng thời điểm tạo synchronized spikes và kéo dài overload.
6. Circuit breaker không giải quyết gì?
Không chữa dependency hoặc bảo đảm fallback đúng; chỉ fail fast và hạn chế calls khi failure vượt policy.
7. Half-open cần giới hạn probes?
Nếu mọi thread cùng probe, dependency vừa hồi phục bị stampede và circuit mở lại.
8. Bulkhead liên quan pool thế nào?
Pool/concurrency riêng giới hạn blast radius; một dependency chậm không chiếm hết threads/connections chung.
9. Vì sao queue vô hạn nguy hiểm?
Không tạo capacity; biến overload thành latency, memory và deadline-expired zombie work.
10. Load shedding khi nào đúng?
Khi capacity/budget cạn; reject work ít ưu tiên sớm để bảo vệ successful throughput của critical traffic.

Architecture và consistency

11. Modular monolith tốt hơn khi nào?
Team/domain còn nhỏ, invariants cần transaction, deployment autonomy chưa có giá trị và platform/on-call chưa đủ.
12. Shared DB phá autonomy thế nào?
Services coupling schema, lock, migration, transaction và release; ownership dữ liệu không rõ.
13. Sync vs async trade-off?
Sync immediate/simple nhưng temporal coupling; async absorb burst/decouple nhưng duplicate, order, eventual state và operations.
14. Gateway nên và không nên làm gì?
Routing, edge auth/rate limit/observability; tránh chứa domain orchestration và retry policy làm khuếch đại.
15. CAP nên giải thích thế nào?
Trong network partition, một replicated operation phải trade consistency với availability; không phải chọn hai trong ba mọi lúc.
16. Read-your-writes dưới replica lag?
Primary read, sticky/session token, version-aware wait hoặc local optimistic view với trạng thái rõ.
17. Split brain là gì?
Hai phía cùng nhận write như authority khi mất liên lạc, tạo divergent histories cần fencing/quorum/conflict resolution.
18. Eventual consistency UX thế nào?
Expose PENDING/PROCESSING/FAILED, operation ID và refresh/notification; không báo completed trước invariant hoàn tất.
19. Reconciliation có vai trò gì?
So source-of-truth với effects/views, phát hiện missing/duplicate/stuck và repair idempotently.
20. Clock timestamp có tạo global order?
Không; skew/jump và concurrent events. Dùng version/sequence/logical ordering trong scope cần thiết.

Workflow và operations

21. Saga compensation khác rollback?
Là business action mới, có thể fail/không đảo ngược hoàn toàn và cần idempotency/operator recovery.
22. Choreography vs orchestration?
Choreography ít coordinator nhưng flow implicit; orchestration rõ state/recovery nhưng thêm coordinator và central logic.
23. Outbox giải quyết và không giải quyết gì?
Atomic business update + publish intent; không bảo đảm single delivery, external consumer effect hay global ordering.
24. Polling outbox vs CDC?
Polling đơn giản/portable nhưng query/latency; CDC hiệu quả hơn nhưng thêm log connector và operational coupling.
25. CQRS có bắt buộc Event Sourcing?
Không. CQRS chỉ tách command/query models; Event Sourcing dùng events làm source of truth.
26. Event Sourcing failure khó nào?
Event schema evolution, projection rebuild, poison event, privacy deletion, tooling và storage growth.
27. Rollout database an toàn?
Expand compatible schema, deploy mixed versions, migrate/backfill, switch reads/writes, rồi contract sau khi không còn consumer cũ.
28. Trace không thay metric vì sao?
Trace sample từng request/path; metrics cho aggregate rate/error/latency/SLO và alert coverage.
29. Scale caller có thể làm DB tệ hơn?
Nhiều instances tạo thêm connections/concurrency/retries vào bottleneck đã saturated, giảm useful throughput.
30. Thiết kế Order–Payment–Inventory?
State machine durable, idempotent commands, local transactions + outbox, bounded retry, compensation/forward recovery, reconciliation và operator path.

Câu hỏi nền tảng bổ sung

31. Service communication nên sync hay async? RestTemplate khác WebClient thế nào?
Chọn synchronous request-response khi caller cần kết quả ngay và chấp nhận temporal coupling; đặt deadline, connection pool, retry có điều kiện và circuit/bulkhead. Chọn asynchronous messaging khi cần absorb burst, decouple thời điểm xử lý hoặc fan-out, đổi lại phải xử lý duplicate, ordering, eventual consistency, DLQ và observability. RestTemplate là HTTP client blocking kiểu cũ; WebClient hỗ trợ non-blocking/reactive và streaming, nhưng gọi HTTP bằng WebClient vẫn là logical request-response chứ không tự biến kiến trúc thành event-driven.
32. Microservices khác monolithic architecture ở trade-off nào?
Monolith có một deployment/process boundary, call nội bộ và transaction đơn giản hơn; modular monolith vẫn có thể giữ domain boundaries rõ. Microservices cho independent deployment/scaling/ownership theo service nhưng thêm network failure, eventual consistency, versioning, observability, security và platform/on-call cost. Bắt đầu monolith khi team/domain nhỏ và chỉ tách service khi autonomy, scale hoặc isolation benefit đã lớn hơn distributed-systems tax.

Follow-up theo tình huống

33. Gọi block() trên WebClient có vấn đề gì?
block() biến call thành chờ đồng bộ; trên WebFlux event-loop nó có thể chặn số ít worker, gây throughput collapse hoặc lỗi blocking detection. Trong ứng dụng imperative có thể dùng WebClient rồi block có kiểm soát nhưng mất phần lớn lợi ích reactive; code mới cần synchronous API nên cân nhắc RestClient, còn reactive flow phải giữ non-blocking end-to-end.
34. Deadline nên truyền qua chuỗi gọi synchronous thế nào?
Caller đặt end-to-end deadline; mỗi hop lấy remaining budget để cấu hình queue/connect/read timeout và không retry vượt thời gian còn lại. Nếu mỗi service dùng timeout độc lập dài, chain tạo zombie work và retry amplification sau khi upstream đã bỏ cuộc. Propagate cancellation khi khả thi và ghi trace deadline/attempt.
35. Chuyển từ REST sang message queue cần bổ sung guarantee nào?
Phải định nghĩa durable publish, delivery semantics, ordering scope, schema version, idempotent consumer, retry/DLQ, backpressure và reconciliation. Queue chỉ buffer burst; nếu arrival rate dài hạn vượt service rate thì backlog vẫn tăng. API/UX cần trạng thái PENDING và operation ID thay vì giả vờ kết quả tức thời.
36. Nên tách microservice theo tiêu chí nào?
Ưu tiên bounded context/data ownership, invariant, change cadence, team ownership và nhu cầu scale/failure isolation. Không tách theo technical layer như controller-service-repository. Một boundary tốt giảm cross-service chatty calls và cho service deploy độc lập mà không sửa schema/service khác.
37. Dấu hiệu của distributed monolith là gì?
Services phải deploy cùng nhau, gọi synchronous thành chuỗi dài, chia sẻ database/schema, dùng shared library lockstep và một thay đổi nhỏ cần phối hợp nhiều team. Hệ thống nhận failure/latency của distributed system nhưng không có autonomy. Có thể gộp lại modular monolith hoặc sửa ownership/contracts trước khi tách thêm.
38. Independent scaling có luôn đủ lý do tách microservice không?
Không. Có thể scale module hot bằng cache, queue, worker pool hoặc read replica mà chưa tách deployment boundary. Tách service đáng giá khi scaling profile ổn định và lợi ích isolation/autonomy bù được network, data consistency, observability và on-call cost. Cần metric bottleneck và migration path cụ thể.