Part 02 · Testing & Code Quality · 2.1.10

Flaky-test operations, Git và code review

Test suite là một production system nội bộ: cần telemetry, ownership, SLO và quy trình xử lý failure thay vì chỉ retry để dashboard xanh.


1. Flaky taxonomy và evidence

NhómDấu hiệuEvidence cần thu
Time/randomPhụ thuộc timezone, clock, random seed hoặc iteration.Timezone, clock, seed, input và iteration.
Order/shared statePass riêng lẻ nhưng fail trong suite hoặc parallel.Thứ tự test, global config, data key và cleanup.
ConcurrencyRace, deadlock hoặc timing-sensitive failure.Thread dump, timing, executor state và interleaving.
InfrastructureContainer, port, CPU/memory hoặc network không ổn định.Container log, inspect output, resource và network metrics.
TimeoutDuration biến động hoặc queue bị saturation.Duration history, percentile và saturation metrics.

2. Incident flow

  1. Giữ artifact, seed, test order, commit và environment.
  2. Phân loại product defect, test defect hay infrastructure defect.
  3. Tái hiện bằng lặp có kiểm soát hoặc stress với input/seed được ghi lại.
  4. Sửa root cause và thêm regression evidence.
  5. Theo dõi recurrence sau merge, không đóng incident ngay khi một lần retry pass.

Rerun có thể thu evidence nhưng không biến lần chạy xanh thành kết luận. Quarantine phải có owner, ticket, deadline và vẫn hiển thị tín hiệu; nếu không, quarantine sẽ thành nghĩa địa test.

3. Suite observability và scale

Đo duration percentile, failure history, retry count, queue time, slowest tests và resource use. Sharding theo historical duration thường tốt hơn chia đều số file vì test cost không đồng nhất.

Chỉ bật parallel sau khi cô lập port, schema, tenant, filesystem và global runtime state. Parallel hóa một suite đang shared-state chỉ làm flaky nhanh hơn và khó tái hiện hơn.

Operational SLO: theo dõi flaky rate, time-to-triage và time-to-repair như chỉ số vận hành; test đỏ không có owner là failure của quy trình, không chỉ của code.

4. Git và review flow

Merge giữ topology của lịch sử; rebase replay commit và đổi identity, vì vậy không rewrite shared public history tùy tiện. PR nhỏ giảm review latency và test blast radius.

Reviewer nên ưu tiên correctness, security, data loss, compatibility, concurrency, failure handling, observability và test evidence trước style. Mô tả PR cần nói rõ why, trade-off, test đã chạy và cách verify failure path.

5. Operational anti-patterns

6. Checklist tự đánh giá

Nguồn tham khảo