HPA, node capacity và controlled disruption
Autoscaling là một delayed feedback loop: metric phải phản ánh đúng bottleneck, workload cần thời gian warm-up, và hạ tầng/downstream phải còn capacity. Scale nhiều Pod hơn không tự tạo database connections, quota, node capacity hay throughput phía sau.
1. Scaling
Trong Kubernetes có nhiều vòng điều khiển khác nhau: Horizontal Pod Autoscaler (HPA) thay đổi số replica; Vertical Pod Autoscaler (VPA) điều chỉnh resource requests; node autoscaler provision hoặc consolidate Nodes. Các vòng này có delay riêng và có thể tương tác xấu nếu signal, requests hoặc policy không nhất quán.
| Layer | Điều chỉnh | Signal chính | Delay/risk |
|---|---|---|---|
| HPA | Số replica | CPU, memory, custom/external metric | Metric lag, startup, stabilization, downstream saturation |
| VPA | CPU/memory requests | Historical/observed usage | Update có thể cần recreate/evict Pod tùy mode |
| Node autoscaler | Node capacity | Pending/unschedulable Pods và scheduling constraints | Cloud provisioning, quota, image pull, topology/storage |
2. HPA
HPA tính desired replicas từ observed metrics so với target. Với CPU utilization, signal phụ thuộc CPU requests; requests thiếu hoặc sai làm tỷ lệ utilization và quyết định scale sai. Tùy workload, queue depth, concurrency, request rate hoặc custom SLI có thể phản ánh pressure tốt hơn CPU.
HPA cần xử lý missing metrics, Pod chưa Ready, tolerance và stabilization. autoscaling/v2 cho phép cấu hình behavior.scaleUp/scaleDown, rate policies và stabilization window. Scale-down thường cần chậm hơn scale-up để tránh churn khi metric dao động.
- Đừng scale theo metric đẹp mà không causal: CPU thấp không có nghĩa queue không backlog; CPU cao do GC/warm-up không nhất thiết cần thêm replica.
- Readiness phải đúng: Pod chưa thực sự phục vụ mà Ready quá sớm sẽ nhận traffic và có thể làm HPA hiểu sai.
- Min/max replicas là capacity policy: min bảo vệ cold-start/headroom; max bảo vệ cost và downstream.
- Observe lag: metric collection + controller loop + scheduling + startup tạo tổng reaction time; burst ngắn hơn reaction time phải được hấp thụ bằng headroom, queue hoặc load shedding.
3. VPA and node scaling
VPA dùng usage history để recommend hoặc update resource requests. Update requests của Pod đang chạy có thể dẫn tới recreate/eviction tùy capability/mode, nên cần hiểu disruption budget và startup cost. HPA và VPA cùng điều khiển CPU/memory dễ tạo feedback conflict; nếu dùng cả hai, nên tách responsibility rõ hoặc dùng metric khác cho HPA.
Node autoscaler phản ứng khi Pods không thể schedule trên capacity hiện có và phải xét resource requests cùng affinity, topology, taints/tolerations và volume constraints. Provision Node không tức thời: cloud quota/capacity, VM boot, bootstrap, networking và image pull đều kéo dài failure window. Vì vậy sudden burst vẫn cần headroom, queueing hoặc shedding.
- Requests quá thấp: scheduler pack quá chặt, node provisioning không phản ánh thực tế và runtime có thể throttle/OOM.
- Requests quá cao: Pod khó schedule, node cost tăng và consolidation bị hạn chế.
- Kiểm tra pending Pods theo reason, autoscaler provisioning failures, node launch latency và available headroom.
- Giữ quota/capacity dự phòng ở cloud provider nếu SLO không chấp nhận chờ provisioning.
4. Disruption
Disruption có hai nhóm: voluntary như drain, upgrade, consolidation; và involuntary như node crash, kernel panic, network partition hoặc resource pressure. PodDisruptionBudget (PDB) giới hạn số Pod của một application có thể bị voluntary disruption cùng lúc; PDB không bảo vệ khỏi node failure và có thể làm maintenance bị block nếu application không còn đủ healthy replicas.
Eviction API tôn trọng PDB và terminationGracePeriodSeconds. Điều này khác với direct deletion theo cách có thể bypass policy của eviction flow. Khi thiết kế maintenance automation, dùng eviction-compatible flow thay vì coi mọi cách xóa Pod là tương đương.
| Tình huống | PDB giúp? | Điều cần thiết thêm |
|---|---|---|
| Planned node drain/upgrade | Có, với API-initiated eviction | Enough replicas/capacity, topology spread, healthy readiness |
| Node hardware/cloud failure | Không ngăn failure | Replica distribution, recovery capacity, dependency HA |
| Pod OOM/resource pressure | Không phải shield chung | Requests/limits, capacity, application resilience |
| Maintenance khi replica đã unhealthy | Có thể block eviction | Fix health hoặc explicit incident procedure |
5. Drain and shutdown
cordon ngăn Pod mới được schedule lên Node; drain dùng eviction để di chuyển workload có thể di chuyển, với các nuance cho DaemonSet/static Pod. Drain an toàn chỉ khi cluster còn capacity ở nơi khác và workload được trải theo failure domain hợp lý.
Graceful shutdown là phối hợp giữa traffic routing và process lifecycle: readiness nên chuyển false trước khi process thực sự mất khả năng phục vụ; EndpointSlice/load balancer cần thời gian ngừng gửi traffic; application xử lý SIGTERM, hoàn thành/abort request đang chạy trong deadline; preStop chỉ dùng khi thực sự cần orchestration bổ sung.
- Đo termination latency p50/p95/p99 và đặt
terminationGracePeriodSecondsdựa trên workload thực tế. - Kiểm tra long-lived connections, WebSocket, consumer partition ownership và state rebalance.
- Đừng drain nhiều failure domains cùng lúc nếu replica/topology không chịu được.
- Trước maintenance lớn, xác nhận PDB status, current healthy replicas và spare schedulable capacity.
6. Overload caveat
Scale application Pods có thể tăng database connections, cache miss, fan-out RPC và retry traffic. Khi bottleneck nằm downstream, replica growth biến một incident capacity thành retry/amplification storm. Bảo vệ hệ thống bằng bounded concurrency, timeouts, circuit breaking, load shedding, queue backpressure và max replica phù hợp với downstream budget.
Scale-down cũng có cost: long-lived connections/state cần rebalance, caches lạnh lại và shutdown hàng loạt có thể tạo reconnect storm. Dùng stabilization và rate limit để tránh oscillation giữa scale-up/scale-down.
7. Failure windows và capacity signals
| Failure window | Signal nên theo dõi | Mitigation |
|---|---|---|
| Metric → HPA decision | Metric age/errors, desired vs current replicas | Reliable metrics pipeline, causal metric, sane tolerance |
| Replica → schedulable Pod | Pending Pods, scheduler reasons | Requests/topology đúng, spare node headroom |
| Unschedulable → new Node ready | Provisioning latency/failures, cloud quota | Warm capacity, multiple node pools, quota planning |
| Pod start → Ready | Startup/readiness latency, image pull, init failures | Small images, startup probes, dependency resilience |
| Eviction → replacement healthy | PDB disruptionsAllowed, unavailable replicas | Enough replicas, controlled drain concurrency |
Capacity review nên theo dõi request utilization, pending Pods, node allocatable headroom, HPA saturation ở maxReplicas, dependency concurrency/connection usage, queue age và error/timeout rate. Một HPA liên tục ở max là capacity alarm chứ không chỉ là autoscaler state.
8. Security và operational safety
Autoscaling và disruption controllers có quyền thay đổi workload hoặc node lifecycle, vì vậy RBAC phải theo least privilege và cluster-autoscaler/cloud credentials cần được bảo vệ như infrastructure credentials. Admission/policy nên ngăn cấu hình nguy hiểm như requests trống ở workload production, PDB selector quá rộng hoặc max replica vượt downstream budget.
- Audit ai thay HPA/VPA/PDB, deployment replicas và node pool limits.
- Rate-limit hoặc change-review cho policy làm tăng cost rất lớn hoặc xóa capacity quá nhanh.
- Không đặt secret trong autoscaler arguments/logs; bảo vệ metrics APIs vì metric sai có thể lái scaling sai.
- Maintenance automation phải có dry-run/preview, stop condition và visibility vào blocked evictions.
9. Rollback và recovery
Khi thay scaling policy, canary theo workload/namespace trước, lưu cấu hình HPA/VPA/PDB cũ và define rollback trigger: error rate tăng, oscillation, pending Pods kéo dài, downstream saturation hoặc cost spike. Nếu autoscaling gây amplification, ưu tiên ổn định hệ thống: cap replica/concurrency, giảm ingress, bảo vệ downstream, rồi mới tối ưu policy.
Khi drain bị block, không vội bypass PDB. Xác định vì sao healthy replicas không đủ, kiểm tra capacity thay thế và topology. Chỉ dùng forced procedure khi incident/runbook cho phép và đã chấp nhận availability risk.