Part 12 · Kubernetes · 12.1.09

HPA, node capacity và controlled disruption

Autoscaling là một delayed feedback loop: metric phải phản ánh đúng bottleneck, workload cần thời gian warm-up, và hạ tầng/downstream phải còn capacity. Scale nhiều Pod hơn không tự tạo database connections, quota, node capacity hay throughput phía sau.

1. Scaling

Trong Kubernetes có nhiều vòng điều khiển khác nhau: Horizontal Pod Autoscaler (HPA) thay đổi số replica; Vertical Pod Autoscaler (VPA) điều chỉnh resource requests; node autoscaler provision hoặc consolidate Nodes. Các vòng này có delay riêng và có thể tương tác xấu nếu signal, requests hoặc policy không nhất quán.

Mental model: demand tăng → metric thay đổi → workload autoscaler tạo thêm Pods → scheduler tìm chỗ chạy → nếu thiếu capacity thì node autoscaler provision Nodes → Pod start/warm-up → readiness mở traffic. Mỗi bước đều tạo một failure window.
LayerĐiều chỉnhSignal chínhDelay/risk
HPASố replicaCPU, memory, custom/external metricMetric lag, startup, stabilization, downstream saturation
VPACPU/memory requestsHistorical/observed usageUpdate có thể cần recreate/evict Pod tùy mode
Node autoscalerNode capacityPending/unschedulable Pods và scheduling constraintsCloud provisioning, quota, image pull, topology/storage

2. HPA

HPA tính desired replicas từ observed metrics so với target. Với CPU utilization, signal phụ thuộc CPU requests; requests thiếu hoặc sai làm tỷ lệ utilization và quyết định scale sai. Tùy workload, queue depth, concurrency, request rate hoặc custom SLI có thể phản ánh pressure tốt hơn CPU.

HPA cần xử lý missing metrics, Pod chưa Ready, tolerance và stabilization. autoscaling/v2 cho phép cấu hình behavior.scaleUp/scaleDown, rate policies và stabilization window. Scale-down thường cần chậm hơn scale-up để tránh churn khi metric dao động.

Không coi HPA là overload protection. Nếu database hoặc dependency đã saturated, tăng replica có thể tăng connection count, retries và pressure nhanh hơn throughput.

3. VPA and node scaling

VPA dùng usage history để recommend hoặc update resource requests. Update requests của Pod đang chạy có thể dẫn tới recreate/eviction tùy capability/mode, nên cần hiểu disruption budget và startup cost. HPA và VPA cùng điều khiển CPU/memory dễ tạo feedback conflict; nếu dùng cả hai, nên tách responsibility rõ hoặc dùng metric khác cho HPA.

Node autoscaler phản ứng khi Pods không thể schedule trên capacity hiện có và phải xét resource requests cùng affinity, topology, taints/tolerations và volume constraints. Provision Node không tức thời: cloud quota/capacity, VM boot, bootstrap, networking và image pull đều kéo dài failure window. Vì vậy sudden burst vẫn cần headroom, queueing hoặc shedding.

4. Disruption

Disruption có hai nhóm: voluntary như drain, upgrade, consolidation; và involuntary như node crash, kernel panic, network partition hoặc resource pressure. PodDisruptionBudget (PDB) giới hạn số Pod của một application có thể bị voluntary disruption cùng lúc; PDB không bảo vệ khỏi node failure và có thể làm maintenance bị block nếu application không còn đủ healthy replicas.

Eviction API tôn trọng PDB và terminationGracePeriodSeconds. Điều này khác với direct deletion theo cách có thể bypass policy của eviction flow. Khi thiết kế maintenance automation, dùng eviction-compatible flow thay vì coi mọi cách xóa Pod là tương đương.

Tình huốngPDB giúp?Điều cần thiết thêm
Planned node drain/upgradeCó, với API-initiated evictionEnough replicas/capacity, topology spread, healthy readiness
Node hardware/cloud failureKhông ngăn failureReplica distribution, recovery capacity, dependency HA
Pod OOM/resource pressureKhông phải shield chungRequests/limits, capacity, application resilience
Maintenance khi replica đã unhealthyCó thể block evictionFix health hoặc explicit incident procedure

5. Drain and shutdown

cordon ngăn Pod mới được schedule lên Node; drain dùng eviction để di chuyển workload có thể di chuyển, với các nuance cho DaemonSet/static Pod. Drain an toàn chỉ khi cluster còn capacity ở nơi khác và workload được trải theo failure domain hợp lý.

Graceful shutdown là phối hợp giữa traffic routing và process lifecycle: readiness nên chuyển false trước khi process thực sự mất khả năng phục vụ; EndpointSlice/load balancer cần thời gian ngừng gửi traffic; application xử lý SIGTERM, hoàn thành/abort request đang chạy trong deadline; preStop chỉ dùng khi thực sự cần orchestration bổ sung.

6. Overload caveat

Scale application Pods có thể tăng database connections, cache miss, fan-out RPC và retry traffic. Khi bottleneck nằm downstream, replica growth biến một incident capacity thành retry/amplification storm. Bảo vệ hệ thống bằng bounded concurrency, timeouts, circuit breaking, load shedding, queue backpressure và max replica phù hợp với downstream budget.

Scale-down cũng có cost: long-lived connections/state cần rebalance, caches lạnh lại và shutdown hàng loạt có thể tạo reconnect storm. Dùng stabilization và rate limit để tránh oscillation giữa scale-up/scale-down.

7. Failure windows và capacity signals

Failure windowSignal nên theo dõiMitigation
Metric → HPA decisionMetric age/errors, desired vs current replicasReliable metrics pipeline, causal metric, sane tolerance
Replica → schedulable PodPending Pods, scheduler reasonsRequests/topology đúng, spare node headroom
Unschedulable → new Node readyProvisioning latency/failures, cloud quotaWarm capacity, multiple node pools, quota planning
Pod start → ReadyStartup/readiness latency, image pull, init failuresSmall images, startup probes, dependency resilience
Eviction → replacement healthyPDB disruptionsAllowed, unavailable replicasEnough replicas, controlled drain concurrency

Capacity review nên theo dõi request utilization, pending Pods, node allocatable headroom, HPA saturation ở maxReplicas, dependency concurrency/connection usage, queue age và error/timeout rate. Một HPA liên tục ở max là capacity alarm chứ không chỉ là autoscaler state.

8. Security và operational safety

Autoscaling và disruption controllers có quyền thay đổi workload hoặc node lifecycle, vì vậy RBAC phải theo least privilege và cluster-autoscaler/cloud credentials cần được bảo vệ như infrastructure credentials. Admission/policy nên ngăn cấu hình nguy hiểm như requests trống ở workload production, PDB selector quá rộng hoặc max replica vượt downstream budget.

9. Rollback và recovery

Khi thay scaling policy, canary theo workload/namespace trước, lưu cấu hình HPA/VPA/PDB cũ và define rollback trigger: error rate tăng, oscillation, pending Pods kéo dài, downstream saturation hoặc cost spike. Nếu autoscaling gây amplification, ưu tiên ổn định hệ thống: cap replica/concurrency, giảm ingress, bảo vệ downstream, rồi mới tối ưu policy.

Khi drain bị block, không vội bypass PDB. Xác định vì sao healthy replicas không đủ, kiểm tra capacity thay thế và topology. Chỉ dùng forced procedure khi incident/runbook cho phép và đã chấp nhận availability risk.

Production checklist: metric phản ánh bottleneck; requests hợp lý; HPA behavior có stabilization; max replica gắn downstream budget; node provisioning latency nằm trong SLO; PDB không single-point block maintenance; topology đủ failure-domain resilience; graceful shutdown được test; dashboard có pending/HPA/PDB/downstream signals; rollback rõ.
Tài liệu chính thức: Horizontal Pod Autoscaling · Node Autoscaling · Disruptions · API-initiated Eviction · Safely Drain a Node