Part 12 · Kubernetes · Execution Labs

Kubernetes: execution labs

Quan sát reconciliation bằng objects, events và endpoints; chứng minh rollout, probes, scheduling, autoscaling, disruption và storage bằng traffic cùng evidence có thể kiểm tra lại.

Nguyên tắc lab: ghi cluster version, rendered manifests, timestamps và evidence trước/sau failure injection. Không chỉ xác nhận “pod chạy”; phải nối object state với traffic, latency, events và giới hạn downstream.

Lab A · Probe và rollout

Objective

Chứng minh readiness/startup gating, rolling update và graceful termination bằng traffic liên tục, Deployment conditions, ReplicaSets, EndpointSlices và events.

Prerequisites

Baseline rollout policy

strategy:
  rollingUpdate:
    maxSurge: 1
    maxUnavailable: 0
minReadySeconds: 10
progressDeadlineSeconds: 300

Steps

  1. Deploy v1 với ba replicas và tạo continuous traffic có version/request ID.
  2. Rollout v2 với startup delay; readiness chỉ true sau initialization.
  3. Inject v2 readiness failure; quan sát Deployment condition, ReplicaSets, EndpointSlices và events.
  4. Sửa readiness rồi rollout lại; assert không có response từ unready pod.
  5. Delete một pod giữa request; kiểm tra preStop, termination grace và request drain.

Failure injection

Làm readiness của v2 thất bại đủ lâu để rollout không tiến triển, sau đó xóa một pod đang phục vụ traffic. Không dùng liveness để mô phỏng lỗi readiness vì liveness có semantics restart khác.

Verification

Expected evidence

Failure window: `maxUnavailable: 0` không đồng nghĩa zero-downtime tuyệt đối. Capacity còn phụ thuộc startup time, readiness accuracy, node capacity, external load balancer propagation và graceful shutdown của ứng dụng.

Lab B · Requests, limits và scheduling

Objective

Nối resource requests/limits với scheduler placement, QoS, CPU throttling, OOMKilled và application latency thay vì chỉ nhìn trạng thái Pod.

Prerequisites

Steps

  1. Đặt requests/limits khác nhau; xem scheduler placement và QoS class.
  2. Tạo pod unschedulable do CPU, memory hoặc affinity; chẩn đoán bằng events.
  3. Gây CPU throttling và memory OOMKilled; nối container state với app latency/log.
  4. Không sửa bằng cách tăng limit mù; ghi capacity model giải thích resource budget.

Failure injection

Tạo ít nhất một case scheduler không tìm được node phù hợp, một CPU-bound case chạm limit, và một memory-bound case bị OOMKilled.

Verification

Expected evidence

Lab C · HPA và burst

Objective

Đo toàn bộ autoscaling lag từ metric detection đến pod Ready, kiểm tra HPA trong burst ngắn và ràng buộc scale-out bằng downstream budget.

Prerequisites

Steps

  1. Cấu hình HPA theo CPU, sau đó thử custom concurrency/backlog metric.
  2. Gửi burst ngắn hơn startup time; đo detection, scheduling, image pull và readiness.
  3. Quan sát tương tác giữa HPA và rolling update, đồng thời ghi scale-in stabilization.
  4. Tính `replicas × DB pool`; cap scaling theo downstream budget.

Failure injection

Dùng burst ngắn nhưng đủ lớn để tải tăng nhanh hơn khả năng pod mới trở thành Ready. Có thể lặp lại khi một phần node capacity không còn để thấy scheduler delay.

Verification

Expected evidence

Overload caveat: HPA phản ứng sau khi tín hiệu tải xuất hiện. Nếu startup time dài hơn burst hoặc downstream đã bão hòa, scale-out có thể đến quá muộn hoặc khuếch đại sự cố. Cần queue/backpressure, admission/rate limit và capacity headroom phù hợp.

Lab D · Disruption và storage

Objective

Phân biệt voluntary disruption với node failure, xác minh phạm vi bảo vệ của PDB và chứng minh recovery của stateful workload qua PVC/attach/topology cùng backup/restore thật.

Prerequisites

Steps

  1. Tạo PDB rồi drain node; quan sát voluntary eviction.
  2. Mô phỏng node failure để thấy PDB không bảo đảm mọi availability.
  3. Với stateful workload, delete/reschedule pod; kiểm tra PVC attach và zonal constraints.
  4. Thực hiện backup/restore data thay vì gọi PVC là backup.

Failure injection

So sánh hai failure path: drain qua eviction API và mất node/involuntary disruption. Với storage, buộc pod reschedule sang nơi có thể làm lộ topology/attach constraint.

Verification

Expected evidence

Recovery limitation: PVC/PV durability không tự tạo application-consistent backup. RPO/RTO phụ thuộc storage backend, snapshot semantics, database flush/quiesce, topology và thời gian reattach/restore.

Deliverables

Operational closeout: với mỗi lab, ghi giả định, failure window, capacity signal, security boundary cần thiết và rollback/recovery action. Nếu một bước không chạy được do cluster/provider limitation, ghi evidence và limitation thay vì thay bằng kết luận suy đoán.
Tài liệu tham khảo: Kubernetes Probes · HPA Walkthrough · Disruptions & PDB · Kubernetes Storage