Part 12 · Kubernetes · 12.2

30 câu hỏi Kubernetes

Ôn theo API/reconciliation call flow, observed evidence và user-visible outcome. Khi trả lời phỏng vấn hoặc debug production, ưu tiên nói rõ desired state → controller/runtime action → evidence quan sát được → failure mode.

Cách dùng: tự trả lời trước khi mở từng mục. Với các câu về rollout, autoscaling, storage hay disruption, đừng dừng ở định nghĩa; hãy nói thêm dấu hiệu quan sát, failure window và cách recovery.

Control plane và workloads

1. Control plane components?
API server là front door của cluster; etcd lưu cluster state; scheduler chọn node cho Pod chưa được bind; controller manager chạy reconciliation loops; kubelet trên node quản lý lifecycle Pod qua container runtime. Một object được API server chấp nhận chưa có nghĩa workload đã ready.
2. Reconciliation nghĩa gì?
Là vòng lặp level-triggered và idempotent: controller đọc desired state từ API, so với observed state, rồi tạo/cập nhật/xóa tài nguyên để tiến tới convergence. Vì vậy hệ thống là eventually convergent chứ không phải một lệnh đồng bộ “apply xong là chạy xong”.
3. spec khác status?
spec mô tả desired intent; statusconditions mô tả observed state do controller cập nhật. Khi debug cần để ý observedGeneration để biết status đang phản ánh generation mới nhất hay dữ liệu cũ.
4. Finalizer làm object Terminating?
Khi object có finalizer, delete sẽ đặt deletion timestamp nhưng giữ object tồn tại cho tới khi controller hoàn tất cleanup rồi remove finalizer. Nếu bị kẹt Terminating, hãy tìm controller owner và external dependency; không nên xóa finalizer mù vì có thể bỏ sót cleanup và rò tài nguyên.
5. Deployment→Pod flow?
Deployment controller tạo/điều chỉnh ReplicaSet; ReplicaSet controller tạo Pod; scheduler chọn node và bind Pod; kubelet pull image/start containers; readiness quyết định Pod có được publish vào EndpointSlice để nhận traffic hay không.
6. Pod khác container?
Pod là scheduling unit của Kubernetes, có thể chứa một hoặc nhiều containers chia sẻ network namespace, volumes và lifecycle coupling. Container là process/runtime unit bên trong Pod; Kubernetes schedule Pod chứ không schedule từng container độc lập.
7. Deployment vs StatefulSet?
Deployment phù hợp replicas có thể thay thế lẫn nhau, thường cho stateless workloads. StatefulSet cung cấp identity ổn định như ordinal, DNS và PVC association. Nó không tự biến ứng dụng thành distributed database hay tự cung cấp replication/consensus cho dữ liệu.
8. Job exactly-once?
Không. Job/Pod có thể retry và một task có thể chạy lặp trong một số failure windows. Side effect bên ngoài như payment, email hay ghi DB cần idempotency key, checkpoint/deduplication hoặc reconciliation riêng.
9. Rollout complete đủ chưa?
Chưa. Deployment báo rollout complete chủ yếu nói controller đã đạt desired replicas/availability theo Kubernetes. Vẫn phải kiểm business SLI, error rate, latency, canary result, schema/API compatibility và dependency health trước khi coi release thành công.
10. Zero-downtime rollout cần gì?
Cần đủ surge/headroom, readiness phản ánh khả năng serve thật, graceful drain/termination, contract/schema tương thích giữa old-new versions, và đủ replicas/topology để chịu mất instance trong rollout. Nếu downstream hoặc database migration không backward-compatible thì Kubernetes rollout một mình không đảm bảo zero downtime.

Resources, network và storage

11. Requests khác limits?
requests là input quan trọng cho scheduling và một số autoscaling decisions; limits là runtime cap/enforcement. CPU vượt limit thường bị throttling; memory vượt limit có thể dẫn tới OOM kill. Requests quá thấp cũng làm bin-packing và HPA utilization misleading.
12. CPU limit gây tail latency?
Có thể. Cgroup CPU quota có thể throttle container theo quota period khi workload burst vượt limit, dù node vẫn còn CPU khả dụng. Request/limit quá chặt dễ làm p95/p99 tăng trong các burst ngắn.
13. OOMKilled debug?
Xem container last state/exit reason, configured memory limit, working set/trend, traffic pattern và node pressure. Với runtime như JVM cần phân biệt heap, native memory, metaspace, thread stacks, direct buffers/page cache; đồng thời kiểm leak hoặc concurrency tăng bất thường.
14. Pod Pending?
Bắt đầu từ Events và scheduler messages: thiếu CPU/memory, taint không được tolerate, affinity/topology constraints, PVC chưa bind, image pull secret/config issue hoặc các scheduling constraints khác. Pending không đồng nghĩa chỉ thiếu resource.
15. Taint/toleration?
Taint làm node repel Pod; toleration chỉ cho phép Pod có thể được schedule lên node đó, không bắt buộc Pod phải lên đó. Nếu cần placement bắt buộc, kết hợp node selector/affinity hoặc topology constraints phù hợp.
16. Running nhưng no traffic?
Kiểm readiness trước; sau đó Service selector, EndpointSlice, port/targetPort, process listener, DNS, NetworkPolicy và Gateway/Ingress/controller path. Pod ở phase Running chỉ nói container process đang chạy, không nói traffic path end-to-end đang tốt.
17. Service load balancing ở đâu?
Service tạo virtual service abstraction; dataplane implementation như kube-proxy hoặc eBPF-based CNI chuyển traffic tới các ready endpoints trong EndpointSlice. Exact packet path phụ thuộc implementation, mode và network plugin.
18. Ingress object có đủ?
Không. Ingress resource cần một Ingress controller thực thi nó. TLS termination, timeout, retry, rewrite và annotation semantics phụ thuộc controller implementation; ở hệ thống mới cũng cần cân nhắc Gateway API khi phù hợp.
19. NetworkPolicy không tác dụng?
Các nguyên nhân thường gặp: CNI không hỗ trợ/enforce NetworkPolicy, selector sai, hiểu nhầm policy additive semantics, hoặc traffic path/node component nằm ngoài expectation. Debug bằng cách xác nhận policy-selected Pods, namespace labels, direction ingress/egress và CNI capability.
20. PVC Pending?
Kiểm StorageClass, CSI provisioner, capacity/quota, requested access mode, topology/zone constraints, volume binding mode và Events của PVC/Pod. Với WaitForFirstConsumer, PVC có thể chờ scheduling context trước khi volume được provision/bind.

Probes, scaling, security và ops

21. Readiness/liveness/startup?
Readiness quyết định Pod có nên nhận traffic; liveness quyết định khi nào kubelet nên restart container; startup probe bảo vệ ứng dụng khởi động chậm bằng cách trì hoãn liveness/readiness cho tới khi startup thành công.

English interview answer: Readiness controls traffic eligibility, liveness decides whether Kubernetes should restart the container, and startup protects slow initialization. I keep liveness focused on unrecoverable local failure rather than transient downstream health.

22. Liveness gọi DB?
Thường không nên nếu DB là shared transient dependency. Khi DB chập chờn, liveness fail trên cả fleet có thể gây restart storm và làm incident nặng hơn. Readiness có thể phản ánh dependency theo contract phục vụ traffic, còn liveness nên tập trung vào process thực sự bị wedged.
23. preStop có cộng thêm grace?
Không nên coi preStop là thời gian cộng thêm độc lập. Hook chạy trong termination flow và tiêu vào terminationGracePeriodSeconds cùng thời gian app xử lý TERM/drain, nên phải budget tổng thời gian shutdown cho phù hợp.
24. HPA và requests?
Với CPU utilization target, HPA tính utilization tương quan với CPU request. Request sai hoặc thiếu có thể khiến metric/desired replicas không phản ánh capacity thật. Ngoài CPU, HPA có thể dùng memory, custom hoặc external metrics tùy setup.
25. Autoscaling không cứu overload?
Autoscaling có detection delay, provisioning delay và warm-up delay; downstream còn có hard capacity riêng. Production cần headroom, queue/backpressure, concurrency limit, rate limit/load shedding và capacity planning chứ không chỉ HPA.
26. PDB bảo vệ gì?
PodDisruptionBudget giới hạn mức disruption được phép trong các voluntary disruptions đi qua eviction flow, ví dụ drain. Nó không ngăn node crash và không trực tiếp điều khiển Deployment rolling update. PDB quá chặt cũng có thể làm drain/maintenance bị block.
27. Secret có encrypted mặc định?
Secret values trong manifest/API representation thường chỉ base64-encoded, không phải encryption. Cần encryption at rest cho Kubernetes API data, RBAC least privilege, audit, rotation và có thể external secret manager theo threat model.
28. Namespace có hard isolation?
Không. Namespace là scope/organization boundary chứ không tự tạo hard multitenant isolation. Cần kết hợp RBAC, NetworkPolicy, ResourceQuota/LimitRange, admission/policy, Pod security controls và trong threat model mạnh hơn có thể cần separate nodes hoặc clusters.
29. CrashLoopBackOff flow?
Xem kubectl describe/Events, current và previous logs, last state/exit code, command/args, config/env, mounts/permissions, probes và resources. CrashLoopBackOff là backoff behavior sau restart lặp lại, không phải root cause.
30. Node drain safe?
Cordon node rồi evict workloads qua drain; kiểm PDB, capacity/topology ở nơi khác, termination grace và stateful behavior. Cần hiểu ngoại lệ/nuance với DaemonSet, static Pods, local data và workloads không đủ replicas; drain an toàn phải có verification và rollback/abort criteria.
Production reminder: khi gặp incident, đừng sửa theo tên resource trước. Hãy xác định failure window, control-plane evidence, node/runtime evidence, traffic/storage path và user-visible impact; sau đó chọn hành động ít phá hủy nhất và giữ đường rollback.