Part 11 · Docker · 11.1.09

PID 1, process và signals trong container

Container lifecycle gắn trực tiếp với process PID 1. Nếu entrypoint hoặc wrapper xử lý sai, signal dừng có thể không đến ứng dụng, child process có thể trở thành zombie, và graceful shutdown có thể bị biến thành forced kill.

Mental model: Docker không “dừng một VM nhỏ”. Docker yêu cầu process chính trong container dừng. Khi process PID 1 kết thúc, container kết thúc.

1. Process & Signals

Container dùng các cơ chế isolation của Linux, trong đó PID namespace làm cho process tree bên trong container có cách nhìn riêng về PID. Process chính của container thường xuất hiện là PID 1 bên trong namespace đó. Docker gửi signal dừng đến process chính, vì vậy cách process này nhận, forward và xử lý signal quyết định chất lượng shutdown.

Signal là cơ chế kernel dùng để thông báo sự kiện cho process, ví dụ SIGTERM yêu cầu kết thúc có kiểm soát, SIGINT thường tương ứng với interrupt từ terminal, còn SIGKILL buộc process kết thúc và không thể bị catch hay ignore.

Signal / trạng tháiÝ nghĩa vận hànhỨng dụng nên làm gì?
SIGTERMSignal dừng mặc định nếu image/container không cấu hình stop signal khác.Ngừng nhận work mới, drain work đang chạy, flush state cần thiết rồi exit trong grace period.
SIGINTInterrupt; thường gặp khi chạy foreground và người dùng nhấn Ctrl+C.Nếu phù hợp, xử lý như một yêu cầu shutdown có kiểm soát.
SIGKILLForced kill; không thể trap.Không có cơ hội cleanup. Hệ thống phải chịu được crash/restart và recover từ trạng thái đã persist.
PID 1 exitContainer kết thúc.Không để wrapper thoát sớm trong khi app thật còn chạy phía sau.

2. Process vs thread

Container cô lập process tree/PID namespace, nhưng thread vẫn là execution task thuộc cùng process và chia sẻ address space cùng nhiều process resources. Một Java application thường chỉ có một JVM process nhưng có nhiều threads như request workers, GC threads, scheduler và background workers.

Cgroup accounting và limits áp vào workload của container, bao gồm CPU/memory mà process và threads tiêu thụ. Tuy nhiên lifecycle của container vẫn xoay quanh process chính: nếu PID 1 kết thúc thì container dừng, dù trước đó JVM có hàng chục hay hàng trăm thread.

Đừng nhầm thread count với container process count. Một JVM có thể có rất nhiều threads mà vẫn chỉ là một application process. Khi debug shutdown, cần xem cả process tree lẫn thread state.

3. ENTRYPOINT: exec form, shell form và wrapper

Docker hỗ trợ exec form và shell form cho ENTRYPOINT/CMD. Với application process, exec form là lựa chọn an toàn hơn vì executable được chạy trực tiếp thay vì bị bọc bởi shell.

KiểuVí dụHệ quả với signal
Exec formENTRYPOINT ["java","-jar","app.jar"]Java process trở thành process chính/PID 1 và nhận signal trực tiếp.
Shell formENTRYPOINT java -jar app.jarThường chạy qua /bin/sh -c; shell là PID 1 và application là child. Signal có thể không được forward như mong đợi.
Wrapper + execexec java -jar app.jarShell được thay thế bởi application; application trở thành process chính và nhận signal.

Wrapper script đúng mục đích

Wrapper chỉ nên tồn tại khi thật sự cần chuẩn bị runtime như render config, kiểm tra dependency, migrate nhỏ có kiểm soát hoặc drop privilege. Sau bước chuẩn bị, dùng exec để thay wrapper bằng application process.

#!/bin/sh
set -eu

# prepare runtime configuration here
exec java -jar /app/app.jar

Nếu container thật sự cần tạo nhiều child processes, hoặc application không reap child process đúng cách, một init process nhỏ có thể giúp forward signals và reap zombies. Docker hỗ trợ docker run --init; Compose cũng có init: true.

4. Stop lifecycle

Khi chạy docker stop, Docker gửi configured stop signal đến process chính. Nếu không cấu hình khác, signal mặc định là SIGTERM. Docker chờ grace period; nếu process chưa thoát khi hết timeout thì daemon gửi SIGKILL.

Trên Docker Engine, nếu container không có timeout riêng, tài liệu CLI hiện tại ghi default do daemon quyết định và mặc định là 10 giây cho Linux containers, 30 giây cho Windows containers. Production orchestrator có thể có termination grace period khác, vì vậy ứng dụng không nên hard-code giả định rằng luôn có đúng 10 giây.

Shutdown sequence nên có

  1. Nhận SIGTERM hoặc stop signal đã cấu hình.
  2. Chuyển readiness/traffic state để không nhận request hoặc job mới.
  3. Reject hoặc queue có kiểm soát các work mới.
  4. Drain in-flight requests/jobs trong thời gian hữu hạn.
  5. Flush/checkpoint state cần thiết; đóng connection/pool nếu framework yêu cầu.
  6. Exit trước khi grace period hết.
Failure window quan trọng: nếu drain + flush lâu hơn grace period, SIGKILL sẽ cắt ngang cleanup. Vì vậy shutdown hook phải có bounded work, timeout rõ ràng và dữ liệu phải chịu được crash recovery.

Java shutdown hooks

Java có shutdown hooks và nhiều framework như Spring Boot tích hợp graceful shutdown, nhưng hook không phải nơi thực hiện công việc vô hạn. Nếu dependency bị treo hoặc flush chờ mãi, container vẫn có thể bị forced kill. Đặt timeout cho outbound calls, thread pools và resource cleanup để tổng shutdown budget nhỏ hơn platform grace period.

5. Diagnostics

Khi container shutdown chậm hoặc exit bất thường, đừng chỉ nhìn một exit code. Kết hợp container state, process tree, logs, OOM state, thread count, open file descriptors và signal handling.

# state / exit / OOM metadata
docker inspect <container>

# process tree from Docker view
docker top <container>

# inspect PID 1 and child processes inside the container
docker exec <container> ps -ef

# send an explicit signal for controlled testing
docker kill --signal=TERM <container>
Triệu chứngKhả năngCách phân biệt
Stop luôn chờ hết timeoutShell form/wrapper không forward, app ignore signal, shutdown hook treoXem PID 1, process tree, log thời điểm SIGTERM và elapsed shutdown.
Exit code 137Process nhận SIGKILL (128+9); có thể do stop timeout hoặc OOM killKiểm tra OOM metadata/kernel events và xem có preceding docker stop/termination event hay không.
Exit code 143Process kết thúc do SIGTERM (128+15), thường phù hợp với controlled terminationĐối chiếu deployment/stop event và application shutdown logs.
Zombie processes tăngPID 1 không reap child processesXem process state/PPID; cân nhắc fix app hoặc dùng init process.
Container exit nhưng child work dở dangWrapper/app thoát trước background workKiểm tra ownership của child process và lifecycle contract.
Exit-code caution: 137 và 143 là manh mối, không phải root cause. Exit 137 không tự động đồng nghĩa OOM; cần correlation với OOMKilled/cgroup/kernel/runtime events.

6. Operational signals và capacity

Graceful shutdown là một capacity problem trong thời gian ngắn: khi instance bắt đầu drain, phần traffic/work của nó phải chuyển sang các instance còn lại. Nếu rollout đồng thời quá nhiều replicas, phần còn lại có thể overload và làm shutdown chậm hơn.

7. Security và reliability boundaries

Signal handling không yêu cầu container chạy privileged. Tránh tăng Linux capabilities chỉ để “sửa” shutdown. Giữ least privilege, user không phải root khi phù hợp, và chỉ cấp quyền cần thiết cho runtime.

Đừng dựa vào graceful shutdown để bảo đảm dữ liệu. SIGKILL, host crash, kernel OOM, power loss hoặc node failure đều có thể bỏ qua cleanup. Dữ liệu quan trọng cần atomic persistence, idempotency/retry và recovery logic độc lập với shutdown hook.

8. Recovery và rollback khi rollout có lỗi shutdown

Nếu phiên bản mới làm termination chậm hoặc forced-kill tăng, rollback nên ưu tiên giảm blast radius thay vì tiếp tục rollout. Tạm dừng deployment, giữ replicas khỏe, thu thập process/thread dump nếu có thể, rồi rollback image/config. Sau rollback, xác minh exit reason, request/job loss/duplicate và backlog trước khi mở rollout lại.

Production checklist:

9. Verification mini-lab

Objective: chứng minh khác biệt giữa shell form và exec form, đồng thời kiểm tra signal path đến process chính.

Prerequisites: Docker Engine và một image/application có thể log khi nhận SIGTERM.

  1. Build variant A bằng shell-form entrypoint và variant B bằng exec-form entrypoint.
  2. Run từng container, dùng docker top để xác định PID 1/process tree.
  3. Gọi docker stop -t 5 <container> và đo thời gian dừng.
  4. Failure injection: với variant A, giữ shell wrapper không forward signal; quan sát container chỉ chết sau timeout nếu app không nhận TERM.
  5. Verification: variant B hoặc wrapper có exec phải log việc nhận TERM và exit trước 5 giây.

Expected evidence: process tree trước stop, application log lúc nhận TERM, elapsed stop time, final exit code và docker inspect state.

10. References

Dockerfile reference — ENTRYPOINT / exec vs shell form
docker container stop — stop signal and timeout
docker run --init — signal forwarding and zombie reaping
Docker — Run multiple processes in a container
Linux signal(7)

Research scope: Docker Engine/Dockerfile behavior and Linux signals as documented in September 2026. Orchestrators can impose different termination grace periods, so deployment-platform settings remain the source of truth for the effective production budget.