PID 1, process và signals trong container
Container lifecycle gắn trực tiếp với process PID 1. Nếu entrypoint hoặc wrapper xử lý sai, signal dừng có thể không đến ứng dụng, child process có thể trở thành zombie, và graceful shutdown có thể bị biến thành forced kill.
1. Process & Signals
Container dùng các cơ chế isolation của Linux, trong đó PID namespace làm cho process tree bên trong container có cách nhìn riêng về PID. Process chính của container thường xuất hiện là PID 1 bên trong namespace đó. Docker gửi signal dừng đến process chính, vì vậy cách process này nhận, forward và xử lý signal quyết định chất lượng shutdown.
Signal là cơ chế kernel dùng để thông báo sự kiện cho process, ví dụ SIGTERM yêu cầu kết thúc có kiểm soát, SIGINT thường tương ứng với interrupt từ terminal, còn SIGKILL buộc process kết thúc và không thể bị catch hay ignore.
| Signal / trạng thái | Ý nghĩa vận hành | Ứng dụng nên làm gì? |
|---|---|---|
SIGTERM | Signal dừng mặc định nếu image/container không cấu hình stop signal khác. | Ngừng nhận work mới, drain work đang chạy, flush state cần thiết rồi exit trong grace period. |
SIGINT | Interrupt; thường gặp khi chạy foreground và người dùng nhấn Ctrl+C. | Nếu phù hợp, xử lý như một yêu cầu shutdown có kiểm soát. |
SIGKILL | Forced kill; không thể trap. | Không có cơ hội cleanup. Hệ thống phải chịu được crash/restart và recover từ trạng thái đã persist. |
| PID 1 exit | Container kết thúc. | Không để wrapper thoát sớm trong khi app thật còn chạy phía sau. |
2. Process vs thread
Container cô lập process tree/PID namespace, nhưng thread vẫn là execution task thuộc cùng process và chia sẻ address space cùng nhiều process resources. Một Java application thường chỉ có một JVM process nhưng có nhiều threads như request workers, GC threads, scheduler và background workers.
Cgroup accounting và limits áp vào workload của container, bao gồm CPU/memory mà process và threads tiêu thụ. Tuy nhiên lifecycle của container vẫn xoay quanh process chính: nếu PID 1 kết thúc thì container dừng, dù trước đó JVM có hàng chục hay hàng trăm thread.
3. ENTRYPOINT: exec form, shell form và wrapper
Docker hỗ trợ exec form và shell form cho ENTRYPOINT/CMD. Với application process, exec form là lựa chọn an toàn hơn vì executable được chạy trực tiếp thay vì bị bọc bởi shell.
| Kiểu | Ví dụ | Hệ quả với signal |
|---|---|---|
| Exec form | ENTRYPOINT ["java","-jar","app.jar"] | Java process trở thành process chính/PID 1 và nhận signal trực tiếp. |
| Shell form | ENTRYPOINT java -jar app.jar | Thường chạy qua /bin/sh -c; shell là PID 1 và application là child. Signal có thể không được forward như mong đợi. |
Wrapper + exec | exec java -jar app.jar | Shell được thay thế bởi application; application trở thành process chính và nhận signal. |
Wrapper script đúng mục đích
Wrapper chỉ nên tồn tại khi thật sự cần chuẩn bị runtime như render config, kiểm tra dependency, migrate nhỏ có kiểm soát hoặc drop privilege. Sau bước chuẩn bị, dùng exec để thay wrapper bằng application process.
#!/bin/sh
set -eu
# prepare runtime configuration here
exec java -jar /app/app.jar
Nếu container thật sự cần tạo nhiều child processes, hoặc application không reap child process đúng cách, một init process nhỏ có thể giúp forward signals và reap zombies. Docker hỗ trợ docker run --init; Compose cũng có init: true.
4. Stop lifecycle
Khi chạy docker stop, Docker gửi configured stop signal đến process chính. Nếu không cấu hình khác, signal mặc định là SIGTERM. Docker chờ grace period; nếu process chưa thoát khi hết timeout thì daemon gửi SIGKILL.
Trên Docker Engine, nếu container không có timeout riêng, tài liệu CLI hiện tại ghi default do daemon quyết định và mặc định là 10 giây cho Linux containers, 30 giây cho Windows containers. Production orchestrator có thể có termination grace period khác, vì vậy ứng dụng không nên hard-code giả định rằng luôn có đúng 10 giây.
Shutdown sequence nên có
- Nhận
SIGTERMhoặc stop signal đã cấu hình. - Chuyển readiness/traffic state để không nhận request hoặc job mới.
- Reject hoặc queue có kiểm soát các work mới.
- Drain in-flight requests/jobs trong thời gian hữu hạn.
- Flush/checkpoint state cần thiết; đóng connection/pool nếu framework yêu cầu.
- Exit trước khi grace period hết.
SIGKILL sẽ cắt ngang cleanup. Vì vậy shutdown hook phải có bounded work, timeout rõ ràng và dữ liệu phải chịu được crash recovery.Java shutdown hooks
Java có shutdown hooks và nhiều framework như Spring Boot tích hợp graceful shutdown, nhưng hook không phải nơi thực hiện công việc vô hạn. Nếu dependency bị treo hoặc flush chờ mãi, container vẫn có thể bị forced kill. Đặt timeout cho outbound calls, thread pools và resource cleanup để tổng shutdown budget nhỏ hơn platform grace period.
5. Diagnostics
Khi container shutdown chậm hoặc exit bất thường, đừng chỉ nhìn một exit code. Kết hợp container state, process tree, logs, OOM state, thread count, open file descriptors và signal handling.
# state / exit / OOM metadata
docker inspect <container>
# process tree from Docker view
docker top <container>
# inspect PID 1 and child processes inside the container
docker exec <container> ps -ef
# send an explicit signal for controlled testing
docker kill --signal=TERM <container>
| Triệu chứng | Khả năng | Cách phân biệt |
|---|---|---|
| Stop luôn chờ hết timeout | Shell form/wrapper không forward, app ignore signal, shutdown hook treo | Xem PID 1, process tree, log thời điểm SIGTERM và elapsed shutdown. |
| Exit code 137 | Process nhận SIGKILL (128+9); có thể do stop timeout hoặc OOM kill | Kiểm tra OOM metadata/kernel events và xem có preceding docker stop/termination event hay không. |
| Exit code 143 | Process kết thúc do SIGTERM (128+15), thường phù hợp với controlled termination | Đối chiếu deployment/stop event và application shutdown logs. |
| Zombie processes tăng | PID 1 không reap child processes | Xem process state/PPID; cân nhắc fix app hoặc dùng init process. |
| Container exit nhưng child work dở dang | Wrapper/app thoát trước background work | Kiểm tra ownership của child process và lifecycle contract. |
6. Operational signals và capacity
Graceful shutdown là một capacity problem trong thời gian ngắn: khi instance bắt đầu drain, phần traffic/work của nó phải chuyển sang các instance còn lại. Nếu rollout đồng thời quá nhiều replicas, phần còn lại có thể overload và làm shutdown chậm hơn.
- Theo dõi termination count, forced-kill count, shutdown duration và tỷ lệ shutdown vượt budget.
- Theo dõi in-flight requests/jobs, queue lag, active connections và thread pool saturation trong deployment.
- Giữ đủ headroom để service vẫn đạt SLO khi một số replicas đang terminating.
- Đặt alert cho restart loop, OOM kill, zombie growth và file-descriptor/thread growth.
7. Security và reliability boundaries
Signal handling không yêu cầu container chạy privileged. Tránh tăng Linux capabilities chỉ để “sửa” shutdown. Giữ least privilege, user không phải root khi phù hợp, và chỉ cấp quyền cần thiết cho runtime.
Đừng dựa vào graceful shutdown để bảo đảm dữ liệu. SIGKILL, host crash, kernel OOM, power loss hoặc node failure đều có thể bỏ qua cleanup. Dữ liệu quan trọng cần atomic persistence, idempotency/retry và recovery logic độc lập với shutdown hook.
8. Recovery và rollback khi rollout có lỗi shutdown
Nếu phiên bản mới làm termination chậm hoặc forced-kill tăng, rollback nên ưu tiên giảm blast radius thay vì tiếp tục rollout. Tạm dừng deployment, giữ replicas khỏe, thu thập process/thread dump nếu có thể, rồi rollback image/config. Sau rollback, xác minh exit reason, request/job loss/duplicate và backlog trước khi mở rollout lại.
ENTRYPOINT/CMDdùng exec form, hoặc wrapper kết thúc bằngexec.- PID 1 nhận được stop signal và child processes được reap.
- Readiness/traffic drain xảy ra trước khi process exit.
- Shutdown work có timeout và nhỏ hơn grace period.
- Test forced kill để chứng minh recovery không phụ thuộc cleanup.
- Phân biệt được stop-timeout kill và OOM kill khi gặp exit 137.
- Rollout có đủ spare capacity và có rollback criteria dựa trên shutdown duration/error rate.
9. Verification mini-lab
Objective: chứng minh khác biệt giữa shell form và exec form, đồng thời kiểm tra signal path đến process chính.
Prerequisites: Docker Engine và một image/application có thể log khi nhận SIGTERM.
- Build variant A bằng shell-form entrypoint và variant B bằng exec-form entrypoint.
- Run từng container, dùng
docker topđể xác định PID 1/process tree. - Gọi
docker stop -t 5 <container>và đo thời gian dừng. - Failure injection: với variant A, giữ shell wrapper không forward signal; quan sát container chỉ chết sau timeout nếu app không nhận TERM.
- Verification: variant B hoặc wrapper có
execphải log việc nhận TERM và exit trước 5 giây.
Expected evidence: process tree trước stop, application log lúc nhận TERM, elapsed stop time, final exit code và docker inspect state.
10. References
docker container stop — stop signal and timeout
docker run --init — signal forwarding and zombie reaping
Docker — Run multiple processes in a container
Linux signal(7)
Research scope: Docker Engine/Dockerfile behavior and Linux signals as documented in September 2026. Orchestrators can impose different termination grace periods, so deployment-platform settings remain the source of truth for the effective production budget.