Bug: Some apps don't react to SIGTERM and are killed after the 10 s stop timeout
-
While auditing one of my own packages I noticed this line in
journalctl -u docker, once for every update of that app:Container failed to exit within 10s of signal 15 - using the forceMy
start.shended withexec gosu cloudron:cloudron python3 server.py, so the app itself was PID 1, and PID 1 only receives SIGTERM if it installs a handler for it. To be clear upfront: nothing broke. The update takes 10 seconds longer and requests that are in flight get cut off, and for an app with SQLite in WAL mode that is harmless. I'm only posting because it turned out not to be limited to my own packages.I then looked at how long the stop takes for every app on my three servers (Cloudron 10.0.5), using the time between
stopContaineranddeleteContainerin the app task logs since 1 September:App PID 1 Stops that took 10 s or more Surfer (4 installs) node /app/code/server.js …16 of 16 MiroTalk SFU (3 installs) npm start12 of 12 IP2Location node /app/code/index.js4 of 4 Release Bell node /app/code/index.js2 of 2 For comparison, the apps with apache2 or supervisord as PID 1 (WordPress, Nextcloud, FreeScout, EspoCRM) stop in under 4 seconds, and so does another app that runs plain
nodeas PID 1 but apparently handles the signal itself. A throwaway container from the Surfer image shows the same thing in isolation: plainnodeas PID 1 takes 10.1 s fordocker stop, with aprocess.on('SIGTERM')handler 0.1 s, and withdocker run --initalso 0.1 s.In my own two packages I changed the last line of
start.shtoexec /usr/bin/tini -- gosu cloudron:cloudron python3 server.py(tini is already incloudron/base). The stop went from 10.2 s to 0.2 s and the line is gone from the docker log.Is this something you'd rather handle per package, or is the hard kill simply fine for these apps? I can imagine it matters little for something like Surfer, I was mostly surprised to see it in the log.
-
not sure if this is worth optimizing with the extra risk that apps have less time to commit potential state to the database or disk. It is pretty standard on linux to tell the app to stop (like pressing ctrl+c in a process on the terminal). Depending on how many file/socket handles or other things are outstanding the runtimes can take their time. So instead of killing them with force earlier, the work then shifts to optimize shutdown behavior in each app, like listening to SIGTERM and trying to keep track of outstanding work, then deciding which is relevant and which isnt. All this seems more risky given the benefit of a rare action being faster, but maybe I am just not restarting apps often.
-
I think I explained it poorly, because I'm not suggesting to kill anything earlier or to shorten the timeout. The 10 s grace period is fine and should stay.
My point is that the apps in the table never get the "please stop" in the first place. With the app as PID 1 and no handler installed, the kernel doesn't deliver the SIGTERM, so the app isn't using those 10 seconds to commit state: it just keeps running as if nothing happened and is then killed with SIGKILL. That is the forced kill you'd want to avoid, and for these apps it happens on every stop. With an init as PID 1 the signal does arrive and the app gets the same 10 seconds to finish, exactly the ctrl+c behaviour you describe.
I agree the gain is small, so I'm fine leaving it at this. I mainly wanted to make sure it's a known thing.
-
Ah I see what you mean. Indeed might be worth looking into, taking surfer as an example, the nodejs process ends up as PID 1 but the default nodejs SIGTERM handler will not gracefully shutdown things like the database connection or close the expressjs socket, thus those keep hanging in there and preventing a clean shutdown. So it seems the solution for that is to keep track within the app and call
close()on the relevant things.Though this brings up the question why using
tinihere changes anything in behavior about the shutdown? -
It's not the open handles, as far as I can tell. I tried it with a node process that has nothing open at all (
node -e "setInterval(()=>{},1000)"as PID 1 in a throwaway container from the Surfer image) and it behaves the same: it survives the SIGTERM and is killed after 10 s.What happens is that node's default SIGTERM handler doesn't shut anything down itself, it resets the handler and raises the signal again with the default action. For a normal process the default action is "terminate", so node exits right away, open sockets or not. But for PID 1 the kernel ignores signals that only have the default action, so the re-raised signal is dropped and node just keeps running. In
/proc/1/statusyou can see the handler inSigCgtbefore the SIGTERM and gone after it, while the process is still there.That's also why tini makes a difference: node is then no longer PID 1, so the default action applies again and it exits in 0.1 s. Your idea works just as well and is the cleaner one: with a
process.on('SIGTERM')handler in the app the signal is delivered even as PID 1, and the app can call close() on what matters before it exits. -
N nebulon marked this topic as a regular topic
Hello! It looks like you're interested in this conversation, but you don't have an account yet.
Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.
With your input, this post could be even better 💗
Register Login