Better Cloudron release rollout
-
My Cloudron auto-updated to 10.0.05 and left me with a crashed system, which I had to manually recover.
Nginx failed with a bad write.
So I had to fix the bad write, bring up Niginx and then manually repair 80 apps (retry configure task).
This smells of something avoidable, so I asked my AI to summarise it :Yes, this failure could have been avoided. While hitting OS limits is common as a system grows, the way the system responded to those limits is where the "failure" actually occurred.
From a systems engineering perspective, there are three specific areas where Cloudron's update process failed you, and how they could approach it differently:
1. Lack of "Atomic Writes" (The Biggest Issue)
The most critical failure was that the update process wrote the Nginx configuration files directly. When the process crashed due to the
inotifylimit, it left a file that was half-written.The better approach: Atomic File Replacement.
The software should write the new configuration to a temporary file (e.g.,app.conf.tmp) and only once the write is 100% complete, use arename()system call to overwrite the original file.- Why this matters: A
renameoperation is atomic in Linux. If the system crashes during the write, you are left with the old working config and a garbage.tmpfile, rather than a broken config that prevents the entire web server from starting.
2. Absence of Pre-flight Resource Checks
The update began without verifying if the environment could support the operation for a high-density instance (80 apps).
The better approach: Resource Validation.
Before starting an update, the Cloudron Manager should check the currentsysctlandulimitsettings against the number of installed apps.- Example: If
apps > 50andmax_user_watches < 100,000, the update should either:- Automatically attempt to increase the limit.
- Pause and warn the user: "Your system limits are too low for the number of apps you have. Please increase inotify limits before proceeding."
3. "Blind" Service Restarts
The system attempted to restart Nginx without verifying that the new configuration was valid.
The better approach: Configuration Validation (Canary Testing).
The update manager should never runsystemctl restart nginxwithout first runningnginx -tinternally.- The logic: If
nginx -treturns a non-zero exit code (syntax error), the manager should stop immediately, roll back the changed files to the previous backup, and alert the user. This prevents a single corrupted file from taking down all 80 apps.
Summary: User vs. Developer Responsibility
- Your part: As a power user running 180 apps, you are operating in a "high-density" scenario. In the future, it is always a good idea to "over-provision" kernel limits (
sysctl) when you scale up. - Cloudron's part: The platform should be "crash-safe." A kernel limit should cause a slowdown or a failed update, but it should never cause permanent configuration corruption that requires manual SSH intervention to fix.
In short: You hit the limit (common), but the software "broke the furniture" on its way down (avoidable).
With Cloudron 10.1 on the way, is there any merit to these AI recommendations ?
- Why this matters: A
-
And :
location ~ '^/($|settings|app|api)(/|$)' {This is not a system limit issue; this is a data corruption/template bug caused by the update.
Somewhere in the Cloudron Manager's process, a string that should have been plain text (the Nginx regex) was HTML-encoded before being saved to the database. When the Manager writes the config file, it is literally printing HTML code into a plain-text Nginx configuration.
Nginx sees ' and has no idea what it is. Because the characters are invalid, the Nginx parser gets confused and reports that the location directive was never properly opened with a {, even though the { is right there.
-
Sorry, Cloudron team, I'm out of my depth here :
- appsmonitor installed and ran fine on 9.2.0
- it won't start on Cloudron 10.0.5
- uninstalled it and reinstalled it : same issue.
No —
applogs.txtnginx lines contain no'and no'. So apostrophes/entities are not the problem.But the logs do expose the real suspect. Line 417 (
writeAppLocationNginxConfig) shows the generated proxyAuth nginx location:"path":"!regexp:^/(public|health)(/|$)" "location":"~ ^(?!(regexp:\\^/\\(public\\|health\\)\\(/\\|\\$\\)))"Decoded, that location is
~ ^(?!(regexp:\^/\(public\|health\)\(/\|\$\)))— the literalregexp:prefix leaked into the pattern and every char is backslash-escaped, so it matches a string, not a regexp. That is malformed and is what changed between box versions; it is not a'/'issue. -
Hello @timconsidine
I just installed your community app on Cloudron 10.0.5 and also got the same issue.
The latest Cloudron box code that is not yet released does not have this problem.
It should be this commit https://git.cloudron.io/platform/box/-/commit/6705b1c1107ec43ce09eff1945cd84ef9c60ff07 that fixes it the problem which is part of Cloudron 10.1 which is not yet released.
But maybe you can hotfix your box with these changes for now. -
Sorry, Cloudron team, I'm out of my depth here :
- appsmonitor installed and ran fine on 9.2.0
- it won't start on Cloudron 10.0.5
- uninstalled it and reinstalled it : same issue.
No —
applogs.txtnginx lines contain no'and no'. So apostrophes/entities are not the problem.But the logs do expose the real suspect. Line 417 (
writeAppLocationNginxConfig) shows the generated proxyAuth nginx location:"path":"!regexp:^/(public|health)(/|$)" "location":"~ ^(?!(regexp:\\^/\\(public\\|health\\)\\(/\\|\\$\\)))"Decoded, that location is
~ ^(?!(regexp:\^/\(public\|health\)\(/\|\$\)))— the literalregexp:prefix leaked into the pattern and every char is backslash-escaped, so it matches a string, not a regexp. That is malformed and is what changed between box versions; it is not a'/'issue.But the logs do expose the real suspect. Line 417 (
writeAppLocationNginxConfig) shows the generated proxyAuth nginx location:"path":"!regexp:^/(public|health)(/|$)" "location":"~ ^(?!(regexp:\\^/\\(public\\|health\\)\\(/\\|\\$\\)))"Decoded, that location is
~ ^(?!(regexp:\^/\(public\|health\)\(/\|\$\)))— the literalregexp:prefix leaked into the pattern and every char is backslash-escaped, so it matches a string, not a regexp. That is malformed and is what changed between box versions; it is not a'/'issue.That’s literally the bug I reported here: https://forum.cloudron.io/post/129507 ?
And BTW it’s not an issue for my own app using
proxyAuth. -
The nginx validation bug/corruption bug is fixed. To fix it, you can simply delete that bad nginx config file and systemctl restart nginx. Then the dashboard will be up and you can proceed to repair the apps. Unfortunate that a bug in code brings down the whole system.
-
But the logs do expose the real suspect. Line 417 (
writeAppLocationNginxConfig) shows the generated proxyAuth nginx location:"path":"!regexp:^/(public|health)(/|$)" "location":"~ ^(?!(regexp:\\^/\\(public\\|health\\)\\(/\\|\\$\\)))"Decoded, that location is
~ ^(?!(regexp:\^/\(public\|health\)\(/\|\$\)))— the literalregexp:prefix leaked into the pattern and every char is backslash-escaped, so it matches a string, not a regexp. That is malformed and is what changed between box versions; it is not a'/'issue.That’s literally the bug I reported here: https://forum.cloudron.io/post/129507 ?
And BTW it’s not an issue for my own app using
proxyAuth.@necrevistonnezr what cloudron version are you on ? 9.2.0 ? or 10.0.5 ?
-
The nginx validation bug/corruption bug is fixed. To fix it, you can simply delete that bad nginx config file and systemctl restart nginx. Then the dashboard will be up and you can proceed to repair the apps. Unfortunate that a bug in code brings down the whole system.
The nginx validation bug/corruption bug is fixed.
Thank you @girish
In what cloudron version ?
Do I need https://git.cloudron.io/platform/box/-/commit/6705b1c1107ec43ce09eff1945cd84ef9c60ff07 ?I tried installing appsmonitor on the cloudron demo box and I get the same error.
Apologies, I am not understanding how to deploy the fix.
Seems like I have to wait for10.0.1or hack my box (which I'm not so keen to do). -
I restricted functionality (no public view facility) in my AppsMonitor app so it works with 10.0.5. (Simpler proxyAuth allowed path).
Given the additional functionality I have added to AppsMonitor, the public page (conceived as a status page) may not be a good idea anyway, so I might remove it permanently anyway. Other ways to do a status page.
Will think about it. -
@necrevistonnezr what cloudron version are you on ? 9.2.0 ? or 10.0.5 ?
@necrevistonnezr what cloudron version are you on ? 9.2.0 ? or 10.0.5 ?
10.0.5
Hello! It looks like you're interested in this conversation, but you don't have an account yet.
Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.
With your input, this post could be even better 💗
Register Login