Large app backup (~70GB) fails at rotate-copy step with `NoSuchKey` — IONOS S3 read-after-write consistency, but Cloudron has no retry
-
Cloudron version: 10.0.4
Backup storage: IONOS S3-compatible Object StorageSymptom:
The backup upload itself completes successfully, but the immediately following server-side rotate-copy (UploadPartCopy) fails withNoSuchKey, causing the whole backup task to fail withOld backup not found: snapshot/app_<id>.tar.gz.enc. Reproduced on two consecutive daily runs, always on our largest app (~70GB snapshot); other apps (4–70GB) in the same runs succeed.Timeline (Task 11372, millisecond precision):
04:32:30.199 upload stats logged: 72,608,517,172 bytes transferred 04:32:30.285 backupupload: upload completed. error: null (CompleteMultipartUpload -> 200 OK) 04:32:30.466 Copying (multipart) snapshot/app_c08ad55d-...tar.gz.enc (rotate-copy starts, 181ms later) 04:32:30.534 Copying part 1 - ... bytes=0-1073741823 04:32:30.535 Copying part 2 - ... bytes=1073741824-2147483647 04:32:30.535 Copying part 3 - ... bytes=2147483648-3221225471 04:32:30.680 Aborting multipart copy (145ms after part requests) 04:32:30.748 storage/s3: copy error: NoSuchKey: UnknownError 04:32:30.748 copy to .../app_registry.20zen.de_v1.13.0.tar.gz.enc errored. error: Old backup not foundAnalysis:
CompleteMultipartUploadreturns success at .285, but the immediately followingUploadPartCopyon the exact same key returnsNoSuchKey~150ms later. This looks like a read-after-write consistency delay on the IONOS backend for very large multipart objects — AWS S3 itself has guaranteed strict read-after-write consistency (including for multipart uploads) since Dec 2020, so this is non-standard behavior on the storage backend side.That said, Cloudron's rotate-copy step (
s3.js: copyInternal) has no resilience for this: no retry, no backoff, no existence check before issuing the copy. A single transient 404 immediately fails the whole backup task, even though a short retry would very likely succeed.Request: Could a short retry/backoff (or a
HeadObjectcheck before copying) be added around the rotate-copy step for large/multipart-uploaded objects, to tolerate brief backend propagation delays on non-AWS S3-compatible providers?(I will also open a ticket at IONOS, but this little improvement would make Cloudron Backups more stable for similar situations with other providers too.)
Relevant stack:
BoxError: Old backup not found: snapshot/app_c08ad55d-….tar.gz.enc at throwError (file:///home/yellowtent/box/src/storage/s3.js:568:49) at copyInternal (file:///home/yellowtent/box/src/storage/s3.js:636:16) at process.processTicksAndRejections (node:internal/process/task_queues:104:5) at async Object.copy (file:///home/yellowtent/box/src/storage/s3.js:670:12) at async Object.copy (file:///home/yellowtent/box/src/backupformat/tgz.js:294:5) -
IONOS responded this:
"We are actively investigating the reported behavior. To support our Cloud Engineering team in conducting deeper log analysis within the storage cluster, we kindly request that you provide the following details:
- Target bucket name & IONOS S3 endpoint used during testing
- Request IDs (x-amz-request-id and x-amz-id-2 from the HTTP response headers) for both the CompleteMultipartUpload and the failed UploadPartCopy call
- Timestamps and timezone of the affected runs (or new, millisecond-precision logs if the error can be reproduced again)
While our Engineering team investigates the merging and propagation behavior for large multipart objects in the backend, we recommend testing one of the following client-side workarounds:
-
HeadObject polling with exponential backoff: Before executing UploadPartCopy on newly merged objects larger than 50 GB, insert a short polling loop using HeadObject (with retries on NoSuchKey / 404 Not Found), combined with exponential backoff and jitter. This ensures object metadata is globally visible across all gateway nodes before dependent copy operations are performed.
-
Insert a fixed delay: If implementing a polling mechanism in your Cloudron workflow is not straightforward, inserting a static delay of 1–2 seconds between CompleteMultipartUpload and subsequent copy actions can effectively prevent this race condition issue.
Please let us know whether applying one of these workarounds resolves the issue in your backup runs, and feel free to send us the requested logs once they are available."
I'm not sure if I can get the x-amz-request-id and x-amz-id-2 headers for them.
What do you think about their suggestions? -
IONOS responded this:
"We are actively investigating the reported behavior. To support our Cloud Engineering team in conducting deeper log analysis within the storage cluster, we kindly request that you provide the following details:
- Target bucket name & IONOS S3 endpoint used during testing
- Request IDs (x-amz-request-id and x-amz-id-2 from the HTTP response headers) for both the CompleteMultipartUpload and the failed UploadPartCopy call
- Timestamps and timezone of the affected runs (or new, millisecond-precision logs if the error can be reproduced again)
While our Engineering team investigates the merging and propagation behavior for large multipart objects in the backend, we recommend testing one of the following client-side workarounds:
-
HeadObject polling with exponential backoff: Before executing UploadPartCopy on newly merged objects larger than 50 GB, insert a short polling loop using HeadObject (with retries on NoSuchKey / 404 Not Found), combined with exponential backoff and jitter. This ensures object metadata is globally visible across all gateway nodes before dependent copy operations are performed.
-
Insert a fixed delay: If implementing a polling mechanism in your Cloudron workflow is not straightforward, inserting a static delay of 1–2 seconds between CompleteMultipartUpload and subsequent copy actions can effectively prevent this race condition issue.
Please let us know whether applying one of these workarounds resolves the issue in your backup runs, and feel free to send us the requested logs once they are available."
I'm not sure if I can get the x-amz-request-id and x-amz-id-2 headers for them.
What do you think about their suggestions?Hello @dsp76
HeadObject polling with exponential backoff: Before executing UploadPartCopy on newly merged objects larger than 50 GB, insert a short polling loop using HeadObject (with retries on NoSuchKey / 404 Not Found), combined with exponential backoff and jitter. This ensures object metadata is globally visible across all gateway nodes before dependent copy operations are performed.
This is an interesting suggestion.
Will need to understand it more in depth tho. -
Hi @james I think it means that - once Cloudron receives an error like NoSuchKey oder 404 not found when polling the object directly after upload - it still retries a number of times with exponentially growing delays added with some random delays ("Jitter"). So the error might only be temporal.
Hello! It looks like you're interested in this conversation, but you don't have an account yet.
Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.
With your input, this post could be even better 💗
Register Login