Cloudron makes it easy to run web apps like WordPress, Nextcloud, GitLab on your server. Find out more or install now.


Skip to content
  • Categories
  • Recent
  • Tags
  • Popular
  • Bookmarks
  • Search
Skins
  • Light
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • Dark
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • Default (No Skin)
  • No Skin
Collapse
Brand Logo

Cloudron Forum

Offical apps | Community apps | Demo | Docs | Install
  1. Cloudron Forum
  2. Feature Requests
  3. Restoring an app from an rsync backup downloads everything again, even when almost all files are already there

Restoring an app from an rsync backup downloads everything again, even when almost all files are already there

Scheduled Pinned Locked Moved Feature Requests
backuprestorersync
3 Posts 2 Posters 57 Views 2 Watching
  • Oldest to Newest
  • Newest to Oldest
  • Most Votes
Reply
  • Reply as topic
Log in to reply
This topic has been deleted. Only users with topic management privileges can see it.
  • imc67I
    imc67I
    imc67
    translator
    wrote last edited by imc67
    #1

    This morning the Immich update to 1.102.3 broke machine learning (topic 16034), so I restored the backup from just before the update. That brought back something I've raised before: backups are fast, restores are very slow.

    Immich, sshfs + rsync to a Hetzner Storage Box Files Data Time
    Backup this morning (incremental) 151,807 1.1 GB transferred 55 s
    Restore of the backup before the update 151,807 291 GB ~1.5 h

    Almost all of those 291 GB were already on the server and unchanged. The restore first empties the data directory and then fetches every file again, one by one. During that time the app is down, and my phone can't upload photos.

    I understand the earlier answers: Cloudron's rsync format isn't real rsync, and most setups use object storage, where comparing with local files isn't straightforward (13362, 13499). But in 13499 girish mentioned that backup integrity with checksums could make it possible to check whether local and remote files are "the same". That is now in place: every backup has a .backupinfo with a sha256 per file.

    What I have in mind is simple: don't empty the data directory, but sync the backup into it. Files whose size and sha256 match the .backupinfo stay where they are, only missing or changed files are downloaded, and files that aren't in the backup are removed. In this case that would have been the photos uploaded after the backup, a few changed thumbnails and the database dump, roughly the size of the incremental backup instead of 291 GB. For sshfs targets that support real rsync (a Hetzner Storage Box does), even plain rsync -a --delete from the snapshot would work. The obvious downside is that the restore then has to read and hash the local files first, but on a local disk that is far faster than fetching 150k files one by one.

    It matters even more for a full server restore. At this speed my largest server would take well over a day, while the incremental backups take minutes. I've also followed the sshfs read-speed discussion (13852).

    Is something like this possible now that the integrity data is there, at least for the rsync format?

    1 Reply Last reply
    4
    • girishG
      girishG
      girish
      Staff
      wrote last edited by
      #2

      Thanks for the reminder! Forgot about this. Yeah, so an idea would be to hash the file locally and only download if something changed. There is a CPU vs network optimization. So, it might make sense for large files and avoid re-downloading then again.

      1 Reply Last reply
      1
      • imc67I
        imc67I
        imc67
        translator
        wrote last edited by
        #3

        Thanks! One data point from today's restore: the time was driven by the number of files, not their size. The ~100k small thumbnails took most of it, since every file costs a round trip over sshfs, so skipping unchanged small files would help just as much. A quick size + mtime check before hashing would probably keep the CPU side cheap too.

        1 Reply Last reply
        1

        Hello! It looks like you're interested in this conversation, but you don't have an account yet.

        Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

        With your input, this post could be even better 💗

        Register Login
        Reply
        • Reply as topic
        Log in to reply
        • Oldest to Newest
        • Newest to Oldest
        • Most Votes


        • Login

        • Don't have an account? Register

        • Login or register to search.
        • First post
          Last post
        0
        • Categories
        • Recent
        • Tags
        • Popular
        • Bookmarks
        • Search