From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from gate001.proxmox.com (gate001.proxmox.com [45.144.208.40]) by lore.proxmox.com (Postfix) with ESMTPS id F2E691FF0E1 for ; Mon, 10 Aug 2026 07:02:40 +0200 (CEST) Received: from gate001.proxmox.com (localhost.localdomain [127.0.0.1]) by gate001.proxmox.com (Proxmox) with ESMTP id 213DC215CA; Mon, 10 Aug 2026 07:02:40 +0200 (CEST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786338149; x=1786942949; darn=lists.proxmox.com; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to:content-type; bh=BWqrjxHfFN7vn+hUOear3vWe0ZfPS3G7omOeYFoHiJo=; b=aFnklaPwmKi8Z+CDVQbxP+W1RYNYwxI95T6Av1rHH8TrXRpT2YypmDvqwBh8OlELr6 FGopNexMvApI1AaA3Tbx9KfaA6yINrgn131lizlBgZZW8vUAjyEOF7/cZmLfSlLnDRSx TqW2eJ/t8BEp6hbUKIdJUS93hlwpn/UZBaqt51DTQfJS/mBhH1DmzlJy+ATTE7iLt7aZ 4Rge/Gi8bdP9DeuzTRToUgmNmHE1MfV8tAw3RIxHQ/wx2v8zLMYPgZWg2fmoC1RxU7H4 QrkNDjOoCnxS7EQ4jaQ4qJOq09IlMa1kYwe6Yw+3NC3VLVwVfK7v+8teJ4OBWjvdN14V 5qGQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786338149; x=1786942949; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=BWqrjxHfFN7vn+hUOear3vWe0ZfPS3G7omOeYFoHiJo=; b=hRaZCtOdCfJjyDcrwVlgPHzembsvM0tSe1LekgT6BnwhsLBwzpIdqkQrgTTE+yh3kb cH43gNCNuWJRfobfkWmmdiss6e98tFtQACCH82i0Ck/mQkaS+tt66Su1Z+3+r9eJd88V Mi7DtjIYCVZpTUjBT5ucp9ALBehPuppJa1gVid/WE+GzW3ef1Gbs1RlMoFQyYvDtuBhC drN3VLUu8Bmr5UFlITQLTO8oz/3AP+w1s2BYHyn7V9jEJLPLs3+G6fnldqMCm6TNvdiJ xrf8fE4Q9CDLDi+gUnlFaCXWnU+yPyVrJZupW7eZQP6Ypfx+2jTQxIEBAGWD6wkzZvsD lz/w== X-Gm-Message-State: AOJu0YzqIwODFK6JI49rjPNZszjDdzA4LPFU2hz/FEOjmdpMQWwhqaUj NWq82E7P7TZ4MpALMKCKgigvaptWRkKwPHttK3CYB8Ys0y0gyOU/lycxuBpkHcWU X-Gm-Gg: AR+sD13+R76j3iqKLDbjCM9nPmiCJnCZHrDXxQJnacmt4DVZoPKuexNy7vTGsXv13a0 roQTKTEwAeT/d56NKQGY3Lwo27oIh2u7nWY5XrxfaphJixPGzACm/gIsuy6eiEBwzpRrl/yBW6E 1ay13qTXzYt3TUmnLprPdNfKlPz0tQ8cNIDxVGSYnjRPtTL1cDu+QYV0HapFXTMt+pNL+iITDPs nei+wYRGzql+JOuJvU1aZ6Oh67E9KyzG96CXak3xBarGrI/JDzDlJGUXqLukqaXZwmSUdjrkFH0 m+zbLBzOJB2NlFb6qnKRkOsLT57C75u+g9CzFW2hzR4IQTTiHjnl+4IXTK5A53438PHnzT7EJJt 6yms4nK7+lE1mPU7HH50lNqcFveF5MqNVVZBogSwAXwG3UWtmK7veeoFsNSmCtAOYE6yez5+8dB 7NSLLwIanALkM2Brw8VS7gyOYexbp5br1K60iWH9LNb8J1XK58xnsTGZZpklpQ88q5TaAymqDXE drA65aB2IZwuhHDNiWVwZcoNJdrqs0jTQXt/cxAI4fC3Fp/JUi86mefI0meP8zI X-Received: by 2002:adf:e010:0:10b0:47f:8ce7:8684 with SMTP id ffacd0b85a97d-47fec509a94mr45912856f8f.6.1786338148458; Sun, 09 Aug 2026 22:02:28 -0700 (PDT) From: Enrico Plantulli To: pbs-devel@lists.proxmox.com Subject: [PATCH proxmox proxmox-backup 0/3] datastore: gc: prefetch chunk metadata before phase 1 Date: Mon, 10 Aug 2026 07:01:57 +0200 Message-ID: <20260810050201.347124-1-plantulli@gmail.com> X-Mailer: git-send-email 2.47.3 MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-SPAM-LEVEL: Spam detection results: 1 DKIM_SIGNED 0.1 Message has a DKIM or DK signature, not necessarily valid DKIM_VALID -0.1 Message has at least one valid DKIM or DK signature DKIM_VALID_AU -0.1 Message has a valid DKIM or DK signature from author's domain DKIM_VALID_EF -0.1 Message has a valid DKIM or DK signature from envelope-from domain DMARC_PASS -0.1 DMARC pass policy FORGED_GMAIL_RCVD 1 'From' gmail.com does not match 'Received' headers FREEMAIL_FROM 0.001 Sender email is commonly abused enduser mail provider KAM_NUMSUBJECT 0.5 Subject ends in numbers excluding current years RCVD_IN_DNSWL_NONE -0.0001 Sender listed at https://www.dnswl.org/, no trust SPF_HELO_NONE 0.001 SPF: HELO does not publish an SPF Record SPF_PASS -0.001 SPF: sender matches SPF record Message-ID-Hash: 6LIU2ISLKKKL37MQU4NKR5ZOKDUPHKXV X-Message-ID-Hash: 6LIU2ISLKKKL37MQU4NKR5ZOKDUPHKXV X-MailFrom: plantulli@gmail.com X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; emergency; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header CC: Enrico Plantulli X-Mailman-Version: 3.3.10 Precedence: list List-Id: Proxmox Backup Server development discussion List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: Hi, garbage collection phase 1 resolves each chunk from its digest and calls utimensat() on it directly, so it never iterates the chunk directories. On a cold store that means every chunk costs an independent, serialized metadata read, and those reads are issued in digest order - which is uncorrelated with the on-disk layout of the inodes by construction, since the digest picks the directory while the inode number follows creation order. How uncorrelated: on datastore A below, the 1059 chunks in a single .chunks/ subdirectory map to 1059 *distinct* 16 KiB metadnode blocks, zero shared, with a median gap of 77966 between consecutive object ids. Most of the numbers below come from two datastores, please keep them apart: Datastore A: 72M chunks, 159 TiB. This is the one that motivates the patch: phase 1 has been running at 39.2 chunks/s, which extrapolates to more than 21 days for phase 1 alone; full GC cycles on this store have taken between 21 and 66 days. No number below is a before/after for this patch on A. Datastore B: 65536 chunk directories, of which the 32768 sampled below hold 5231473 chunks - about half the store. All readdir/lstat numbers below are from B. Measured on ZFS on rotational disks (7x raidz1 of 4 HDDs, ZFS 2.4.3-pve1, PBS 4.2.5, cold cache): readdir over those 32768 directories: 270.9 s, 869726 physical reads lstat() over the same 5231473 entries: 80.1 s, 0 physical reads That is 0.166 physical reads per chunk brought in. A pass over all 65536 directories costs about twice the readdir figure; the pass pays an openat+getdents even for empty directories, so its fixed cost is always on 65536 directories. Physical reads are the sum of the completed-read counters in /proc/diskstats over the 28 pool members, so one 16 KiB logical metadata read shows up as more than one device read: 869726 physical reads against 634549 ARC misses, 1.37 per miss. Of those misses 600605 were prefetch and 33944 demand (94.7% prefetch); the demand figure is within 4% of the 32768 ZAP headers getdents has to read anyway. The readdir pass was run twice from a cold ARC and the two physical read counts agreed to within 0.01%. lstat() is used as a stand-in for the read side of the utimensat() that phase 1 does; it does not perform the atime write. zfs_readdir() calls dmu_prefetch_dnode() for the entries it returns, which is where the win comes from - getdents itself needs only the directory entry, not the inode. The same effect showed up on a third store: after a readdir-only pass over 6979381 chunks, an lstat() over all of them cost zero physical reads. Patch 1 (proxmox) adds a gc-chunk-metadata-prefetch tuning option. Patch 2 (proxmox-backup) implements the pass, wires it up and documents it. Patch 3 (proxmox-backup) adds the GUI field. Patches 2 and 3 depend on patch 1, so proxmox-backup needs a pbs-api-types release and a dependency bump before it builds. I did not touch debian/changelog or Cargo.toml. Testing status: the series is compile-tested (cargo build, clippy, fmt, test of the touched crates) against current master of both repositories, using the pbs-api-types path override already present in Cargo.toml, and I have since run it end-to-end once on datastore B in production, with binaries rebuilt from the 4.2.5-1 tag plus this series: the prefetch walked 10462700 entries in 27m40s (about half of them already cached from an earlier experiment; a fully cold pass extrapolates to roughly 55m), phase 1 then completed in 11m59s against a 2h14m baseline measured on the same store four days earlier - about 13700 chunks/s warm against 440-563 chunks/s cold - and the ZFS dirty-data throttle never engaged (dmu_tx_dirty_delay stayed at 0). The full collection, including a phase 2 that swept 9.85M chunks through the nightly backup window, finished in 6h21m with TASK OK and removed 436 GiB. I have not tested the S3 code path, and datastore A has not run with the option yet. Design notes and open questions, on which I would appreciate guidance: * The win is conditional on the prefetched metadata surviving in the cache until phase 1 consumes it. On ZFS that is roughly 512 bytes of dnode per chunk: about 2.5 GiB for the 5.2M chunks sampled from datastore B (twice that for the full store) and about 34 GiB for datastore A. On a store whose dnodes do not fit in the ARC the pass does not pay off. And even where they fit, whether the tail of the prefetch survives until a phase 1 that runs for hours reaches it - against the index files phase 1 itself reads, the dnodes it dirties and concurrent backup traffic - is an open question on a store of A's size. If it does not, the better design would be prefetch windows interleaved with marking, which this simple one-pass version does not attempt. Treat the datastore A figures as motivation, not as a measured result of this patch. * The warm figures bound the read side only. Phase 1 also dirties every chunk's dnode through utimensat(), and in digest order two touches that share a 16 KiB metadnode block are millions of operations apart, so copy-on-write rewrites the block once per touch instead of once per block: on datastore A that is ~72M block rewrites, on the order of 1 TiB of metadnode churn plus a multiple of that in dirtied indirect blocks. This is the same volume phase 1 writes today, only compressed into less wall time - a faster phase 1 even coalesces the indirect block updates better. But it does cap the marking rate well below what the lstat() figure alone would suggest: with the default 4 GiB zfs_dirty_data_max the ZFS write throttle starts to shape the rate in the low thousands of utimensat() per second. So the end state I expect on datastore A is phase 1 bound by metadata write-back at a few thousand chunks per second - two orders of magnitude above the 39.2 chunks/s it does today, but not the tens of thousands the warm read numbers alone might suggest. On datastore B the end-to-end run stayed below any throttling (~13700 chunks/s with dmu_tx_dirty_delay at 0), so the cap was not reached there; for A this remains an estimate, not a measurement. * Opt-in was the conservative choice. Note that tying it to chunk-order would not be conservative at all: chunk-order defaults to `inode`, so that would effectively enable the pass everywhere. If you want it on by default, I would rather make the option default to true than key it off chunk-order. * The pass is not free, but it is also not a whole extra pass: GC phase 2 already walks the same directories with the same iterator, so the prefetch can warm phase 2 as well - though after an hours-long phase 1 that rewrites the dnodes it touches, how much of that warmth is left for phase 2 is equally open. * The pass is best effort: a failure is logged and ignored, except for abort and shutdown requests, which are re-checked in the error path so that cancelling the task still stops the collection. * It reuses get_chunk_store_iterator(), the same iterator phase 2 uses, so there is no second implementation of the chunk directory walk. The per-entry hex filtering is redundant for this use, but reusing the iterator seemed better than duplicating the walk. * The pass is sequential, one directory at a time, so it keeps the queue depth of the pool low - the same limitation that makes phase 1 slow in the first place. Parallelising it over ranges of the 65536 subdirectories would likely help further on wide pools, but it would need a second walk implementation instead of reusing get_chunk_store_iterator(), so I left it out of this first version. Happy to add it if you would take it. * This is complementary to the LRU cache added in [0] for #5331: that commit removes redundant atime updates, this one makes the remaining ones cheap. It does not change what phase 1 marks, only what it has to wait for. * A larger version of the same idea would be to have phase 1 itself walk in inode order, the way chunk-order=inode already does for verify. That is a much more invasive change and I did not attempt it; the readdir pass gets most of the benefit for a fraction of the risk. [0] https://git.proxmox.com/?p=proxmox-backup.git;a=commit;h=03143eee0a59cf319be0052e139f7e20e124d572 Diffstat over the whole series: proxmox: pbs-api-types/src/datastore.rs | 10 ++++++++++ 1 file changed, 10 insertions(+) proxmox-backup: docs/storage.rst | 12 ++++++++++++ pbs-datastore/src/chunk_store.rs | 38 ++++++++++++++++++++++++++++++++++++++ pbs-datastore/src/datastore.rs | 23 +++++++++++++++++++++++ www/Utils.js | 6 ++++++ www/datastore/OptionView.js | 16 ++++++++++++++++ 5 files changed, 95 insertions(+) Thanks, Enrico Plantulli