From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from gate001.proxmox.com (gate001.proxmox.com [45.144.208.40]) by lore.proxmox.com (Postfix) with ESMTPS id 62D6C1FF0E1 for ; Mon, 10 Aug 2026 17:05:51 +0200 (CEST) Received: from gate001.proxmox.com (localhost.localdomain [127.0.0.1]) by gate001.proxmox.com (Proxmox) with ESMTP id 1C07921799; Mon, 10 Aug 2026 17:05:51 +0200 (CEST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786374333; x=1786979133; darn=lists.proxmox.com; h=to:in-reply-to:cc:references:message-id:date:subject:mime-version :from:content-transfer-encoding:content-type:from:to:cc:subject:date :message-id:reply-to:content-type; bh=OoZWEodJDzVVT7pMt/kUCOVIVop2x0LojaypZLX+7CE=; b=PCPUZy88tuYXrpqS37XyqOXOX4Ko6EUmKViMCXCGzxVCxdwsez76r//cIyTmmYJKK5 5eFO73sU1X6Blx+mbFx5CV2wwB0EEo3pn76mxy6dJYOc/rhylFfQrRHw6xlud1AB/pc8 lp6algxV0Pmrd6Rh4Zo3KRlEKxMk08rZWyefWpkPcLnWtJjNM0M1yWPCT0D0QlqCDSUW /TOyYk2+LHwpEdH0JBjYxoI5z7Dn3ahhLptvoQ2bojE9SM/hBNkwkoWh6rMrw7iKiJnf TBfhYwT9guLXEo6f9rJLGFdwatudm47Knjl95AVncKbG6sIhxA85ymUfa8HMdXCmtTFm dWrg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786374333; x=1786979133; h=to:in-reply-to:cc:references:message-id:date:subject:mime-version :from:content-transfer-encoding:content-type:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=OoZWEodJDzVVT7pMt/kUCOVIVop2x0LojaypZLX+7CE=; b=srqNKP1mYrRmDlP6oOg1LR/gq+gccAp/mNCxzZBosUAYpioZ8nyc/YYZ1AmCy48ZZY yDVXFPpRzINkmq+YYln6BtizIqGKFqT/Ur0+8pjP9LOeGMKOrUZO1BQf4bf9n8XVEFVu o4xXdL+j8yX0ICrwz2QzjTPQj2cCvgckZTJ4S7gnnQKfgvulJriSl60YznrrPBZqOags GZPwX+yTgILoClFymxr7Cf9qcjzrVN5wF4903aOZ9OYrj/rOjsd41F7Zr4kzSQ1Y2qSk Lq2NAYCqASHzQQWCL9EzDk3uGKF0R4eabougGY92L+XgayhE090eGpWmdR6BcFvj6Wuf ELzg== X-Gm-Message-State: AOJu0Yxvb/mqChFypoX4+6aHbS02vCv0tWodtldlsN4KgAyM4bbAiloE ehQRbq7Pvef5sU/9y9EZCfuFAcFLmJ1x0FwdiZy2wn9m3U7LWPuo73uh X-Gm-Gg: AR+sD10po1qOrO8NJ1KzUwg8GpxFFgiz6wwcOIkmQWRy+qWYEAg4A/aLjFLhhUd1jov lVCROgkofcqG1lY2MJ+H+nhWO5NcLfg0wYFKLutl2+VYi7Re2z6yX2x3t88PVEatoL9eE7Zvr2r JLJz8g4QwWXizDFMXxOa0clGeicuhd9s+EE2M9kzyxeuyWkpm9UkI9z23QVwgVrWmStteJKzjbN nnVNocqYkKdQ9Di7J70CYLhRLhxp9+zd3zQ/5oCrZ+hNNoW86ObxS/2ZHSGyrtODJCp0gpqKf3r 3kdHu+Un+6Z3gcFtKZ+haREPl9URiK2qPJuU5HKFKj2p9mtWTcpIvNzBYtSbOlUqqOJIoXdQTvw N4xF0JMqS0v8pyxuZOXZixxqk+qk7cAjxPrI0CDuG2vVcmkNRsJh4ZrFk/bee9QCznVUI7y5VpS 17O/XCvKagju46BhhoZCg4dhVwuFZ2B1kIG3jRSCzOJeW6mzxK3BFtY4R7h4Ax9K8/t8IM6s6lR ssn+KJ7meqqMmqa28XPvS26KYwqF76U80chetCrdMKk/3yqPA== X-Received: by 2002:a05:600c:1d24:b0:496:c9cd:e7ab with SMTP id 5b1f17b1804b1-4995e08452bmr308461325e9.5.1786374332530; Mon, 10 Aug 2026 08:05:32 -0700 (PDT) Content-Type: text/plain; charset=utf-8 Content-Transfer-Encoding: quoted-printable From: Enrico Plantulli Mime-Version: 1.0 (1.0) Subject: Re: [PATCH proxmox-backup 2/3] datastore: gc: optionally prefetch chunk metadata before phase 1 Date: Mon, 10 Aug 2026 17:05:20 +0200 Message-Id: <1F771D00-6120-43B8-8A23-9B16B7F7F60A@gmail.com> References: In-Reply-To: To: Christian Ebner X-Mailer: iPhone Mail (23G71) X-SPAM-LEVEL: Spam detection results: 0 DKIM_SIGNED 0.1 Message has a DKIM or DK signature, not necessarily valid DKIM_VALID -0.1 Message has at least one valid DKIM or DK signature DKIM_VALID_AU -0.1 Message has a valid DKIM or DK signature from author's domain DKIM_VALID_EF -0.1 Message has a valid DKIM or DK signature from envelope-from domain DMARC_PASS -0.1 DMARC pass policy FREEMAIL_FROM 0.001 Sender email is commonly abused enduser mail provider KAM_NUMSUBJECT 0.5 Subject ends in numbers excluding current years RCVD_IN_DNSWL_NONE -0.0001 Sender listed at https://www.dnswl.org/, no trust SPF_HELO_NONE 0.001 SPF: HELO does not publish an SPF Record SPF_PASS -0.001 SPF: sender matches SPF record Message-ID-Hash: G2M2RCBUZRDLOA4PWWFE4OT6IFGS2S74 X-Message-ID-Hash: G2M2RCBUZRDLOA4PWWFE4OT6IFGS2S74 X-MailFrom: plantulli@gmail.com X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; emergency; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header CC: pbs-devel@lists.proxmox.com X-Mailman-Version: 3.3.10 Precedence: list List-Id: Proxmox Backup Server development discussion List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: Already done yesterday ! Cheers Enrico > Il giorno 10 ago 2026, alle ore 12:19, Christian Ebner ha scritto: >=20 > =EF=BB=BFOn 8/10/26 7:02 AM, Enrico Plantulli wrote: >> Phase 1 of garbage collection resolves each chunk from its digest and >> calls utimensat() on it directly. It never iterates the chunk >> directories, so on a cold store every chunk costs an independent, >> serialized metadata read, issued in digest order - which is uncorrelated >> with the on-disk layout of the inodes by construction, since the digest >> picks the directory while the inode number follows creation order. >> On file systems that prefetch inode metadata while iterating a directory >> this can be avoided almost entirely: iterating the chunk directories >> once brings in that metadata in one go, and phase 1 then finds it >> cached. >> Measured on a ZFS datastore on spinning disks (7x raidz1 of 4 HDDs, ZFS >> 2.4.3, cold cache), over 32768 of its 65536 chunk directories, holding >> 5231473 chunks: >> readdir over those 32768 directories: 270.9 s, 869726 physical reads >> lstat() over the same 5231473 entries: 80.1 s, 0 physical reads >> That is 0.166 physical reads per chunk brought in. 94.7% of the >> resulting ARC misses were prefetch misses, and zfs_readdir() calls >> dmu_prefetch_dnode() for the entries it returns, which is where the win >> comes from - getdents itself needs only the directory entry, not the >> inode. The pass only helps as long as that metadata is still cached when >> phase 1 uses it. >> Note that the pass removes the read stalls only: phase 1 still >> dirties every chunk's dnode via utimensat(), so once the reads are >> cached the marking rate is expected to cap at the pool's metadata >> write-back rate, not at readdir/lstat speed. >> The pass is not free and it does not help everywhere, so it is opt-in >> via the new gc-chunk-metadata-prefetch tuning option, and it is skipped >> for S3 backed datastores: there the local chunk store only holds the >> cached chunks and the in-use markers, not the full chunk set, so a full >> pass over the local chunk directories is not what phase 1 has to wait >=20 > comment: This is however not correct, for S3 the marker file is present fo= r each chunk known to be present on the backend. So the same logic and gain c= an be expected for these as well. The difference in handling is here mostly i= n phase 2, where the chunks as present on the S3 backend must be listed and d= eleted if no longer required. >=20 >> for. A failing prefetch is logged and ignored, except for abort and >> shutdown requests, which are re-checked in the error path: the pass is >> an optimization and must never become a new way for garbage collection >> to fail, but it must not swallow a cancellation either. >> Signed-off-by: Enrico Plantulli >> --- >> docs/storage.rst | 12 ++++++++++ >> pbs-datastore/src/chunk_store.rs | 38 ++++++++++++++++++++++++++++++++ >> pbs-datastore/src/datastore.rs | 23 +++++++++++++++++++ >> 3 files changed, 73 insertions(+) >> diff --git a/docs/storage.rst b/docs/storage.rst >> index 2c6ba0a..575ece9 100644 >> --- a/docs/storage.rst >> +++ b/docs/storage.rst >> @@ -686,6 +686,18 @@ There are some tuning related options for the datast= ore that are more advanced: >> cache slots, 1048576 (=3D 1024 * 1024) being the default, 8388608 (=3D= 8192 * >> 1024) the maximum value. >> +* ``gc-chunk-metadata-prefetch``: Prefetch chunk metadata before GC pha= se 1: >> + Phase 1 reaches each chunk directly by digest and never iterates the c= hunk >> + directories. On file systems that prefetch inode metadata while iterat= ing a >> + directory, such as ZFS, a single pass over the chunk directories befor= e >> + phase 1 brings in that metadata in one go, instead of leaving phase 1 t= o >> + fault it in one chunk at a time. This can reduce phase 1 runtime >> + substantially on large datastores on rotational disks, at the cost of o= ne >> + extra pass. The pass only pays off if the metadata is still cached onc= e >> + phase 1 uses it: on ZFS that is roughly 512 bytes of dnode per chunk, s= o a >> + datastore whose dnodes do not fit into the ARC will not benefit. It is= >> + disabled by default and has no effect on S3 backed datastores. >=20 > comment: please move the docs to a standalone patch, independent on how th= e final implementation will look like. >=20 >> + >> * ``default-verification-workers`` and ``default-verification-readers``:= >> Define the default number of threads used for verification and reading= of chunks, >> respectively. By default, 4 threads are used for verification and 1 th= read is used >> diff --git a/pbs-datastore/src/chunk_store.rs b/pbs-datastore/src/chunk_s= tore.rs >> index 6d2ffdb..b595d15 100644 >> --- a/pbs-datastore/src/chunk_store.rs >> +++ b/pbs-datastore/src/chunk_store.rs >> @@ -302,6 +302,44 @@ impl ChunkStore { >> Ok(is_bad) >> } >> + /// Iterate all chunk directories once, discarding the entries. >> + /// >> + /// No chunk file is opened, read, written or stat'ed, so no chunk a= ccess time is touched. >> + /// The readdir does update the access time of the up to 65536 `.chu= nks/` subdirectories, >> + /// which garbage collection never looks at: phase 2 only ever unlin= ks regular files, and >> + /// the atime cutoff applies to chunks, not to directories. >> + /// >> + /// The only purpose is to let the file system fault in the chunk in= ode metadata in bulk, so >> + /// that garbage collection phase 1 finds it cached instead of fault= ing it in one chunk at a >> + /// time in digest order, which is uncorrelated with the on-disk lay= out of the inodes. >> + /// >> + /// Returns the number of chunk directory entries seen (chunks, bad c= hunks and markers). >> + pub fn prefetch_chunk_metadata(&self, worker: &dyn WorkerTaskContext= ) -> Result { >> + let mut last_percentage =3D 0; >> + let mut chunk_count =3D 0; >> + >> + for (entry, percentage, _chunk_ext) in self.get_chunk_store_iter= ator()? { >> + if last_percentage !=3D percentage { >> + last_percentage =3D percentage; >> + info!("prefetched {percentage}% ({chunk_count} entries)"= ); >> + } >> + >> + worker.check_abort()?; >> + worker.fail_on_shutdown()?; >> + >> + if let Err(err) =3D entry { >> + bail!( >> + "chunk iterator on chunk store '{}' failed - {err}",= >> + self.name, >> + ); >> + } >> + >> + chunk_count +=3D 1; >> + } >> + >> + Ok(chunk_count) >> + } >> + >> fn get_chunk_store_iterator( >> &self, >> ) -> Result< >> diff --git a/pbs-datastore/src/datastore.rs b/pbs-datastore/src/datastore= .rs >> index c3db522..da77134 100644 >> --- a/pbs-datastore/src/datastore.rs >> +++ b/pbs-datastore/src/datastore.rs >> @@ -2459,6 +2459,29 @@ impl DataStore { >> 1024 * 1024 >> }; >> + if tuning.gc_chunk_metadata_prefetch.unwrap_or(false) { >> + if s3_client.is_some() { >> + info!("Chunk metadata prefetch not supported for S3 back= ed datastore, skipping."); >> + } else { >> + info!("Start GC chunk metadata prefetch"); >> + let start =3D std::time::Instant::now(); >> + // Best effort only: this is an optimization, it must ne= ver fail the collection. >> + match self.inner.chunk_store.prefetch_chunk_metadata(wor= ker) { >> + Ok(chunk_count) =3D> info!( >> + "Chunk metadata prefetch done, {chunk_count} ent= ries in {}", >> + TimeSpan::from(start.elapsed()), >> + ), >> + Err(err) =3D> { >> + // an abort or shutdown request must not be swal= lowed by the >> + // best effort handling below >> + worker.check_abort()?; >> + worker.fail_on_shutdown()?; >> + warn!("Chunk metadata prefetch failed, continuin= g without it - {err}"); >> + } >> + } >> + } >> + } >> + >> info!("Start GC phase1 (mark used chunks)"); >> self.mark_used_chunks( >> base-commit: 5d95fc20eeae6df1936216b1ab1ab6952472fdcf >=20 >=20