From: Ciro Iriarte <cyruspy@gmail.com>
To: pve-devel@lists.proxmox.com
Subject: [RFC PATCH storage 3/5] dir: implement copy-offload via reflink (FICLONE)
Date: Mon, 20 Jul 2026 07:30:08 -0700 (PDT) [thread overview]
Message-ID: <20260720.3.copyoffload@cyruspy.gmail.com> (raw)
In-Reply-To: <20260720.0.copyoffload@cyruspy.gmail.com>
A filesystem that supports FICLONE -- XFS with reflink=1, btrfs, or ZFS with block
cloning -- can produce a full copy by sharing extents copy-on-write. The result is
INDEPENDENT immediately: the extents are reference counted, so deleting the source
does not affect the copy.
That makes this the cheapest member of the 'copy-offload-atomic' class. There is no
data movement and nothing to wait for, so copy_image_status() reports 'complete' on
the first poll and the asynchronous machinery simply does not engage. Today a full
clone on a directory storage byte-copies the whole image through qemu-img convert
even when the filesystem underneath could share the extents for free.
Three things are refused rather than half-done:
- A qcow2 that HAS A BACKING FILE. A byte-identical copy of an overlay inherits the
dependency: it passes qemu-img check and then breaks the moment the base is
removed, which is exactly the "readable but not independent" failure the hook's
contract exists to prevent. Only a real convert can flatten it, so the feature is
not advertised for such volumes -- deliberately not merely rejected in prepare,
because failing there would abort the clone instead of letting it take the normal
host-side path.
- Format conversion. FICLONE copies bytes; qcow2 -> raw needs qemu-img.
- A source snapshot, and a target on a different filesystem.
prepare() creates the target file rather than only choosing a name. The caller
releases the storage lock between prepare and start, so a name that was merely
chosen could be taken by a concurrent allocation, and the caller's rollback would
then delete that other operation's volume. An empty file costs nothing.
start() uses 'cp --reflink=always' so an unsupported filesystem fails loudly instead
of silently degrading into a full byte copy that would block a caller which may be
holding a guest frozen.
Verified on a loop-backed XFS (reflink=1) by driving the hooks directly: the clone
is byte-identical, status is 'complete' on the first poll, and the clone still
matches after the source is deleted. The guards were checked too -- a snapshot
source and a format conversion are both refused, a qcow2 with a backing file is not
advertised, and a plain qcow2 is. Independence and zero space use were separately
confirmed on XFS, btrfs and ZFS, where a 256 MiB clone added 0 MiB of allocation.
Generated-By: Claude (https://claude.ai)
Signed-off-by: Ciro Iriarte <ciro.iriarte+software@gmail.com>
Co-Authored-By: Claude <noreply@anthropic.com>
---
src/PVE/Storage/DirPlugin.pm | 140 +++++++++++++++++++++++++++++++++++
1 file changed, 140 insertions(+)
diff --git a/src/PVE/Storage/DirPlugin.pm b/src/PVE/Storage/DirPlugin.pm
index 80c4a03..cbc7357 100644
--- a/src/PVE/Storage/DirPlugin.pm
+++ b/src/PVE/Storage/DirPlugin.pm
@@ -8,6 +8,7 @@ use Encode qw(decode encode);
use File::Path;
use File::Spec;
use IO::File;
+use JSON;
use POSIX;
use PVE::Storage::Plugin;
@@ -95,6 +96,8 @@ sub options {
bwlimit => { optional => 1 },
preallocation => { optional => 1 },
'snapshot-as-volume-chain' => { optional => 1, fixed => 1 },
+ 'copy-offload' => { optional => 1 },
+ 'copy-offload-timeout' => { optional => 1 },
};
}
@@ -323,4 +326,141 @@ sub volume_qemu_snapshot_method {
return $scfg->{'snapshot-as-volume-chain'} ? 'mixed' : 'qemu';
}
+
+# qemu_img_info() returns raw JSON text, so decode before use. Returns the backing
+# filename, or undef when there is none / the image cannot be inspected.
+my sub qcow2_backing_file {
+ my ($path) = @_;
+ my $json = eval { PVE::Storage::Common::qemu_img_info($path, undef, 10) };
+ return undef if $@ || !$json;
+ my $info = eval { decode_json($json) };
+ return undef if $@ || ref($info) ne 'HASH';
+ return $info->{'backing-filename'};
+}
+
+sub volume_has_feature {
+ my ($class, $scfg, $feature, $storeid, $volname, $snapname, $running, $opts) = @_;
+
+ if ($feature eq 'copy-offload-atomic') {
+ # reflink clones a whole file; there is no way to pick a snapshot out of one
+ return 0 if $snapname;
+
+ my ($vtype, undef, undef, undef, undef, undef, $format) =
+ eval { $class->parse_volname($volname) };
+ return 0 if $@ || !defined($vtype) || $vtype ne 'images';
+ return 0 if $format ne 'raw' && $format ne 'qcow2';
+
+ # A qcow2 with a backing file cannot be flattened by a byte-identical copy, so
+ # do not advertise it -- otherwise the clone would fail in prepare instead of
+ # quietly taking the normal host-side path.
+ if ($format eq 'qcow2') {
+ my $path = eval { $class->filesystem_path($scfg, $volname) };
+ return 0 if $@ || !defined($path);
+ return 0 if qcow2_backing_file($path);
+ }
+
+ return 1;
+ }
+
+ return $class->SUPER::volume_has_feature(
+ $scfg, $feature, $storeid, $volname, $snapname, $running, $opts,
+ );
+}
+
+# ---- storage-offloaded full copy via reflink -------------------------------------
+#
+# A filesystem that supports FICLONE (XFS with reflink=1, btrfs, ZFS with block
+# cloning) can produce a full copy by sharing extents copy-on-write. Unlike an RBD
+# clone or a ZFS clone-from-snapshot, the result is INDEPENDENT straight away: the
+# extents are reference counted, so deleting the source does not affect the copy.
+#
+# That makes this the cheapest possible member of the 'copy-offload-atomic' class --
+# instant, no extra space, and nothing to wait for. copy_image_status() reports
+# 'complete' on the first poll because there is no background work.
+
+# Same filesystem? FICLONE cannot cross one, and the caller may be copying between
+# two different storages that happen to be directories.
+my sub same_filesystem {
+ my ($a, $b) = @_;
+ my $da = (stat($a))[0];
+ my $db = (stat($b))[0];
+ return defined($da) && defined($db) && $da == $db;
+}
+
+sub copy_image_prepare {
+ my (
+ $class, $scfg, $storeid, $volname,
+ $target_scfg, $target_storeid, $target_vmid, $snap, $opts,
+ ) = @_;
+
+ die "copy offload cannot copy from a snapshot\n" if defined($snap);
+
+ my ($vtype, undef, undef, undef, undef, undef, $format) = $class->parse_volname($volname);
+ die "copy offload only handles VM images, not '$vtype'\n" if $vtype ne 'images';
+
+ # FICLONE copies bytes; it cannot convert between formats.
+ my $target_format = $opts->{format} // $format;
+ die "copy offload cannot convert '$format' to '$target_format'\n"
+ if $target_format ne $format;
+
+ my $path = $class->filesystem_path($scfg, $volname);
+
+ # A qcow2 with a backing file is NOT independent, and a byte-identical copy of it
+ # inherits that dependency -- it would pass qemu-img check and then break when the
+ # base is removed. Only a real convert can flatten it.
+ die "copy offload cannot flatten '$volname': it has a backing file\n"
+ if $format eq 'qcow2' && qcow2_backing_file($path);
+
+ my $target_dir = $class->get_subdir($target_scfg, 'images') . "/$target_vmid";
+ mkpath $target_dir;
+
+ die "copy offload requires source and target on the same filesystem\n"
+ if !same_filesystem($path, $target_dir);
+
+ # the trailing 1 adds the format suffix; without it the volname does not parse
+ my $name =
+ $class->find_free_diskname($target_storeid, $target_scfg, $target_vmid, $format, 1);
+
+ # Reserve the name by creating the file. The caller drops the storage lock between
+ # prepare and start, so a name that was merely chosen could be taken by a
+ # concurrent allocation -- and the caller's rollback would then delete somebody
+ # else's volume. An empty file costs nothing and makes the target freeable.
+ my $target_path = "$target_dir/$name";
+ my $fh = IO::File->new($target_path, O_WRONLY | O_CREAT | O_EXCL, 0640)
+ or die "unable to reserve '$target_path' - $!\n";
+ close($fh);
+
+ return "$target_vmid/$name";
+}
+
+sub copy_image_start {
+ my (
+ $class, $scfg, $storeid, $volname,
+ $target_scfg, $target_storeid, $target_volname, $snap,
+ ) = @_;
+
+ my $src = $class->filesystem_path($scfg, $volname);
+ my $dst = $class->filesystem_path($target_scfg, $target_volname);
+
+ # --reflink=always so an unsupported filesystem fails loudly rather than silently
+ # turning this into a full byte copy that blocks the caller -- which may be holding
+ # a guest frozen.
+ eval { PVE::Tools::run_command(['/bin/cp', '--reflink=always', '--', $src, $dst]) };
+ if (my $err = $@) {
+ unlink($dst);
+ die "reflink copy of '$volname' failed - $err";
+ }
+
+ return;
+}
+
+sub copy_image_status {
+ my ($class, $scfg, $storeid, $volname, $source) = @_;
+
+ # FICLONE is synchronous and the extents are reference counted, so the copy is
+ # already independent of its source. Nothing to poll and nothing to clean up.
+ return { state => 'complete' };
+}
+
+
1;
--
2.54.0
next prev parent reply other threads:[~2026-07-20 14:31 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-20 14:29 [RFC PATCH storage, qemu-server 0/5] offload full clone to the storage backend Ciro Iriarte
2026-07-20 14:29 ` [RFC PATCH storage 1/5] storage: add asynchronous copy-offload hook for full copies Ciro Iriarte
2026-07-20 14:29 ` [RFC PATCH storage 2/5] rbd: implement copy-offload via snapshot + clone + flatten Ciro Iriarte
2026-07-20 14:30 ` Ciro Iriarte [this message]
2026-07-20 14:30 ` [RFC PATCH storage 4/5] btrfs, lvmthin: implement copy-offload (atomic class) Ciro Iriarte
2026-07-20 14:30 ` [RFC PATCH qemu-server 5/5] use storage copy-offload for full clone Ciro Iriarte
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260720.3.copyoffload@cyruspy.gmail.com \
--to=cyruspy@gmail.com \
--cc=pve-devel@lists.proxmox.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.