From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from gate001.proxmox.com (gate001.proxmox.com [IPv6:2a0f:8001:1:32::40]) by lore.proxmox.com (Postfix) with ESMTPS id ADB071FF130 for ; Mon, 20 Jul 2026 16:30:20 +0200 (CEST) Received: from gate001.proxmox.com (localhost.localdomain [127.0.0.1]) by gate001.proxmox.com (Proxmox) with ESMTP id 04BA121568; Mon, 20 Jul 2026 16:30:17 +0200 (CEST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784557774; x=1785162574; darn=lists.proxmox.com; h=content-transfer-encoding:content-type:mime-version:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=2VA6C88ECJxZtoF2kpPHvSDYC8QgbVpGcNZIobC78Uw=; b=S5s67hkCmfWNmqYMfodbllqdBsEkPTZ2RcO0b79oXTSx8ONklTxTZGRgygYkBI7Yxp wtE3aFkv5gfK6U1SfdvB1YMd9mNoV4EZ10bUWLpnT85uE1VtYhiO6pPpED13mYZu7xjv 2SknEPsikE5gUxf494f/zxgrxa7lGBeCha9DIUhjNs+J8oYigABR5m+i2+l9AEVi12dt K1QsEZsRTPnlSouZGtuY9hDFaLC8s6uQ0/94e199+gMhJ8h/X40bqvOOkiriUpXJb0FU HSW6a8F+K0Un/cOlBPxa14+6aaUrsmKl3rWWVrb8gwd5naMpr1c740tD5EENoUT4KHkU muOA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784557774; x=1785162574; h=content-transfer-encoding:content-type:mime-version:message-id :subject:cc:to:from:date:x-gm-gg:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to:content-type; bh=2VA6C88ECJxZtoF2kpPHvSDYC8QgbVpGcNZIobC78Uw=; b=Er+ePs2fpzCZ+7FaBQrPj5ivCn2W+DDb8oUuH/bDK6x7U3hRCbev4Ok7jGxALpcQpD VCAOntRSoLsLAkTv37eaPhOxe8Tfv3Yk3Kf19pfacGD+8AuuNKGxSuOA4d2CEAjHzr+D qEfpwwW+5idSaFe/BxoI+5R5x7VvNMlrUpqWE21Bxws2aUtWzVkKe2CC3dcOnMYyWLqL Farpwa09Se0aJ7xGwws89r64xRlxZK9JiaLTZpkgMOIiTVUhtVyfUkMbeTpSBZTCzbNK atNIWrQCOzn01RYUUAFshnuu57yZx56HUHNOzYjcOPgD+7qaAk4+HDV2z/ahNN3KL3kf zBUQ== X-Gm-Message-State: AOJu0YzhnZbWUZdzko0m7AOoXZRFSRoY2ZGika5I10wqarJIHKnet29F YzGb3UOKEE/+Qw+PFRzXR1Ovw2oNJsrl3o5oQXiZGzHkH/vplUY4dQMHyMDyuOFkNOA= X-Gm-Gg: AR+sD12BSr3T9jfGXGX9r3fBO5P1VXF/VcVCSLBqOImR092cFjxrTVPfH1/yexBh3MB ESzCYXmf8TSS3OFTQnUM7y4zLKGLxEkxJBKkjqYOJytiAUvwTQ+Vrm85eybFK3Jj0Ijv7tXCDKz Ga8hA+OOZrNWXx9s1uliaxezQdwtoAjT//XBxpPwa+HUw9Z/OCM1FtrvHk2Q4hvX9WSlPHPpPNp lEAS40gYVWToxI3rq8UjXXqa1xwbaEz0EiOWqiQlfp8rKTKfi3DuQ8tkLNqysUyDClyAL52jK3O BG46zhF8MaGNqWJ8b8/NoKrz4x5OgdzE6ZLGS1FmqS9xnK2tJFCQ6D4Y1Rr5nTZnRFn2FumZVH2 aqnypm6VSIiXm1vi9uKGVEnGcAo8pV45pb4/CPTUiDFW0Va2IBSarHEIboG/WFWbcPFdeHhTYeq sHBCYSUc4= X-Received: by 2002:a17:902:ea07:b0:2c9:ed16:8d8d with SMTP id d9443c01a7336-2cf349ea664mr169121385ad.38.1784557774187; Mon, 20 Jul 2026 07:29:34 -0700 (PDT) Date: Mon, 20 Jul 2026 07:29:33 -0700 (PDT) From: Ciro Iriarte To: pve-devel@lists.proxmox.com Subject: [RFC PATCH storage, qemu-server 0/5] offload full clone to the storage backend Message-ID: <20260720.0.copyoffload@cyruspy.gmail.com> MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit X-SPAM-LEVEL: Spam detection results: 0 AWL 0.089 Adjusted score from AWL reputation of From: address DKIM_SIGNED 0.1 Message has a DKIM or DK signature, not necessarily valid DKIM_VALID -0.1 Message has at least one valid DKIM or DK signature DKIM_VALID_AU -0.1 Message has a valid DKIM or DK signature from author's domain DKIM_VALID_EF -0.1 Message has a valid DKIM or DK signature from envelope-from domain DMARC_PASS -0.1 DMARC pass policy FREEMAIL_FROM 0.001 Sender email is commonly abused enduser mail provider RCVD_IN_DNSWL_NONE -0.0001 Sender listed at https://www.dnswl.org/, no trust SPF_HELO_NONE 0.001 SPF: HELO does not publish an SPF Record SPF_PASS -0.001 SPF: sender matches SPF record Message-ID-Hash: U2RMAKYVNCMSHNJLHZ54I5PCHNQWXOAP X-Message-ID-Hash: U2RMAKYVNCMSHNJLHZ54I5PCHNQWXOAP X-MailFrom: cyruspy@gmail.com X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; emergency; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header X-Mailman-Version: 3.3.10 Precedence: list List-Id: Proxmox VE development discussion List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: This implements the contract discussed in "[RFC storage, qemu-server] offload full/live clone to storage backend" (2026-06-30). Still RFC: three questions from that thread are open, listed at the end. Fiona asked to see the qemu-server side alongside the storage patches, so 5/5 is that consumer. It is included for context and is the least finished part. Design ------ Three phases rather than one call: copy_image_prepare() runs under the target storage lock, reserves a real and freeable target volume, returns its volname copy_image_start() fixes the source's point in time and begins the copy copy_image_status() polled outside both, reports 'complete' only when the copy is INDEPENDENT of the source A synchronous hook is not workable. cfs_lock_storage() sets alarm(60) around the locked code, and an array-side clone can run for minutes or hours, so the worker is killed and the lock broken. The existing host path already holds the lock only for vdisk_alloc() and runs qemu-img convert outside it. The split also gives the caller something cheap enough to put inside a guest freeze. start() does the minimum that pins the source; the data movement is reported by status(). Allocation is not cheap enough for that, which is why it is a separate phase. 'complete' means independent, not readable. An RBD clone is readable the moment it exists but still depends on its parent snapshot until the flatten finishes, and deleting the source before that destroys it. Distinguishing the two is the main thing the status phase exists for. Two features rather than a capability argument on volume_has_feature(), per Fiona's suggestion: copy-offload-atomic and copy-offload-bulk. Only the atomic class has implementations here. Enabled per storage with copy-offload, and ignored unless both storages are instances of the same plugin type. Implemented ----------- rbd snapshot + clone v2, flatten handed to the Ceph manager dir reflink (FICLONE), where the filesystem supports it btrfs subvolume snapshot lvmthin thin snapshot The last three are instant and independent immediately. RBD is the case that motivated the asynchronous shape. Verified on Ceph (Tentacle 20.2.1), on loop-backed XFS with reflink=1, and on a loop-backed btrfs and thin pool. In each case the copy is byte-identical, a copy taken from a snapshot holds the snapshot's content rather than live data, and the copy still reads correctly after the source is deleted outright. src/test/copy_offload_functional_test.pl drives the hooks against real loop devices; it needs root, so it is not part of run_plugin_tests.pl and skips when it cannot run. One thing worth flagging ------------------------ The caller drops the target storage lock between prepare() and start(), so the reserved name has to stay reserved across a window the plugin does not control. Whether it does depends on two regexes that do not obviously belong together: volume listers are anchored, while $get_vm_disk_number is not. A placeholder parked under a prefixed name is therefore invisible to list_images() but still reserves the disk number, and a '.copytmp' suffix is the opposite on both counts. Getting that wrong is not cosmetic. If the parked name stops reserving, a concurrent allocation can take it, and the failing copy's rollback then frees that other volume. btrfs and RBD both had this; btrfs now swaps with renameat2(RENAME_EXCHANGE) so the name is never free, and RBD parks under a prefix instead of removing the placeholder first. copy_offload_naming_test.pm pins the property since nothing about the naming looks load-bearing. This is really an argument for question 3 below. Open questions -------------- 1. move_disk. It needs the bulk path, and may be wanted in the same series rather than after it. 2. copy_images_start(), a batch entry point so a multi-disk guest can start every copy inside one freeze. There is no in-tree user today, so it is not in this series; it would be a droppable last patch. 3. Whether prepare() should return an opaque backend token instead of a volname. It no longer blocks anything here, but it is the general answer to the rollback problem above: a volname alone is not proof of ownership once start() is allowed to move things around outside the lock. A token the rollback verifies would let a plugin refuse to free a volume that is no longer its own, rather than each backend arranging its own naming. Ciro Iriarte (5): storage: add asynchronous copy-offload hook for full copies rbd: implement copy-offload via snapshot + clone + flatten dir: implement copy-offload via reflink (FICLONE) btrfs, lvmthin: implement copy-offload (atomic class) qemu-server: use storage copy-offload for full clone pve-storage: ApiChangeLog | 24 +++ src/PVE/Storage.pm | 237 ++++++++++++++++++++++- src/PVE/Storage/BTRFSPlugin.pm | 250 ++++++++++++++++++++++++ src/PVE/Storage/DirPlugin.pm | 140 ++++++++++++++ src/PVE/Storage/LvmThinPlugin.pm | 189 ++++++++++++++++++ src/PVE/Storage/Plugin.pm | 176 +++++++++++++++++ src/PVE/Storage/RBDPlugin.pm | 318 +++++++++++++++++++++++++++++++ src/test/copy_offload_feature_test.pm | 98 ++++++++++ src/test/copy_offload_functional_test.pl | 286 +++++++++++++++++++++++++++ src/test/copy_offload_naming_test.pm | 82 ++++++++ src/test/copy_offload_test.pm | 90 +++++++++ src/test/run_plugin_tests.pl | 3 + 12 files changed, 1891 insertions(+), 2 deletions(-) qemu-server: src/PVE/API2/Qemu.pm | 17 ++++ src/PVE/QemuServer.pm | 201 ++++++++++++++++++++++++++++++++++++++++- src/PVE/QemuServer/BlockJob.pm | 32 ++++++- 3 files changed, 245 insertions(+), 5 deletions(-)