From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from gate001.proxmox.com (gate001.proxmox.com [IPv6:2a0f:8001:1:32::40]) by lore.proxmox.com (Postfix) with ESMTPS id F3AE81FF145 for ; Sun, 19 Jul 2026 05:35:51 +0200 (CEST) Received: from gate001.proxmox.com (localhost.localdomain [127.0.0.1]) by gate001.proxmox.com (Proxmox) with ESMTP id D61002147C; Sun, 19 Jul 2026 05:35:50 +0200 (CEST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784432139; x=1785036939; darn=lists.proxmox.com; h=references:in-reply-to:subject:to:from:date:message-id:from:to:cc :subject:date:message-id:reply-to:content-type; bh=YWqWQ3Jo6AzvPUnCQmSFjrKhmTC197UtGEYZGeOtb4U=; b=ANq8wpULqcj3tN8yOZjrHHFpiWRaFr+p5J7+u7sZhnBDMugHZHD0IH10u7jiUyu1E1 vFDhrZRaVEvhN2WaZdcWjYhCqCSUoygw4pE4ZO7+zyR/wGFjRe7RW2uf/RqEDXI848yf w6Yf4mKZCNTP0yUctgGxfpDNXETQco313PYW78oFk978YSMB79kBonKIGkyirSvxheQc cqMdT0YQn3GMIjGVYaU9NMnQ0Xu7IFvdTY7fCZhaDnQwSAGtlHjTk87nR57p+KUVK3gT 1+gNE+ZOSgOPCpDVhNJWvzCO8KKn5Mm6oftKICZYt4HEpe2L6kMNdaWFOvE3ERhu/fqd vdsA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784432139; x=1785036939; h=references:in-reply-to:subject:to:from:date:message-id:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=YWqWQ3Jo6AzvPUnCQmSFjrKhmTC197UtGEYZGeOtb4U=; b=ngbBKRjsgfzq0KKSJ16hLHDOxpwUDHKt8QSUN7T3gF1iSS+9fJaThW7QblhbxR9XsD izXAKZNZ+b8E2wUT/HdGHjfNY8q80PCI1IIgqxSq2QnJpBz85IbT3bFeEgBc2Tkk7Xdf vLUwap0yKoaf7r6h+zqqvt2XAolbnMSjg+L6FXQw3uSRyXWmWYY0ymd7wlrDqq2EdrQL fAg4J88qak79qpE2okJSjDBRuwHSx9WiQP4NyexCjg6CzI3XrTqT8vWRrEVwuO2dCHR0 hCwBS38pspfFEoXrfq3DAz3WsFfo3viyj8oN4nNGF78m5xNZu+sjfcf6jWlzTNX9r6qa EsrQ== X-Forwarded-Encrypted: i=1; AHgh+RolMg8jdoyYTdqRxKZXSGmopIdmYicON4LrtXarU0OGEQhHWn+tfaLNTTZRJDjh4ltxWfketT0+grk=@lists.proxmox.com X-Gm-Message-State: AOJu0Yz3wsxlyjxShf7uIx1qEZBolveMobI+yPid1n7aosdeWUA3iut4 dM3L+1bXqusLt0iM+pwQtw2Rzuo+9AUCp/0d3metB31tThzaMgOtF6RT X-Gm-Gg: AfdE7cn8sq8m5iwCuT4ylJ/9w/qA/qTIRDBs0WIlSAUzjJkDCXZiegkWQf7DbiMG6qE zLCFgkYz23TFpg1mYn+lxz0OZ3SsTEXH2lPZPfRxhXCMlgem3menniMjYD9pPKVsc5j12s1yLkG 6CaeJ/hdnfSMoIYSHTh9smMFWyxkViQrou4uBluKTI/0iGOBbVTWiqFqnkeQSDZeIKBrErSvCbR GBrSQu2VMky5ryqOYcSBlcNWZOZrh58O6KQpqIUBLjRCSf1QFMgpkevMudd80D/JM5dNlSES4cG Y70CH8AgW15TxwvUUNFRQIBUiAswLTu6lKf8KvoXgljgVtQFXLwGuGUYpVywaIESnGTInJFZxpo ZfT/v5Nkq4D8RlHqmB1JcmnsDZn3VmhqvY65eSVe5zCXXgh8/IFwsixyA+scxreN02jdARfRp X-Received: by 2002:a17:90b:54c4:b0:38e:3a8:2374 with SMTP id 98e67ed59e1d1-38e4b5408c0mr8687503a91.30.1784432138954; Sat, 18 Jul 2026 20:35:38 -0700 (PDT) Message-ID: <6a5c460a.66c0679a.24976c.4074@mx.google.com> Date: Sat, 18 Jul 2026 20:35:38 -0700 (PDT) From: Ciro Iriarte To: Fiona Ebner , pve-devel@lists.proxmox.com Subject: Re: [RFC storage, qemu-server] offload full/live clone to storage backend In-Reply-To: References: <6a4416f5.01f0a1da.1533d8.2fdf@mx.google.com> <1783430000.wo44ge3ggk.astroid@yuna.none> <90aae63f-8d67-4a26-ab6b-ed32fd1537ce@proxmox.com> <1783500024.h4hi20lbpy.astroid@yuna.none> <6a50ff20.d6ce4a42.63dff.76fa@mx.google.com> X-SPAM-LEVEL: Spam detection results: 0 AWL 0.133 Adjusted score from AWL reputation of From: address DKIM_SIGNED 0.1 Message has a DKIM or DK signature, not necessarily valid DKIM_VALID -0.1 Message has at least one valid DKIM or DK signature DKIM_VALID_AU -0.1 Message has a valid DKIM or DK signature from author's domain DKIM_VALID_EF -0.1 Message has a valid DKIM or DK signature from envelope-from domain DMARC_PASS -0.1 DMARC pass policy FREEMAIL_FROM 0.001 Sender email is commonly abused enduser mail provider RCVD_IN_DNSWL_NONE -0.0001 Sender listed at https://www.dnswl.org/, no trust SPF_HELO_NONE 0.001 SPF: HELO does not publish an SPF Record SPF_PASS -0.001 SPF: sender matches SPF record Message-ID-Hash: MELNN5NM7RJQWCXXN3FRL5NP6OWZ3ZK6 X-Message-ID-Hash: MELNN5NM7RJQWCXXN3FRL5NP6OWZ3ZK6 X-MailFrom: cyruspy@gmail.com X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; emergency; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header X-Mailman-Version: 3.3.10 Precedence: list List-Id: Proxmox VE development discussion List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: On July 13, 2026 4:03 pm, Fiona Ebner wrote: > I guess it's easiest to just have two different features to avoid the > need to change the interface for volume_has_feature(). Done -- `copy-offload-atomic` and `copy-offload-bulk`, volume_has_feature() untouched. Sorry for the delay; the design changed twice while I was building it. Two things I got wrong, both of which changed the shape. First, my argument that a clone needs no host-side mirror only holds for a single disk. For several disks the freeze in monitor() is what puts them at one instant, and an offload started per-disk in clone_disk() misses that rendezvous entirely -- data and journal on separate disks would clone torn. Fabian's point about unifying with the mirror machinery stands for the running case. So I use that rendezvous instead of avoiding it: for a running source clone_disk() only allocates the target and defers the copy's start, and monitor() takes an optional on_frozen callback that runs the deferred starts inside the same freeze that cuts the mirrors over. Offloaded and mirrored disks then share an instant, and mixed VMs need no second freeze. Only the data movement happens after the thaw. Second, a synchronous hook cannot work on shared storage: the plugin is called under cluster_lock_storage(), and PVE::Cluster wraps that in alarm(60), while a full copy runs for minutes or hours. I had missed this. So the hook is now three phases: copy_image_prepare() allocate the target - under the storage lock copy_image_start() pin the source PIT, begin - inside the guest freeze copy_image_status() poll until independent - outside both The freeze is the second reason for the split: allocation is far too expensive to hold a guest frozen across, so only the start runs there. On a VSP the point in time is fixed by one group action, so the freeze holds a single request regardless of disk count. 'complete' means independent of the source, not merely readable -- the distinction that matters for rbd flatten, a zfs clone still tied to its origin, or an array pair that has not dissolved. The series includes an RBD implementation, so the hook can be exercised without vendor hardware: snapshot, clone, flatten. On a Ceph Tentacle cluster the copy finished in 3s, the target had no parent, md5 matched the source, and it still read correctly after I deleted the source. The flatten goes to `ceph rbd task add flatten` rather than a forked child -- cluster-side, survives the worker dying, cancellable. It also makes failure visible: status can tell complete (no parent) from pending (a manager task exists) from failed (still a clone, no task). Without that last state a dead flatten looks exactly like a slow one and the caller polls forever. One bug worth flagging since it was mine and it was silent: prepare received $snap but start did not, and start is what fixes the point in time. RBD therefore snapshotted the current image while advertising the feature for the 'snap' key, so a clone from a snapshot would have copied live data. The parameter had gone to the phase that allocates rather than the one that captures. Fixed, and each entry of a batch carries its own $snap now. A related one: prepare reserved only a name, so a concurrent allocation could take it and rollback would free the other operation's volume -- prepare creates the target now. Open questions: - move-disk. Same mechanics, but it needs converge-to-current, which is the bulk + bitmap/mirror path I have not implemented. Tell me if you would rather they land together. - copy_images_start() starts a batch between one pair of storages, with a default that just loops, so there is nothing to negotiate. It exists because an array can capture several volumes at one instant. Nothing in-tree overrides it: Ceph group snapshots do work, but an image can only be in one group, so grouping would mutate the source's group membership and collide with an admin's own. If that is too speculative without an in-tree user, it is the last commit and drops cleanly. - prepare returns a volname. A token would let a plugin carry reservation state without inventing a name, at the cost of messier rollback. Three commits: the hook (with ApiChangeLog and tests for the eligibility rules), the RBD implementation, the qemu-server consumer. It compiles against a real tree and the storage half is validated end-to-end on Ceph; the freeze bracket has only been exercised offline so far. bulk with a running guest still falls through to the host path. Happy to post the patches now, or after you have looked at the shape. Thanks, Ciro