From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from gate001.proxmox.com (gate001.proxmox.com [IPv6:2a0f:8001:1:32::40]) by lore.proxmox.com (Postfix) with ESMTPS id 550C21FF09B for ; Mon, 14 Sep 2026 16:05:58 +0200 (CEST) Received: from gate001.proxmox.com (localhost.localdomain [127.0.0.1]) by gate001.proxmox.com (Proxmox) with ESMTP id 60466215B0; Mon, 14 Sep 2026 16:05:54 +0200 (CEST) From: =?UTF-8?q?Michael=20K=C3=B6ppl?= To: pve-devel@lists.proxmox.com Subject: [PATCH qemu-server 1/1] fix #8030: snapshot: avoid orphaned vmstate volume on failure Date: Mon, 14 Sep 2026 16:05:47 +0200 Message-ID: <20260914140547.615346-1-m.koeppl@proxmox.com> X-Mailer: git-send-email 2.47.3 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit X-Bm-Milter-Handled: 55990f41-d878-4baa-be0a-ee34c49e34d2 X-Bm-Transport-Timestamp: 1789394735773 X-SPAM-LEVEL: Spam detection results: 0 AWL 0.685 Adjusted score from AWL reputation of From: address DMARC_MISSING 0.1 Missing DMARC policy KAM_DMARC_STATUS 0.01 Test Rule for DKIM or SPF Failure with Strict Alignment (newer systems) RCVD_IN_DNSWL_MED -2.3 Sender listed at https://www.dnswl.org/, medium trust SPF_HELO_NONE 0.001 SPF: HELO does not publish an SPF Record SPF_PASS -0.001 SPF: sender matches SPF record Message-ID-Hash: FVIXPJ5S4GH5OEJPLQOGJKSDR3ZFIISA X-Message-ID-Hash: FVIXPJ5S4GH5OEJPLQOGJKSDR3ZFIISA X-MailFrom: m.koeppl@proxmox.com X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; emergency; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header X-Mailman-Version: 3.3.10 Precedence: list List-Id: Proxmox VE development discussion List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: The state volume was allocated before the running instance was queried for its machine type, CPU argument and network MTUs. Those queries can fail (e.g. the 'query-machines' QMP command, which can run into a timeout while QEMU's is busy [0]). The helper is called from __snapshot_prepare() before write_config(), so the config never records the snapshot, snapshot_create() dies before it gets to its cleanup path. Such a failure leaves behind the just allocated volume. Query the running instance first, which makes the allocation the last step of the function that can fail. [0] https://bugzilla.proxmox.com/show_bug.cgi?id=8030 Signed-off-by: Michael Köppl --- Tested this pretty much as described in the Bugzilla entry, but using a Perl script to connect to the QMP socket instead of socat. src/PVE/QemuConfig.pm | 7 +++++-- 1 file changed, 5 insertions(+), 2 deletions(-) diff --git a/src/PVE/QemuConfig.pm b/src/PVE/QemuConfig.pm index 26f0fda2..2ef6c075 100644 --- a/src/PVE/QemuConfig.pm +++ b/src/PVE/QemuConfig.pm @@ -238,8 +238,6 @@ sub __snapshot_save_vmstate { my $name = "vm-$vmid-state-$snapname"; $name .= ".raw" if $scfg->{path}; # add filename extension for file base storage - my $statefile = - PVE::Storage::vdisk_alloc($storecfg, $target, $vmid, 'raw', $name, $size * 1024); my $runningmachine = PVE::QemuServer::Machine::get_current_qemu_machine($vmid); # get current QEMU -cpu argument to ensure consistency of custom CPU models @@ -249,6 +247,11 @@ sub __snapshot_save_vmstate { my $nets_host_mtu = PVE::QemuServer::Network::get_nets_host_mtu($vmid, $conf); + # allocate only after querying the running instance, nothing below can fail, so a failed + # query cannot leave an orphaned state volume behind + my $statefile = + PVE::Storage::vdisk_alloc($storecfg, $target, $vmid, 'raw', $name, $size * 1024); + if (!$suspend) { $conf = $conf->{snapshots}->{$snapname}; } -- 2.47.3