From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from gate001.proxmox.com (gate001.proxmox.com [45.144.208.40]) by lore.proxmox.com (Postfix) with ESMTPS id A6D571FF0B2 for ; Tue, 22 Sep 2026 10:14:14 +0200 (CEST) Received: from gate001.proxmox.com (localhost.localdomain [127.0.0.1]) by gate001.proxmox.com (Proxmox) with ESMTP id A89B8214F1; Tue, 22 Sep 2026 10:14:10 +0200 (CEST) Mime-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset=UTF-8 Date: Tue, 22 Sep 2026 10:14:03 +0200 Message-Id: To: =?utf-8?q?Michael_K=C3=B6ppl?= , Subject: Re: [PATCH qemu-server 1/1] fix #8030: snapshot: avoid orphaned vmstate volume on failure From: "Jakob Klocker" X-Mailer: aerc 0.20.0 References: <20260914140547.615346-1-m.koeppl@proxmox.com> In-Reply-To: <20260914140547.615346-1-m.koeppl@proxmox.com> X-Bm-Milter-Handled: 55990f41-d878-4baa-be0a-ee34c49e34d2 X-Bm-Transport-Timestamp: 1790064843861 X-SPAM-LEVEL: Spam detection results: 0 AWL 1.603 Adjusted score from AWL reputation of From: address DMARC_MISSING 0.1 Missing DMARC policy KAM_DMARC_STATUS 0.01 Test Rule for DKIM or SPF Failure with Strict Alignment (newer systems) RCVD_IN_DNSWL_MED -2.3 Sender listed at https://www.dnswl.org/, medium trust SPF_HELO_NONE 0.001 SPF: HELO does not publish an SPF Record SPF_PASS -0.001 SPF: sender matches SPF record Message-ID-Hash: I5KUXWUBHXLB7UQZJV6IG2M4VSMHJNQP X-Message-ID-Hash: I5KUXWUBHXLB7UQZJV6IG2M4VSMHJNQP X-MailFrom: j.klocker@proxmox.com X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; emergency; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header X-Mailman-Version: 3.3.10 Precedence: list List-Id: Proxmox VE development discussion List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: Confirmed that an orphaned volume is left behind when the QMP socket is occupied, causing the snapshot to fail because `query-machines` times out. After applying the patch, no orphaned volume remains, which fixes the reported bug. Not adding additional cleanup, but moving the disk allocation below the setup functions that can fail is IMO the cleanest solution here. Consider this: Reviewed-by: Jakob Klocker Tested-by: Jakob Klocker On Mon Sep 14, 2026 at 4:05 PM CEST, Michael K=C3=B6ppl wrote: > The state volume was allocated before the running instance was queried > for its machine type, CPU argument and network MTUs. Those queries can > fail (e.g. the 'query-machines' QMP command, which can run into a > timeout while QEMU's is busy [0]). > > The helper is called from __snapshot_prepare() before write_config(), so > the config never records the snapshot, snapshot_create() dies before it > gets to its cleanup path. Such a failure leaves behind the just > allocated volume. > > Query the running instance first, which makes the allocation the last > step of the function that can fail. > > [0] https://bugzilla.proxmox.com/show_bug.cgi?id=3D8030 > > Signed-off-by: Michael K=C3=B6ppl > [SNIP]