From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from gate001.proxmox.com (gate001.proxmox.com [45.144.208.40]) by lore.proxmox.com (Postfix) with ESMTPS id 397A01FF0E6 for ; Fri, 24 Jul 2026 11:57:22 +0200 (CEST) Received: from gate001.proxmox.com (localhost.localdomain [127.0.0.1]) by gate001.proxmox.com (Proxmox) with ESMTP id EB0BF21492; Fri, 24 Jul 2026 11:57:21 +0200 (CEST) Date: Fri, 24 Jul 2026 11:57:14 +0200 From: Fabian =?iso-8859-1?q?Gr=FCnbichler?= Subject: Re: [PATCH storage] fix #7598: qemu-img resize: tolerate timeout if resize succeeded To: Jakob Klocker , pve-devel@lists.proxmox.com References: <20260603082557.25359-1-j.klocker@proxmox.com> In-Reply-To: <20260603082557.25359-1-j.klocker@proxmox.com> MIME-Version: 1.0 User-Agent: astroid/0.17.0 (https://github.com/astroidmail/astroid) Message-Id: <1784886368.5u9kt3tdeb.astroid@yuna.none> Content-Type: text/plain; charset=utf-8 Content-Transfer-Encoding: quoted-printable X-Bm-Milter-Handled: 55990f41-d878-4baa-be0a-ee34c49e34d2 X-Bm-Transport-Timestamp: 1784887007265 X-SPAM-LEVEL: Spam detection results: 0 AWL -0.209 Adjusted score from AWL reputation of From: address DMARC_MISSING 0.1 Missing DMARC policy KAM_DMARC_STATUS 0.01 Test Rule for DKIM or SPF Failure with Strict Alignment (newer systems) RCVD_IN_DNSWL_LOW -0.7 Sender listed at https://www.dnswl.org/, low trust SPF_HELO_NONE 0.001 SPF: HELO does not publish an SPF Record SPF_PASS -0.001 SPF: sender matches SPF record Message-ID-Hash: 7SQGTVOHAU5RV253YXBAMFL5HWKQJK6N X-Message-ID-Hash: 7SQGTVOHAU5RV253YXBAMFL5HWKQJK6N X-MailFrom: f.gruenbichler@proxmox.com X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; emergency; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header X-Mailman-Version: 3.3.10 Precedence: list List-Id: Proxmox VE development discussion List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: On June 3, 2026 10:25 am, Jakob Klocker wrote: > On slow storages there is a chance the 10 second timeout is triggered > when resizing a volume. If the timeout fires while the resize is in a > certain state near the end, the operation can still complete > successfully even though a timeout error is thrown. In that case the > config is never updated and keeps the old, wrong size. >=20 > Because the config is out of sync, the volume is then displayed with > the wrong size in the web interface. >=20 > Link: https://bugzilla.proxmox.com/show_bug.cgi?id=3D7598 > Signed-off-by: Jakob Klocker > --- > src/PVE/Storage/Common.pm | 12 +++++++++++- > 1 file changed, 11 insertions(+), 1 deletion(-) >=20 > diff --git a/src/PVE/Storage/Common.pm b/src/PVE/Storage/Common.pm > index 3932aee..ee2ea00 100644 > --- a/src/PVE/Storage/Common.pm > +++ b/src/PVE/Storage/Common.pm > @@ -277,7 +277,17 @@ sub qemu_img_resize { > push $cmd->@*, '-f', $format, $path, $size; > =20 > $timeout =3D 10 if !$timeout; I think making this worker-task aware and bumping the timeout in that case would make more sense - storage performance varies wildy, and like I noted in the bug, this only seems to be called from a worker context where the additional time hurts way less than risking running into the timeout just because something is slow. e.g., in the RBD plugin we have a default connect timeout of 60s in workers, in the ZFS plugin we bump ZFS requests to 300s if a smaller timeout is set, and default to a timeout of one hour if none is set. > - run_command($cmd, timeout =3D> $timeout); > + eval { run_command($cmd, timeout =3D> $timeout); }; > + if (my $err =3D $@) { > + > + die $err if $err !~ /got timeout/; > + > + my $info =3D JSON::decode_json(qemu_img_info($path, $format)); > + die $err if !$info; > + > + my $actual_size =3D $info->{'virtual-size'}; > + die $err if !defined($actual_size) || $actual_size < $size; > + } IMHO this doesn't fix the actual bug mentioned.. I guess what you see here if running into this behaviour is that the resize was done, but syncing then takes long and hits the timeout? because what `qemu-img resize` does for raw images is basically just ftruncate(..) fdatasync(..) and just because you read back the updated size in the timeout case (from local, cached metadata), does not necessarily mean it got persisted to the storage (which might be on a different system) - that's what the sync is for after all.. > } > =20 > 1; > --=20 > 2.47.3 >=20 >=20 >=20 >=20 >=20