From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from gate001.proxmox.com (gate001.proxmox.com [IPv6:2a0f:8001:1:32::40]) by lore.proxmox.com (Postfix) with ESMTPS id 9C9131FF09F for ; Thu, 03 Sep 2026 15:15:30 +0200 (CEST) Received: from gate001.proxmox.com (localhost.localdomain [127.0.0.1]) by gate001.proxmox.com (Proxmox) with ESMTP id 6317F215C8; Thu, 03 Sep 2026 15:15:29 +0200 (CEST) Message-ID: Date: Thu, 3 Sep 2026 15:15:22 +0200 MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: SPAM: qemu-img convert runs with cache=unsafe for every block storage except zfspool To: "Ing. Alfonso Kuen Arroyo" , pve-devel@lists.proxmox.com References: <178795574681.28052.11392785472014160013@idkmanager.com> Content-Language: en-US From: Fiona Ebner In-Reply-To: <178795574681.28052.11392785472014160013@idkmanager.com> Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 7bit X-Bm-Milter-Handled: 55990f41-d878-4baa-be0a-ee34c49e34d2 X-Bm-Transport-Timestamp: 1788441319989 X-SPAM-LEVEL: Spam detection results: 0 AWL 0.727 Adjusted score from AWL reputation of From: address DMARC_MISSING 0.1 Missing DMARC policy KAM_DMARC_STATUS 0.01 Test Rule for DKIM or SPF Failure with Strict Alignment (newer systems) RCVD_IN_DNSWL_MED -2.3 Sender listed at https://www.dnswl.org/, medium trust SPF_HELO_NONE 0.001 SPF: HELO does not publish an SPF Record SPF_PASS -0.001 SPF: sender matches SPF record Message-ID-Hash: WUZ4ZRYKKR23XJP7AO45ORF5RMWSVZXZ X-Message-ID-Hash: WUZ4ZRYKKR23XJP7AO45ORF5RMWSVZXZ X-MailFrom: f.ebner@proxmox.com X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; emergency; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header X-Mailman-Version: 3.3.10 Precedence: list List-Id: Proxmox VE development discussion List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: Hi Alfonso, thank you for the report! Am 03.09.26 um 1:13 PM schrieb Ing. Alfonso Kuen Arroyo: > Hello, > > Reporting a data-loss path in disk move / migration that we hit in production on > PVE 9.2.4 (qemu-server 9.1.18). Happy to move this to Bugzilla if that is the > preferred channel -- I do not have an account there yet. Yes, usually the Bugzilla is the preferred place for filing issues. > > THE ISSUE > > PVE/QemuServer/QemuImage.pm protects the cache mode only when the storage is a > ZFS pool: > > 122: $cachemode = 'none' if $src_scfg->{type} eq 'zfspool'; # source (-T) > 149: push @$cmd, '-t', 'none' if $dst_scfg->{type} eq 'zfspool'; # target (-t) > > For every other destination, qemu-img convert falls back to its own default for > -t, which is 'unsafe', i.e. BDRV_O_NO_FLUSH. If a write fails while the data is > still a dirty page, the page is dropped, the errseq is never consumed by an > fsync that would report it, and qemu-img exits 0. PVE then reports TASK OK for a > copy that silently lost data. > > This affects any block destination: iSCSI, FC, LVM over SAN, NVMe-oF, and > third-party storage plugins. It is not specific to any one of them. This is an unfortunate default. > WHY THE EXISTING SAFETY ARGUMENT DOES NOT HOLD > > The pve-devel thread that introduced this (Aug 2016) asked exactly the right > question -- "is this really safe?" -- and the answer was that bdrv_close() sends > a flush. But QEMU's bdrv_co_flush() contains: > > /* But don't actually force it to the disk with cache=unsafe */ > if (bs->open_flags & BDRV_O_NO_FLUSH) { > goto flush_children; > } > > The flush is issued, and cache=unsafe is precisely the mode in which it does > nothing. The safety argument invalidates itself, and it has been in production > for ten years. > > WHAT IT COST US > > Three PVE 9 nodes against a TrueNAS SCALE 25.10.4 NVMe/TCP target. The target > was failing large commands under memory pressure -- a separate bug, ours to fix, > written up at: > > https://github.com/truenas/truenas-proxmox-plugin/issues/96 > > Those failures should have surfaced as a failed migration. Instead: > > - A 2 TiB image copy lost roughly 98 GiB, and qemu-img exited 0. > - The guest then ran 41 hours on that copy, logging 2346 XFS metadata > corruption events, before we noticed. > - All 520 "lost async page write" events were on the destination. > > With -t none the copy would have failed loudly and no data would have been lost. > The storage bug was ours; the silence was not. > > SUGGESTED FIX > > Pass -t none for any destination that is not a file-backed storage where the > performance tradeoff was deliberately accepted. More conservatively: invert the > condition so 'unsafe' is opt-in per storage type, rather than the default for > everything except one. Yes, we might want to use none or writeback depending on the storage type, similar to PBS restore. > > REPRODUCING WITHOUT A FAULTY ARRAY > > dmsetup a 'flakey' or 'error' target under a test LV, qemu-img convert onto it, > and check the exit code. The point of the report is that the exit code is 0. I'm not able to reproduce this directly, could you share more details? [I] root@pve9a1 ~# dmsetup create test-error --table ' 0 8388607 linear /dev/sdu 0 8388607 8388608 error ' 8388608 [I] root@pve9a1 ~# dmsetup create test-flakey --table ' 0 8388608 flakey /dev/sdu 0 1 1 ' 8388608 [I] root@pve9a1 ~# /usr/bin/qemu-img convert -p -n -f qcow2 -O raw /mnt/pve/nfs/images/100/vm-100-disk-1.qcow2 /dev/mapper/test-error qemu-img: error while writing at byte 4292870144: Input/output error [I] root@pve9a1 ~ [1]# /usr/bin/qemu-img convert -p -n -f qcow2 -O raw /mnt/pve/nfs/images/100/vm-100-disk-1.qcow2 /dev/mapper/test-flakey qemu-img: error while writing at byte 1694498816: Input/output error Best Regards, Fiona