all lists on lists.proxmox.com
 help / color / mirror / Atom feed
* qemu-img convert runs with cache=unsafe for every block storage except zfspool
@ 2026-08-28 22:22 Ing. Alfonso Kuen Arroyo
  2026-09-03 13:15 ` SPAM: " Fiona Ebner
  0 siblings, 1 reply; 2+ messages in thread
From: Ing. Alfonso Kuen Arroyo @ 2026-08-28 22:22 UTC (permalink / raw)
  To: pve-devel

Hello,

Reporting a data-loss path in disk move / migration that we hit in production on
PVE 9.2.4 (qemu-server 9.1.18). Happy to move this to Bugzilla if that is the
preferred channel -- I do not have an account there yet.

THE ISSUE

PVE/QemuServer/QemuImage.pm protects the cache mode only when the storage is a
ZFS pool:

  122:  $cachemode = 'none' if $src_scfg->{type} eq 'zfspool';       # source (-T)
  149:  push @$cmd, '-t', 'none' if $dst_scfg->{type} eq 'zfspool';  # target (-t)

For every other destination, qemu-img convert falls back to its own default for
-t, which is 'unsafe', i.e. BDRV_O_NO_FLUSH. If a write fails while the data is
still a dirty page, the page is dropped, the errseq is never consumed by an
fsync that would report it, and qemu-img exits 0. PVE then reports TASK OK for a
copy that silently lost data.

This affects any block destination: iSCSI, FC, LVM over SAN, NVMe-oF, and
third-party storage plugins. It is not specific to any one of them.

WHY THE EXISTING SAFETY ARGUMENT DOES NOT HOLD

The pve-devel thread that introduced this (Aug 2016) asked exactly the right
question -- "is this really safe?" -- and the answer was that bdrv_close() sends
a flush. But QEMU's bdrv_co_flush() contains:

  /* But don't actually force it to the disk with cache=unsafe */
  if (bs->open_flags & BDRV_O_NO_FLUSH) {
          goto flush_children;
  }

The flush is issued, and cache=unsafe is precisely the mode in which it does
nothing. The safety argument invalidates itself, and it has been in production
for ten years.

WHAT IT COST US

Three PVE 9 nodes against a TrueNAS SCALE 25.10.4 NVMe/TCP target. The target
was failing large commands under memory pressure -- a separate bug, ours to fix,
written up at:

  https://github.com/truenas/truenas-proxmox-plugin/issues/96

Those failures should have surfaced as a failed migration. Instead:

  - A 2 TiB image copy lost roughly 98 GiB, and qemu-img exited 0.
  - The guest then ran 41 hours on that copy, logging 2346 XFS metadata
    corruption events, before we noticed.
  - All 520 "lost async page write" events were on the destination.

With -t none the copy would have failed loudly and no data would have been lost.
The storage bug was ours; the silence was not.

SUGGESTED FIX

Pass -t none for any destination that is not a file-backed storage where the
performance tradeoff was deliberately accepted. More conservatively: invert the
condition so 'unsafe' is opt-in per storage type, rather than the default for
everything except one.

REPRODUCING WITHOUT A FAULTY ARRAY

dmsetup a 'flakey' or 'error' target under a test LV, qemu-img convert onto it,
and check the exit code. The point of the report is that the exit code is 0.

Regards,
Alfonso Kuen



^ permalink raw reply	[flat|nested] 2+ messages in thread

* Re: SPAM: qemu-img convert runs with cache=unsafe for every block storage except zfspool
  2026-08-28 22:22 qemu-img convert runs with cache=unsafe for every block storage except zfspool Ing. Alfonso Kuen Arroyo
@ 2026-09-03 13:15 ` Fiona Ebner
  0 siblings, 0 replies; 2+ messages in thread
From: Fiona Ebner @ 2026-09-03 13:15 UTC (permalink / raw)
  To: Ing. Alfonso Kuen Arroyo, pve-devel

Hi Alfonso,

thank you for the report!

Am 03.09.26 um 1:13 PM schrieb Ing. Alfonso Kuen Arroyo:
> Hello,
> 
> Reporting a data-loss path in disk move / migration that we hit in production on
> PVE 9.2.4 (qemu-server 9.1.18). Happy to move this to Bugzilla if that is the
> preferred channel -- I do not have an account there yet.

Yes, usually the Bugzilla is the preferred place for filing issues.

> 
> THE ISSUE
> 
> PVE/QemuServer/QemuImage.pm protects the cache mode only when the storage is a
> ZFS pool:
> 
>   122:  $cachemode = 'none' if $src_scfg->{type} eq 'zfspool';       # source (-T)
>   149:  push @$cmd, '-t', 'none' if $dst_scfg->{type} eq 'zfspool';  # target (-t)
> 
> For every other destination, qemu-img convert falls back to its own default for
> -t, which is 'unsafe', i.e. BDRV_O_NO_FLUSH. If a write fails while the data is
> still a dirty page, the page is dropped, the errseq is never consumed by an
> fsync that would report it, and qemu-img exits 0. PVE then reports TASK OK for a
> copy that silently lost data.
> 
> This affects any block destination: iSCSI, FC, LVM over SAN, NVMe-oF, and
> third-party storage plugins. It is not specific to any one of them.

This is an unfortunate default.

> WHY THE EXISTING SAFETY ARGUMENT DOES NOT HOLD
> 
> The pve-devel thread that introduced this (Aug 2016) asked exactly the right
> question -- "is this really safe?" -- and the answer was that bdrv_close() sends
> a flush. But QEMU's bdrv_co_flush() contains:
> 
>   /* But don't actually force it to the disk with cache=unsafe */
>   if (bs->open_flags & BDRV_O_NO_FLUSH) {
>           goto flush_children;
>   }
> 
> The flush is issued, and cache=unsafe is precisely the mode in which it does
> nothing. The safety argument invalidates itself, and it has been in production
> for ten years.
> 
> WHAT IT COST US
> 
> Three PVE 9 nodes against a TrueNAS SCALE 25.10.4 NVMe/TCP target. The target
> was failing large commands under memory pressure -- a separate bug, ours to fix,
> written up at:
> 
>   https://github.com/truenas/truenas-proxmox-plugin/issues/96
> 
> Those failures should have surfaced as a failed migration. Instead:
> 
>   - A 2 TiB image copy lost roughly 98 GiB, and qemu-img exited 0.
>   - The guest then ran 41 hours on that copy, logging 2346 XFS metadata
>     corruption events, before we noticed.
>   - All 520 "lost async page write" events were on the destination.
> 
> With -t none the copy would have failed loudly and no data would have been lost.
> The storage bug was ours; the silence was not.
> 
> SUGGESTED FIX
> 
> Pass -t none for any destination that is not a file-backed storage where the
> performance tradeoff was deliberately accepted. More conservatively: invert the
> condition so 'unsafe' is opt-in per storage type, rather than the default for
> everything except one.

Yes, we might want to use none or writeback depending on the storage
type, similar to PBS restore.

> 
> REPRODUCING WITHOUT A FAULTY ARRAY
> 
> dmsetup a 'flakey' or 'error' target under a test LV, qemu-img convert onto it,
> and check the exit code. The point of the report is that the exit code is 0.

I'm not able to reproduce this directly, could you share more details?

[I] root@pve9a1 ~# dmsetup create test-error --table '
                   0 8388607 linear /dev/sdu 0
                   8388607 8388608 error
                   ' 8388608
[I] root@pve9a1 ~# dmsetup create test-flakey --table '
                   0 8388608 flakey /dev/sdu 0 1 1
                   ' 8388608
[I] root@pve9a1 ~# /usr/bin/qemu-img convert -p -n -f qcow2 -O raw
/mnt/pve/nfs/images/100/vm-100-disk-1.qcow2 /dev/mapper/test-error
qemu-img: error while writing at byte 4292870144: Input/output error

[I] root@pve9a1 ~ [1]# /usr/bin/qemu-img convert -p -n -f qcow2 -O raw
/mnt/pve/nfs/images/100/vm-100-disk-1.qcow2 /dev/mapper/test-flakey
qemu-img: error while writing at byte 1694498816: Input/output error

Best Regards,
Fiona




^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-09-03 13:15 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-28 22:22 qemu-img convert runs with cache=unsafe for every block storage except zfspool Ing. Alfonso Kuen Arroyo
2026-09-03 13:15 ` SPAM: " Fiona Ebner

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.
Service provided by Proxmox Server Solutions GmbH | Privacy | Legal