From: Dominik Csapak <d.csapak@proxmox.com>
To: Roland Kammerer <roland.kammerer@linbit.com>,
pve-devel@lists.proxmox.com
Subject: Re: [RFC storage/qemu-server] add 'live-migration' hint to volume (de)activation
Date: Thu, 8 Oct 2026 12:48:18 +0200 [thread overview]
Message-ID: <4908cd76-2678-45aa-b101-8bf095c47e16@proxmox.com> (raw)
In-Reply-To: <asYrFw1whXIOWF0-@arm64>
Hi,
thanks for the detailed write-up. I'm not convinced we need a new
storage API for this, mostly because I think the problem sits in the
plugin's access model rather than in a missing signal from PVE. Some
details below.
On 10/7/26 1:29 PM, Roland Kammerer wrote:
[snip]
> During a live migration, the
> target QEMU process opens the disks read-write while the source QEMU
> still has them open, so the resource temporarily needs to allow two
> Primaries.
That is true for the open, but there is no concurrent write access. QEMU
hands over ownership of the images during migration, so there is never
more than one writer during live migration. The only requirement on the
storage is that a flush completed on the source is visible to reads on
the target afterwards.
This also matches the contract of the 'shared' flag, which is "a single
storage with the same contents on all nodes". PVE does not expect a
shared storage to arbitrate writers. Ownership is handled above the
storage layer by the node that owns the guest config, by HA fencing,
and by QEMU during migration. All other storage types (AFAIK) allow
concurrent opens/rw and never learn that a migration is happening.
> Specifically for DRBD we only allow two Primaries when using DRBD
> protocol C (the one with the strongest guarantees). Sometime people
> would like to use weaker guarantees (i.e., protocol A and B), which we
> can only allow when we sacrifice live migration as protocols A and B
> don't allow two Primaries at all. If we would know the live migration
> window, we could temporarily upgrade the connections between these 2
> nodes to protocol C and (also temporarily) allow-two-primaries.
As far as I understand protocols A and B, a write is acknowledged before
it has reached the peer's disk. In that case the source's final flush
can complete before the target's replica has the data, and the target
reads stale data right after activation. So for live migration,
protocol C is a correctness requirement.
Does switching from A/B to C handle this correctly?
> Another problem with setting allow-two-primaries permanently is that it
> allows admins, scripts,... to open the in-use device on a peer node and
> unwillingly altering data by accident.
I understand the wish for that protection, but none of our other shared
storages offer it, and PVE does not rely on it. If the plugin wants to
be stricter than the contract, it can already do so on its own, for
example by allowing the second Primary only for the duration of an
activation on another node.
Also a probably better interface would be a storage plugin api
like 'add/end_shared_access' (or similar) that is called before and
after live migration (per storage; with a list of volumes).
We had some discussion of this internally and we're not convinced that
adding this kind of API for a single storage type is justified.
Can you name any other storage that might profit from this?
(I could only think of lvm + lockd maybe, but we don't use that in
favor of our own cluster wide locking)
Best Regards
Dominik
prev parent reply other threads:[~2026-10-08 10:48 UTC|newest]
Thread overview: 2+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-07 11:20 [RFC storage/qemu-server] add 'live-migration' hint to volume (de)activation Roland Kammerer
2026-10-08 10:48 ` Dominik Csapak [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=4908cd76-2678-45aa-b101-8bf095c47e16@proxmox.com \
--to=d.csapak@proxmox.com \
--cc=pve-devel@lists.proxmox.com \
--cc=roland.kammerer@linbit.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox