From: Roland Kammerer <roland.kammerer@linbit.com>
To: pve-devel@lists.proxmox.com
Subject: Re: [RFC storage/qemu-server] add 'live-migration' hint to volume (de)activation
Date: Thu, 8 Oct 2026 15:39:14 +0200 [thread overview]
Message-ID: <asedAu3C-KiPmo-9@arm64> (raw)
In-Reply-To: <4908cd76-2678-45aa-b101-8bf095c47e16@proxmox.com>
On Thu, Oct 08, 2026 at 12:48:18PM +0200, Dominik Csapak wrote:
> Hi,
>
> thanks for the detailed write-up. I'm not convinced we need a new
> storage API for this, mostly because I think the problem sits in the
> plugin's access model rather than in a missing signal from PVE. Some
> details below.
>
>
> On 10/7/26 1:29 PM, Roland Kammerer wrote:
> [snip]
> > During a live migration, the
> > target QEMU process opens the disks read-write while the source QEMU
> > still has them open, so the resource temporarily needs to allow two
> > Primaries.
>
> That is true for the open, but there is no concurrent write access. QEMU
> hands over ownership of the images during migration, so there is never
> more than one writer during live migration. The only requirement on the
> storage is that a flush completed on the source is visible to reads on
> the target afterwards.
Sure, I know that, otherwise we would have a serious problem and we need
to trust qemu/pve there anyways (with our without the live-migration
hints).
> This also matches the contract of the 'shared' flag, which is "a single
> storage with the same contents on all nodes". PVE does not expect a
> shared storage to arbitrate writers. Ownership is handled above the
> storage layer by the node that owns the guest config, by HA fencing,
> and by QEMU during migration. All other storage types (AFAIK) allow
> concurrent opens/rw and never learn that a migration is happening.
True, that is how they work, they are shared storage. DRBD on the other
hand is a bit special here. It can act "shared"/dual primary, but in
general we would like to keep that window as small as possible. Actually
in the best case, and what DRBD9 describes as the only supported case,
we only want to allow that during live migration. But for that we would
need to know when one happens :).
> > Specifically for DRBD we only allow two Primaries when using DRBD
> > protocol C (the one with the strongest guarantees). Sometime people
> > would like to use weaker guarantees (i.e., protocol A and B), which we
> > can only allow when we sacrifice live migration as protocols A and B
> > don't allow two Primaries at all. If we would know the live migration
> > window, we could temporarily upgrade the connections between these 2
> > nodes to protocol C and (also temporarily) allow-two-primaries.
>
> As far as I understand protocols A and B, a write is acknowledged before
> it has reached the peer's disk. In that case the source's final flush
> can complete before the target's replica has the data, and the target
> reads stale data right after activation. So for live migration,
> protocol C is a correctness requirement.
> Does switching from A/B to C handle this correctly?
>
I assume so, otherwise that would be a problem the DRBD devs have to
solve.
> > Another problem with setting allow-two-primaries permanently is that it
> > allows admins, scripts,... to open the in-use device on a peer node and
> > unwillingly altering data by accident.
>
> I understand the wish for that protection, but none of our other shared
> storages offer it, and PVE does not rely on it.
mhm, that is what I tried to answer above, they are how they are and
can't do better, DRBD could.
> If the plugin wants to be stricter than the contract, it can already
> do so on its own, for example by allowing the second Primary only for
> the duration of an activation on another node.
I tried that, but it was always guess work, it never completely worked
out/felt right. AFAIR activate/deactivate don't have to be strictly
symmetrical, activates can happen for different reasons,... Probably
with more heuristics and more guessing one can come up with something,
but at some point it looked easier to just pass down the information I'm
really interested in - live migrations start/end - to the plugin
and be done with it. Then it would have become trivial in the plugin
without guessing why an activate happens.
> Also a probably better interface would be a storage plugin api
> like 'add/end_shared_access' (or similar) that is called before and
> after live migration (per storage; with a list of volumes).
Yes, I saw that alternative too, the proposal was long enough already,
but certainly this would have been fine as well. "Something that passes
start/end of live migration".
> We had some discussion of this internally and we're not convinced that
> adding this kind of API for a single storage type is justified.
> Can you name any other storage that might profit from this?
> (I could only think of lvm + lockd maybe, but we don't use that in
> favor of our own cluster wide locking)
Fair enough and thanks for the review and comments. Today I learned that
there is a DRBD feature in the pipeline where a write on the second
Primary if in dual primary forces the first one to become Secondary, all
within DRBD. That then will fix the most important part of our scenario:
Not letting people shoot themselves in the foot if one could do better
(e.g., by minimizing the "dangerous window" via the proposed
live-migration hints).
Thanks, rck
prev parent reply other threads:[~2026-10-08 13:39 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-07 11:20 [RFC storage/qemu-server] add 'live-migration' hint to volume (de)activation Roland Kammerer
2026-10-08 10:48 ` Dominik Csapak
2026-10-08 13:39 ` Roland Kammerer [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=asedAu3C-KiPmo-9@arm64 \
--to=roland.kammerer@linbit.com \
--cc=pve-devel@lists.proxmox.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.