* [RFC storage/qemu-server] add 'live-migration' hint to volume (de)activation
@ 2026-10-07 11:20 Roland Kammerer
2026-10-08 10:48 ` Dominik Csapak
0 siblings, 1 reply; 2+ messages in thread
From: Roland Kammerer @ 2026-10-07 11:20 UTC (permalink / raw)
To: pve-devel
Hi,
Disclaimer: parts of this RFC have been written with the help of an LLM.
This results in a relatively verbose RFC, but I kept the details in the
hope they can be useful. At this time I'm mainly interested if the idea
for hints when a live migration starts/ends is a good idea in general or
not. I marked LLM assisted/generated sections where I think it adds some
benefit for the reviewers.
With the hints mechanism introduced in storage plugin APIVER 13 [0]
there is now a clean, well-defined channel to pass contextual
information from the upper layers down to storage plugins on volume
activation. I would like to propose a new hint that tells a plugin that
a volume (de)activation happens in the context of a live migration,
including the role of the local node.
Motivation
==========
Some storage technologies need to know about the concurrent-access
window that a live migration opens. The concrete case is DRBD via the
external linstor-proxmox plugin: DRBD normally allows exactly one
Primary (writable) node per resource. During a live migration, the
target QEMU process opens the disks read-write while the source QEMU
still has them open, so the resource temporarily needs to allow two
Primaries. Today there is no signal for this, so the plugin has to
either keep allow-two-primaries enabled permanently (current situation,
undesirable, as it weakens split-brain protection during normal
operation) or guess from resource state (racy). Other plugins with
single-writer semantics face the same problem.
Another problem with setting allow-two-primaries permanently is that it
allows admins, scripts,... to open the in-use device on a peer node and
unwillingly altering data by accident.
Specifically for DRBD we only allow two Primaries when using DRBD
protocol C (the one with the strongest guarantees). Sometime people
would like to use weaker guarantees (i.e., protocol A and B), which we
can only allow when we sacrifice live migration as protocols A and B
don't allow two Primaries at all. If we would know the live migration
window, we could temporarily upgrade the connections between these 2
nodes to protocol C and (also temporarily) allow-two-primaries.
LLM generated, unverified: Current behavior (as of qemu-server master)
======================================================================
- Migration start: the target node activates all volumes in
vm_start_nolock() before command line generation. This call site
already passes hints (generate_storage_hints()), and the incoming
migration context is in scope there.
- Migration end (success): the source node deactivates the volumes via
phase3_cleanup() -> vm_stop() -> vm_stop_cleanup() ->
deactivate_volumes(). No hints support in that path today.
- Migration end (abort after target start): the source VM keeps running
and never deactivates; instead the *target* tears down its VM via
vm_stop() (which already carries $migratedfrom on this path) ->
vm_stop_cleanup() -> deactivate_volumes().
So the concurrent-access window is opened by exactly one call (hinted
activate on the target) and closed by exactly one call (deactivate on
the source on success, or on the target on abort). The proposal is to
mark all three.
LLM assisted, reviewed to my best knowledge: Proposal
=====================================================
1) pve-storage: add a new hint property
'live-migration' => {
type => 'string',
enum => ['incoming', 'outgoing'],
optional => 1,
description => "The volume belongs to a guest that is currently
being live-migrated. The value denotes the role of the local
node in the migration.",
},
Per the APIVER 13 rules, adding a hint itself needs no version bump.
2) pve-storage: add a $hints parameter to the deactivate_volume() plugin
method (appended after $cache) and to PVE::Storage::deactivate_volumes(),
analogous to activate_volume(). This is a backward-compatible signature
extension: APIVER 17 with APIAGE increased, existing plugins can simply
ignore the additional parameter.
3) qemu-server: pass the hint at the three sites above, guarded by
PVE::Storage::Plugin::is_hint_supported() as usual, plus a versioned
dependency bump on libpve-storage-perl:
- vm_start_nolock(): 'incoming' on activate_volumes() when starting
for an incoming live migration.
- source, after successful migration: 'outgoing' on the deactivation.
- target, when cleaning up a failed incoming migration: 'incoming' on
the deactivation.
For the two deactivation sites this means threading the migration
context through vm_stop()/vm_stop_cleanup() (which already know
$migratedfrom on the abort path), or alternatively explicit
deactivate_volumes() calls in the QemuMigrate cleanup phases. I am
happy to go either way, opinions welcome.
LLM assisted, reviewed to my best knowledge: Semantics for plugins
==================================================================
An activation hinted 'incoming' means a concurrent-activation window
opens: the volume is (or will shortly be) simultaneously active on the
migration source. A deactivation hinted 'incoming' or 'outgoing' means
that window is closed. A plugin that needs to relax single-writer
enforcement can key it off exactly these calls.
In line with the existing hints contract, this stays best-effort: hints
are not guaranteed on every call, and a node crash mid-migration means
the closing deactivation may never happen. Plugins must therefore remain
correct without the hint and self-heal (e.g. linstor-proxmox would record
that it relaxed the setting and clear it on the next unhinted
(de)activation of the volume). The hint makes the common path robust and
removes the need for permanently relaxed settings; it is not a
transactional API.
Consumer
========
linstor-proxmox would use the hint to temporarily allow two Primaries for
the specific resource during the migration window (LINSTOR already knows
which node is Primary and can adjust settings per connection), instead of
requiring the current permanently-enabled configuration.
plugin change alongside as a demonstration.
If the general direction is agreeable, I could try to follow up with patches.
If you think the description is good enough and the implementation is better
done by PVE devs, that certainly is fine as well.
Looking forward to your feedback.
Best regards,
Roland
[0] https://lore.proxmox.com/pve-devel/20251031103709.60233-1-f.weber@proxmox.com/
^ permalink raw reply [flat|nested] 2+ messages in thread
* Re: [RFC storage/qemu-server] add 'live-migration' hint to volume (de)activation
2026-10-07 11:20 [RFC storage/qemu-server] add 'live-migration' hint to volume (de)activation Roland Kammerer
@ 2026-10-08 10:48 ` Dominik Csapak
0 siblings, 0 replies; 2+ messages in thread
From: Dominik Csapak @ 2026-10-08 10:48 UTC (permalink / raw)
To: Roland Kammerer, pve-devel
Hi,
thanks for the detailed write-up. I'm not convinced we need a new
storage API for this, mostly because I think the problem sits in the
plugin's access model rather than in a missing signal from PVE. Some
details below.
On 10/7/26 1:29 PM, Roland Kammerer wrote:
[snip]
> During a live migration, the
> target QEMU process opens the disks read-write while the source QEMU
> still has them open, so the resource temporarily needs to allow two
> Primaries.
That is true for the open, but there is no concurrent write access. QEMU
hands over ownership of the images during migration, so there is never
more than one writer during live migration. The only requirement on the
storage is that a flush completed on the source is visible to reads on
the target afterwards.
This also matches the contract of the 'shared' flag, which is "a single
storage with the same contents on all nodes". PVE does not expect a
shared storage to arbitrate writers. Ownership is handled above the
storage layer by the node that owns the guest config, by HA fencing,
and by QEMU during migration. All other storage types (AFAIK) allow
concurrent opens/rw and never learn that a migration is happening.
> Specifically for DRBD we only allow two Primaries when using DRBD
> protocol C (the one with the strongest guarantees). Sometime people
> would like to use weaker guarantees (i.e., protocol A and B), which we
> can only allow when we sacrifice live migration as protocols A and B
> don't allow two Primaries at all. If we would know the live migration
> window, we could temporarily upgrade the connections between these 2
> nodes to protocol C and (also temporarily) allow-two-primaries.
As far as I understand protocols A and B, a write is acknowledged before
it has reached the peer's disk. In that case the source's final flush
can complete before the target's replica has the data, and the target
reads stale data right after activation. So for live migration,
protocol C is a correctness requirement.
Does switching from A/B to C handle this correctly?
> Another problem with setting allow-two-primaries permanently is that it
> allows admins, scripts,... to open the in-use device on a peer node and
> unwillingly altering data by accident.
I understand the wish for that protection, but none of our other shared
storages offer it, and PVE does not rely on it. If the plugin wants to
be stricter than the contract, it can already do so on its own, for
example by allowing the second Primary only for the duration of an
activation on another node.
Also a probably better interface would be a storage plugin api
like 'add/end_shared_access' (or similar) that is called before and
after live migration (per storage; with a list of volumes).
We had some discussion of this internally and we're not convinced that
adding this kind of API for a single storage type is justified.
Can you name any other storage that might profit from this?
(I could only think of lvm + lockd maybe, but we don't use that in
favor of our own cluster wide locking)
Best Regards
Dominik
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-10-08 10:48 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-07 11:20 [RFC storage/qemu-server] add 'live-migration' hint to volume (de)activation Roland Kammerer
2026-10-08 10:48 ` Dominik Csapak
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox