From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from gate001.proxmox.com (gate001.proxmox.com [IPv6:2a0f:8001:1:32::40]) by lore.proxmox.com (Postfix) with ESMTPS id 4C2571FF0AF for ; Thu, 08 Oct 2026 12:48:33 +0200 (CEST) Received: from gate001.proxmox.com (localhost.localdomain [127.0.0.1]) by gate001.proxmox.com (Proxmox) with ESMTP id A61A0214D8; Thu, 08 Oct 2026 12:48:30 +0200 (CEST) Message-ID: <4908cd76-2678-45aa-b101-8bf095c47e16@proxmox.com> Date: Thu, 8 Oct 2026 12:48:18 +0200 MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Beta Subject: Re: [RFC storage/qemu-server] add 'live-migration' hint to volume (de)activation To: Roland Kammerer , pve-devel@lists.proxmox.com References: Content-Language: en-US From: Dominik Csapak In-Reply-To: Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 7bit X-Bm-Milter-Handled: 55990f41-d878-4baa-be0a-ee34c49e34d2 X-Bm-Transport-Timestamp: 1791456506258 X-SPAM-LEVEL: Spam detection results: 0 AWL 0.377 Adjusted score from AWL reputation of From: address DMARC_MISSING 0.1 Missing DMARC policy KAM_DMARC_STATUS 0.01 Test Rule for DKIM or SPF Failure with Strict Alignment (newer systems) RCVD_IN_DNSWL_MED -2.3 Sender listed at https://www.dnswl.org/, medium trust SPF_HELO_NONE 0.001 SPF: HELO does not publish an SPF Record SPF_PASS -0.001 SPF: sender matches SPF record Message-ID-Hash: IWKJ7HWXA7KI7FTSMGJVIHWIJ5V45TJV X-Message-ID-Hash: IWKJ7HWXA7KI7FTSMGJVIHWIJ5V45TJV X-MailFrom: d.csapak@proxmox.com X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; emergency; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header X-Mailman-Version: 3.3.10 Precedence: list List-Id: Proxmox VE development discussion List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: Hi, thanks for the detailed write-up. I'm not convinced we need a new storage API for this, mostly because I think the problem sits in the plugin's access model rather than in a missing signal from PVE. Some details below. On 10/7/26 1:29 PM, Roland Kammerer wrote: [snip] > During a live migration, the > target QEMU process opens the disks read-write while the source QEMU > still has them open, so the resource temporarily needs to allow two > Primaries. That is true for the open, but there is no concurrent write access. QEMU hands over ownership of the images during migration, so there is never more than one writer during live migration. The only requirement on the storage is that a flush completed on the source is visible to reads on the target afterwards. This also matches the contract of the 'shared' flag, which is "a single storage with the same contents on all nodes". PVE does not expect a shared storage to arbitrate writers. Ownership is handled above the storage layer by the node that owns the guest config, by HA fencing, and by QEMU during migration. All other storage types (AFAIK) allow concurrent opens/rw and never learn that a migration is happening. > Specifically for DRBD we only allow two Primaries when using DRBD > protocol C (the one with the strongest guarantees). Sometime people > would like to use weaker guarantees (i.e., protocol A and B), which we > can only allow when we sacrifice live migration as protocols A and B > don't allow two Primaries at all. If we would know the live migration > window, we could temporarily upgrade the connections between these 2 > nodes to protocol C and (also temporarily) allow-two-primaries. As far as I understand protocols A and B, a write is acknowledged before it has reached the peer's disk. In that case the source's final flush can complete before the target's replica has the data, and the target reads stale data right after activation. So for live migration, protocol C is a correctness requirement. Does switching from A/B to C handle this correctly? > Another problem with setting allow-two-primaries permanently is that it > allows admins, scripts,... to open the in-use device on a peer node and > unwillingly altering data by accident. I understand the wish for that protection, but none of our other shared storages offer it, and PVE does not rely on it. If the plugin wants to be stricter than the contract, it can already do so on its own, for example by allowing the second Primary only for the duration of an activation on another node. Also a probably better interface would be a storage plugin api like 'add/end_shared_access' (or similar) that is called before and after live migration (per storage; with a list of volumes). We had some discussion of this internally and we're not convinced that adding this kind of API for a single storage type is justified. Can you name any other storage that might profit from this? (I could only think of lvm + lockd maybe, but we don't use that in favor of our own cluster wide locking) Best Regards Dominik