From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from gate001.proxmox.com (gate001.proxmox.com [45.144.208.40]) by lore.proxmox.com (Postfix) with ESMTPS id B8E4B1FF0AF for ; Thu, 08 Oct 2026 15:39:34 +0200 (CEST) Received: from gate001.proxmox.com (localhost.localdomain [127.0.0.1]) by gate001.proxmox.com (Proxmox) with ESMTP id 319B821478; Thu, 08 Oct 2026 15:39:31 +0200 (CEST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linbit.com; s=google; t=1791466757; x=1792071557; darn=lists.proxmox.com; h=user-agent:in-reply-to:content-disposition:content-type :mime-version:references:message-id:subject:cc:to:from:date:from:to :cc:subject:date:message-id:reply-to:content-type; bh=NTwezt4AOgOmdzjc1zli857Zk5hMcV7X98/VHFsBFwc=; b=Ec+ATMZaw5zNMpLnVLJ9Hy0L3e7jdzs9kdfpKF4aCkS/yoZI4AhN9E+XsnOMO/vo9J n9eWDUW6rr2ydVbmMX6knGgQs+Oi95/GzB5mR6UgtnAa6twrfi7+N3TX+qNZiq16dZ9d phMZNT36cn6KPUJ+Ui7UYlkn/ItFoVBDyCw/JS9/PCJ4hz7HJqj3/HPeZEzZFkviIKev XBkPsxTAfTp91eC9C6ACroOZ3am0mv6EEc+mMqdoEGXrgQTS19Q1VURGeKfec0402KrA btLrLu1AouhnmEGouRA/JCqqALsH27i3e3/yIBsuIA5m55P1bD0SX0qC6LSqdMIc9HQf 0r6Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1791466757; x=1792071557; h=user-agent:in-reply-to:content-disposition:content-type :mime-version:references:message-id:subject:cc:to:from:date:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=NTwezt4AOgOmdzjc1zli857Zk5hMcV7X98/VHFsBFwc=; b=iStRiKhO9rvgp6Cnl5p9WlxZTihtzIm9PGAcmwUKZCNSB+8SEX0r7mNCb4mUdKvr6F wKEfSx3AFzLlGAvUqDvXDJHesy30MPh9Db7VOZq7nG1qWzJSiQPpnBjcCnWpAxE6RJw4 S82ErNOiarPAR/AzAoXAEir7HLCc01jOBVv9pK06ZUqzomrwkHmBRaj32+FimBtUIV/P fGveOE/CcZonCtV6OkZKSf8SWBxi8zvyQItM6Js8844o9tEpLJmJkIAMMwmvJhdGg0jJ JYbtkcmhGa5Mc5dKBLwe4oqc2g2iKZbD/HS5t81kF8iY80RdlpCllgZJkmyRMglKIt9a cw2Q== X-Gm-Message-State: AFq9FYL3UNHAydVGi3i6aY2OmM/GiiKFgIctfNlqmp5/8NBgqxo8Hq1z /bZVrvEGc1Rm22/W7/ehYHfk4pQco18rO2ssALQgmghnimUZKbhXUCSnDlK6kdjc5/tYv2Ofxc+ JHTOMQdF8tQ== X-Gm-Gg: AYBFou35qudH8wkv3SpxoUtALySMAWxlYk4CicTtmSnZZoXuOwbwRrWFGA/e60xyVm4 HWveOa9wsLyHmXezX9yaymYLGqX45M7lzlYYpEYozXrgyNzNDFe8HLDsymFchJIxEbWSCGVV6aG 9k9S6efvwuqNXk8lnaPZVRCepSy71fTkeu6p4rekqLy9KKd9NJ1JTHvdW7qUEk2rzC4BYuiPUdv BsnCYLR4ai0OimO/hZaxQHov+OKvXqF+n0R/5oXdRyQVSFmQp6PAlc7RIuKacQKaVMjNmwZ2wep yL0iYYXsS4ZFy3zx7HpjC2zHuzXBBhqKrP7kRMA9RPdLkSbiH1AjiiJRhCjWSdzcGSVngTKTgU9 J+Bfto3ZuMOiVj49Xe/1IHx4geSH40aO66vVzvfNCbZEn6ptDSUqjmZssae+m7tkVttaM1N+OKc 650c52o/f6PT+5ClhsU2ujJEkkzJLcWv8dhr4x7TVB/4T3H2yUVG551GAw/K8sgZyJVEJ6nzQYk tIIZKRZe7dvHJ8KcS+iLMHrXrqbTWttQcsft4aQ3KmrOJa5J9c= X-Received: by 2002:a05:6512:6c8:b0:5ba:3c46:bc4e with SMTP id 2adb3069b0e04-5bcd06f3f0bmr2196873e87.47.1791466757265; Thu, 08 Oct 2026 06:39:17 -0700 (PDT) Date: Thu, 8 Oct 2026 15:39:14 +0200 From: Roland Kammerer To: pve-devel@lists.proxmox.com Subject: Re: [RFC storage/qemu-server] add 'live-migration' hint to volume (de)activation Message-ID: References: <4908cd76-2678-45aa-b101-8bf095c47e16@proxmox.com> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <4908cd76-2678-45aa-b101-8bf095c47e16@proxmox.com> User-Agent: Mutt/2.3.3 (2026-06-12) X-SPAM-LEVEL: Spam detection results: 0 AWL 0.400 Adjusted score from AWL reputation of From: address DKIM_SIGNED 0.1 Message has a DKIM or DK signature, not necessarily valid DKIM_VALID -0.1 Message has at least one valid DKIM or DK signature DKIM_VALID_AU -0.1 Message has a valid DKIM or DK signature from author's domain DKIM_VALID_EF -0.1 Message has a valid DKIM or DK signature from envelope-from domain DMARC_PASS -0.1 DMARC pass policy RCVD_IN_DNSWL_NONE -0.0001 Sender listed at https://www.dnswl.org/, no trust SPF_HELO_NONE 0.001 SPF: HELO does not publish an SPF Record SPF_PASS -0.001 SPF: sender matches SPF record Message-ID-Hash: MYNA5MPOFLYAM6QXNDKSV2JELMPUJLAI X-Message-ID-Hash: MYNA5MPOFLYAM6QXNDKSV2JELMPUJLAI X-MailFrom: roland.kammerer@linbit.com X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; emergency; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header X-Mailman-Version: 3.3.10 Precedence: list List-Id: Proxmox VE development discussion List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: On Thu, Oct 08, 2026 at 12:48:18PM +0200, Dominik Csapak wrote: > Hi, > > thanks for the detailed write-up. I'm not convinced we need a new > storage API for this, mostly because I think the problem sits in the > plugin's access model rather than in a missing signal from PVE. Some > details below. > > > On 10/7/26 1:29 PM, Roland Kammerer wrote: > [snip] > > During a live migration, the > > target QEMU process opens the disks read-write while the source QEMU > > still has them open, so the resource temporarily needs to allow two > > Primaries. > > That is true for the open, but there is no concurrent write access. QEMU > hands over ownership of the images during migration, so there is never > more than one writer during live migration. The only requirement on the > storage is that a flush completed on the source is visible to reads on > the target afterwards. Sure, I know that, otherwise we would have a serious problem and we need to trust qemu/pve there anyways (with our without the live-migration hints). > This also matches the contract of the 'shared' flag, which is "a single > storage with the same contents on all nodes". PVE does not expect a > shared storage to arbitrate writers. Ownership is handled above the > storage layer by the node that owns the guest config, by HA fencing, > and by QEMU during migration. All other storage types (AFAIK) allow > concurrent opens/rw and never learn that a migration is happening. True, that is how they work, they are shared storage. DRBD on the other hand is a bit special here. It can act "shared"/dual primary, but in general we would like to keep that window as small as possible. Actually in the best case, and what DRBD9 describes as the only supported case, we only want to allow that during live migration. But for that we would need to know when one happens :). > > Specifically for DRBD we only allow two Primaries when using DRBD > > protocol C (the one with the strongest guarantees). Sometime people > > would like to use weaker guarantees (i.e., protocol A and B), which we > > can only allow when we sacrifice live migration as protocols A and B > > don't allow two Primaries at all. If we would know the live migration > > window, we could temporarily upgrade the connections between these 2 > > nodes to protocol C and (also temporarily) allow-two-primaries. > > As far as I understand protocols A and B, a write is acknowledged before > it has reached the peer's disk. In that case the source's final flush > can complete before the target's replica has the data, and the target > reads stale data right after activation. So for live migration, > protocol C is a correctness requirement. > Does switching from A/B to C handle this correctly? > I assume so, otherwise that would be a problem the DRBD devs have to solve. > > Another problem with setting allow-two-primaries permanently is that it > > allows admins, scripts,... to open the in-use device on a peer node and > > unwillingly altering data by accident. > > I understand the wish for that protection, but none of our other shared > storages offer it, and PVE does not rely on it. mhm, that is what I tried to answer above, they are how they are and can't do better, DRBD could. > If the plugin wants to be stricter than the contract, it can already > do so on its own, for example by allowing the second Primary only for > the duration of an activation on another node. I tried that, but it was always guess work, it never completely worked out/felt right. AFAIR activate/deactivate don't have to be strictly symmetrical, activates can happen for different reasons,... Probably with more heuristics and more guessing one can come up with something, but at some point it looked easier to just pass down the information I'm really interested in - live migrations start/end - to the plugin and be done with it. Then it would have become trivial in the plugin without guessing why an activate happens. > Also a probably better interface would be a storage plugin api > like 'add/end_shared_access' (or similar) that is called before and > after live migration (per storage; with a list of volumes). Yes, I saw that alternative too, the proposal was long enough already, but certainly this would have been fine as well. "Something that passes start/end of live migration". > We had some discussion of this internally and we're not convinced that > adding this kind of API for a single storage type is justified. > Can you name any other storage that might profit from this? > (I could only think of lvm + lockd maybe, but we don't use that in > favor of our own cluster wide locking) Fair enough and thanks for the review and comments. Today I learned that there is a DRBD feature in the pipeline where a write on the second Primary if in dual primary forces the first one to become Secondary, all within DRBD. That then will fix the most important part of our scenario: Not letting people shoot themselves in the foot if one could do better (e.g., by minimizing the "dangerous window" via the proposed live-migration hints). Thanks, rck