From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from gate001.proxmox.com (gate001.proxmox.com [IPv6:2a0f:8001:1:32::40]) by lore.proxmox.com (Postfix) with ESMTPS id CC2111FF0B2 for ; Fri, 25 Sep 2026 14:38:45 +0200 (CEST) Received: from gate001.proxmox.com (localhost.localdomain [127.0.0.1]) by gate001.proxmox.com (Proxmox) with ESMTP id B267F216C4; Fri, 25 Sep 2026 14:38:43 +0200 (CEST) Date: Fri, 25 Sep 2026 14:38:39 +0200 From: Wolfgang Bumiller To: Hannes Laimer Subject: Re: [RFC cluster/manager 00/10] pmxcfs: add a change notification socket Message-ID: References: <20260918144152.575163-1-h.laimer@proxmox.com> <4kp2pm6gasdbailmstjwxrvsc3ozmpiln4tz67pjh4ii3od2eg@ec5poxqld7gh> <7edd3c81-faeb-4ff8-9523-0d2517e490df@proxmox.com> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <7edd3c81-faeb-4ff8-9523-0d2517e490df@proxmox.com> X-Bm-Milter-Handled: 55990f41-d878-4baa-be0a-ee34c49e34d2 X-Bm-Transport-Timestamp: 1790339919518 X-SPAM-LEVEL: Spam detection results: 0 AWL 0.642 Adjusted score from AWL reputation of From: address DMARC_MISSING 0.1 Missing DMARC policy KAM_DMARC_STATUS 0.01 Test Rule for DKIM or SPF Failure with Strict Alignment (newer systems) RCVD_IN_DNSWL_MED -2.3 Sender listed at https://www.dnswl.org/, medium trust SPF_HELO_NONE 0.001 SPF: HELO does not publish an SPF Record SPF_PASS -0.001 SPF: sender matches SPF record Message-ID-Hash: JYGW7634B2LPKY3ZII7C6KR3ACMFSWXX X-Message-ID-Hash: JYGW7634B2LPKY3ZII7C6KR3ACMFSWXX X-MailFrom: w.bumiller@proxmox.com X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; emergency; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header CC: pve-devel@lists.proxmox.com X-Mailman-Version: 3.3.10 Precedence: list List-Id: Proxmox VE development discussion List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: On Fri, Sep 25, 2026 at 02:18:54PM +0200, Hannes Laimer wrote: > On 2026-09-25 14:01, Wolfgang Bumiller wrote: > > On Fri, Sep 18, 2026 at 04:41:42PM +0200, Hannes Laimer wrote: > >> Every daemon that cares about /etc/pve learns about changes by calling > >> cfs_update on each loop iteration and comparing the version vector it > >> gets over the libqb IPC. inotify cannot replace that, since remote > > > > Not strictly true - pmxcfs could add inotify support vis fuse, but it is > > limited to regular files and would not support whole directory watches. > > > >> changes arrive through corosync and never touch the local VFS, so > >> pvestatd, pve-firewall, the HA daemons and pvescheduler all wake up on > >> timers and re-read what usually has not changed. > >> > >> This series adds a push path. pmxcfs gets a Unix stream socket at > >> /run/pve-cluster/pmxcfs.sock, served by a Rust thread linked into the C > >> daemon as a static library with a small C ABI. A client subscribes with > >> named path templates such as nodes/{node}/qemu-server/{vmid}.conf and > >> receives one JSON line per matching mutation, carrying the memdb version > >> as sequence number, the event type, the path and the values the > >> placeholders captured. Connections authorized by group membership never > >> see private paths, as on the IPC and FUSE side. > >> > >> The daemon keeps the last mutations in a fixed size ring and each > >> connection a cursor into it, so the mutating thread only appends and > >> never waits for a client. A slow client is caught up from the ring, one > >> that fell off it gets a single resync event, and a reconnecting client > >> resumes at the last sequence number it saw, also across a restart of > >> pmxcfs when nothing changed meanwhile. A node that takes the whole > >> state from the cluster after a membership change hands its clients a > >> resync event, since no sequence of mutations describes that. A panic in > >> the Rust code only disables the notifier for the rest of the daemon's > >> life. A sequence number identifies a state within one line of history > >> only. A client that missed a resync while disconnected and returns at > >> the same version after a pmxcfs restart resumes as if nothing changed. > >> Carrying the root entry's mtime next to the version would close that. > >> > >> On top of the socket, pve-cluster ships a Perl client with reconnect > >> and resync handling and a hook registry, where a hook is registered > >> like an API method with a path template whose placeholders become the > >> parameters of a run. pve-manager gets the runner that executes those > >> runs in children of a listener hosted by pvescheduler, one run per hook > >> and parameter set in flight and later events for the same set collapsed > >> into a single rerun, so a config update through the API costs one extra > >> run rather than one per save. The listener records the position it has > >> processed up to under /run and resumes there after a reload. No hook > >> ships in this series. A consumer registers one as the example below > >> shows. > >> > >> pve-manager depends on the pve-cluster packages of this series for the > >> new modules and the socket, at build time as well, since its tests run > >> through the check target. > >> > >> Example usage: > >> > >> package PVE::Network::Hooks; > >> > >> use base qw(PVE::Cluster::Hooks); > >> > >> __PACKAGE__->register_hook({ > >> name => 'guest-firewall', > > > > ^ This particular example would have to run in the pve-firewall daemon, > > not pvescheduler, though. > > well, maybe, what the daemon does currently is basically translate > config into nft kernel state periodically. the idea was that it could > stop being periodic, and only one daemon would then work on config > updates.. but could also just run in the fw daemon, not sure if there > are good arguments against having only one daemon that handles such > events We have 2 separate firewall implementations, one is in rust, so pvescheduler taking over the work of the pve-firewall daemon is just a bit weird. It would have to *conditionally* load the firewall packages.