From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from gate001.proxmox.com (gate001.proxmox.com [45.144.208.40]) by lore.proxmox.com (Postfix) with ESMTPS id 8B21F1FF124 for ; Sun, 27 Sep 2026 02:59:23 +0200 (CEST) Received: from gate001.proxmox.com (localhost.localdomain [127.0.0.1]) by gate001.proxmox.com (Proxmox) with ESMTP id E7CF721623; Sun, 27 Sep 2026 02:59:16 +0200 (CEST) From: Kefu Chai To: pve-devel@lists.proxmox.com Subject: [PATCH manager 0/4] cephx: migrate keys with a monitor this cluster does not manage Date: Sun, 27 Sep 2026 08:59:02 +0800 Message-ID: <20260927005906.4184138-1-k.chai@proxmox.com> X-Mailer: git-send-email 2.47.3 MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-Bm-Milter-Handled: 55990f41-d878-4baa-be0a-ee34c49e34d2 X-Bm-Transport-Timestamp: 1790470749633 X-SPAM-LEVEL: Spam detection results: 0 AWL -0.980 Adjusted score from AWL reputation of From: address DMARC_MISSING 0.1 Missing DMARC policy KAM_DMARC_STATUS 0.01 Test Rule for DKIM or SPF Failure with Strict Alignment (newer systems) SPF_HELO_NONE 0.001 SPF: HELO does not publish an SPF Record SPF_PASS -0.001 SPF: sender matches SPF record Message-ID-Hash: 5Q2BYA3GLDYKT3AH23ATXPKLJHZBT2U5 X-Message-ID-Hash: 5Q2BYA3GLDYKT3AH23ATXPKLJHZBT2U5 X-MailFrom: k.chai@proxmox.com X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; emergency; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header CC: Thomas Lamprecht X-Mailman-Version: 3.3.10 Precedence: list List-Id: Proxmox VE development discussion List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: Hi everyone, A user reported an interesting stretch-cluster setup that pve-cephx-rotate-service-keys wasn't ready for: a tiebreaker monitor running on a standalone Proxmox VE host that isn't a node of the cluster itself, joined to the corosync quorum only through a QDevice and never listed under /etc/pve/nodes. All five monitors were in quorum and on the same, capable Ceph version, but the dry run still failed twice over: FAIL: could not reach node 'ch95o-pvewp01': command failed on node 'ch95o-pvewp01': command '/usr/bin/ssh ... perl - ceph mon:ch95o-pvewp01' failed: exit code 255 could not verify the installed Ceph version of mon. on node 'ch95o-pvewp01'; Both failures trace back to the same two assumptions: that a daemon can always be reached with a shell on its host, and that pvestatd on some node of this cluster always knows what's installed there. Neither holds for a monitor we don't manage, so this series walks through getting the helper past both, step by step. First, we teach it to ask a monitor about itself without needing a shell at all: over its admin socket where the helper already has a route, and over 'ceph tell mon.' otherwise. That alone covers the three questions behind the ssh failure above (auth methods required, pending-key support, and current sessions). With the monitor reachable again, the version check is next: rather than insisting on an installed version from pvestatd, which only nodes of this cluster report, we judge such a daemon by the version it runs, since Ceph already reports that for every daemon in the cluster regardless of who manages it. That clears the dry run, so the third step is making the actual rotation work: the 'mon.' key change itself writes to the auth database and reaches every monitor through paxos, so an unmanaged monitor gets migrated right along with the others. Only refreshing its emergency keyring still needs a shell it doesn't have, and the helper now says so plainly instead of refusing the whole rotation. Last, since that emergency keyring (and any host's own 'client.admin' keyring) is left stale by design, we make sure the helper tells the operator exactly which hosts and files need a manual refresh, and how. To check all this against a real cluster and not just the test suite, I built a nested lab matching the reported shape: three Proxmox VE nodes running as VMs, clustered with Ceph and a monitor each on the first two, and a fourth VM running only ceph-mon, joined straight to the Ceph cluster without ever joining the Proxmox VE cluster, playing the tiebreaker. Running the patched helper there end to end, from dry run through the monitor-key rotation to the final client-key notes, matched what the series is meant to do. Together these four steps let the reported cluster's OSD, manager, MDS and monitor keys rotate with the tiebreaker happily in place. 'client.admin' and per-storage keys still need a manual procedure for now, that's a separate piece of work I'm looking at next. As always, happy to hear any thoughts or concerns, thanks for reading! Kefu Chai (4): migrations: cephx: ask a monitor about itself through one route migrations: cephx: judge cipher support by the version a daemon runs migrations: cephx: rotate the monitor key with monitors we do not manage migrations: cephx: name the hosts keeping key copies we cannot write PVE/Ceph/KeyMigration.pm | 19 +- bin/pve-cephx-rotate-service-keys | 254 +++++++++++++++---- test/CephKeyMigrationScript_test.pl | 366 +++++++++++++++++++++++++++- test/CephKeyMigration_test.pl | 17 ++ 4 files changed, 587 insertions(+), 69 deletions(-) -- 2.47.3