public inbox for pve-devel@lists.proxmox.com
 help / color / mirror / Atom feed
* [PATCH manager 0/4] cephx: migrate keys with a monitor this cluster does not manage
@ 2026-09-27  0:59 Kefu Chai
  2026-09-27  0:59 ` [PATCH manager 1/4] migrations: cephx: ask a monitor about itself through one route Kefu Chai
                   ` (3 more replies)
  0 siblings, 4 replies; 5+ messages in thread
From: Kefu Chai @ 2026-09-27  0:59 UTC (permalink / raw)
  To: pve-devel; +Cc: Thomas Lamprecht

Hi everyone,

A user reported an interesting stretch-cluster setup that
pve-cephx-rotate-service-keys wasn't ready for: a tiebreaker monitor
running on a standalone Proxmox VE host that isn't a node of the
cluster itself, joined to the corosync quorum only through a QDevice
and never listed under /etc/pve/nodes. All five monitors were in
quorum and on the same, capable Ceph version, but the dry run still
failed twice over:

    FAIL: could not reach node 'ch95o-pvewp01': command failed on node
    'ch95o-pvewp01': command '/usr/bin/ssh ... perl - ceph mon:ch95o-pvewp01'
    failed: exit code 255

    could not verify the installed Ceph version of mon. on node
    'ch95o-pvewp01';

Both failures trace back to the same two assumptions: that a daemon
can always be reached with a shell on its host, and that pvestatd on
some node of this cluster always knows what's installed there.
Neither holds for a monitor we don't manage, so this series walks
through getting the helper past both, step by step.

First, we teach it to ask a monitor about itself without needing a
shell at all: over its admin socket where the helper already has a
route, and over 'ceph tell mon.<id>' otherwise. That alone covers the
three questions behind the ssh failure above (auth methods required,
pending-key support, and current sessions).

With the monitor reachable again, the version check is next: rather
than insisting on an installed version from pvestatd, which only
nodes of this cluster report, we judge such a daemon by the version
it runs, since Ceph already reports that for every daemon in the
cluster regardless of who manages it.

That clears the dry run, so the third step is making the actual
rotation work: the 'mon.' key change itself writes to the auth
database and reaches every monitor through paxos, so an unmanaged
monitor gets migrated right along with the others. Only refreshing
its emergency keyring still needs a shell it doesn't have, and the
helper now says so plainly instead of refusing the whole rotation.

Last, since that emergency keyring (and any host's own 'client.admin'
keyring) is left stale by design, we make sure the helper tells the
operator exactly which hosts and files need a manual refresh, and how.

To check all this against a real cluster and not just the test suite, I
built a nested lab matching the reported shape: three Proxmox VE nodes
running as VMs, clustered with Ceph and a monitor each on the first two,
and a fourth VM running only ceph-mon, joined straight to the Ceph
cluster without ever joining the Proxmox VE cluster, playing the
tiebreaker. Running the patched helper there end to end, from dry run
through the monitor-key rotation to the final client-key notes, matched
what the series is meant to do.

Together these four steps let the reported cluster's OSD, manager, MDS and
monitor keys rotate with the tiebreaker happily in place.
'client.admin' and per-storage keys still need a manual procedure for
now, that's a separate piece of work I'm looking at next.

As always, happy to hear any thoughts or concerns, thanks for reading!

Kefu Chai (4):
  migrations: cephx: ask a monitor about itself through one route
  migrations: cephx: judge cipher support by the version a daemon runs
  migrations: cephx: rotate the monitor key with monitors we do not
    manage
  migrations: cephx: name the hosts keeping key copies we cannot write

 PVE/Ceph/KeyMigration.pm            |  19 +-
 bin/pve-cephx-rotate-service-keys   | 254 +++++++++++++++----
 test/CephKeyMigrationScript_test.pl | 366 +++++++++++++++++++++++++++-
 test/CephKeyMigration_test.pl       |  17 ++
 4 files changed, 587 insertions(+), 69 deletions(-)

-- 
2.47.3





^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-09-27  1:00 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-27  0:59 [PATCH manager 0/4] cephx: migrate keys with a monitor this cluster does not manage Kefu Chai
2026-09-27  0:59 ` [PATCH manager 1/4] migrations: cephx: ask a monitor about itself through one route Kefu Chai
2026-09-27  0:59 ` [PATCH manager 2/4] migrations: cephx: judge cipher support by the version a daemon runs Kefu Chai
2026-09-27  0:59 ` [PATCH manager 3/4] migrations: cephx: rotate the monitor key with monitors we do not manage Kefu Chai
2026-09-27  0:59 ` [PATCH manager 4/4] migrations: cephx: name the hosts keeping key copies we cannot write Kefu Chai

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Service provided by Proxmox Server Solutions GmbH | Privacy | Legal