* [PATCH manager 0/4] cephx: migrate keys with a monitor this cluster does not manage
@ 2026-09-27 0:59 Kefu Chai
2026-09-27 0:59 ` [PATCH manager 1/4] migrations: cephx: ask a monitor about itself through one route Kefu Chai
` (3 more replies)
0 siblings, 4 replies; 5+ messages in thread
From: Kefu Chai @ 2026-09-27 0:59 UTC (permalink / raw)
To: pve-devel; +Cc: Thomas Lamprecht
Hi everyone,
A user reported an interesting stretch-cluster setup that
pve-cephx-rotate-service-keys wasn't ready for: a tiebreaker monitor
running on a standalone Proxmox VE host that isn't a node of the
cluster itself, joined to the corosync quorum only through a QDevice
and never listed under /etc/pve/nodes. All five monitors were in
quorum and on the same, capable Ceph version, but the dry run still
failed twice over:
FAIL: could not reach node 'ch95o-pvewp01': command failed on node
'ch95o-pvewp01': command '/usr/bin/ssh ... perl - ceph mon:ch95o-pvewp01'
failed: exit code 255
could not verify the installed Ceph version of mon. on node
'ch95o-pvewp01';
Both failures trace back to the same two assumptions: that a daemon
can always be reached with a shell on its host, and that pvestatd on
some node of this cluster always knows what's installed there.
Neither holds for a monitor we don't manage, so this series walks
through getting the helper past both, step by step.
First, we teach it to ask a monitor about itself without needing a
shell at all: over its admin socket where the helper already has a
route, and over 'ceph tell mon.<id>' otherwise. That alone covers the
three questions behind the ssh failure above (auth methods required,
pending-key support, and current sessions).
With the monitor reachable again, the version check is next: rather
than insisting on an installed version from pvestatd, which only
nodes of this cluster report, we judge such a daemon by the version
it runs, since Ceph already reports that for every daemon in the
cluster regardless of who manages it.
That clears the dry run, so the third step is making the actual
rotation work: the 'mon.' key change itself writes to the auth
database and reaches every monitor through paxos, so an unmanaged
monitor gets migrated right along with the others. Only refreshing
its emergency keyring still needs a shell it doesn't have, and the
helper now says so plainly instead of refusing the whole rotation.
Last, since that emergency keyring (and any host's own 'client.admin'
keyring) is left stale by design, we make sure the helper tells the
operator exactly which hosts and files need a manual refresh, and how.
To check all this against a real cluster and not just the test suite, I
built a nested lab matching the reported shape: three Proxmox VE nodes
running as VMs, clustered with Ceph and a monitor each on the first two,
and a fourth VM running only ceph-mon, joined straight to the Ceph
cluster without ever joining the Proxmox VE cluster, playing the
tiebreaker. Running the patched helper there end to end, from dry run
through the monitor-key rotation to the final client-key notes, matched
what the series is meant to do.
Together these four steps let the reported cluster's OSD, manager, MDS and
monitor keys rotate with the tiebreaker happily in place.
'client.admin' and per-storage keys still need a manual procedure for
now, that's a separate piece of work I'm looking at next.
As always, happy to hear any thoughts or concerns, thanks for reading!
Kefu Chai (4):
migrations: cephx: ask a monitor about itself through one route
migrations: cephx: judge cipher support by the version a daemon runs
migrations: cephx: rotate the monitor key with monitors we do not
manage
migrations: cephx: name the hosts keeping key copies we cannot write
PVE/Ceph/KeyMigration.pm | 19 +-
bin/pve-cephx-rotate-service-keys | 254 +++++++++++++++----
test/CephKeyMigrationScript_test.pl | 366 +++++++++++++++++++++++++++-
test/CephKeyMigration_test.pl | 17 ++
4 files changed, 587 insertions(+), 69 deletions(-)
--
2.47.3
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH manager 1/4] migrations: cephx: ask a monitor about itself through one route
2026-09-27 0:59 [PATCH manager 0/4] cephx: migrate keys with a monitor this cluster does not manage Kefu Chai
@ 2026-09-27 0:59 ` Kefu Chai
2026-09-27 0:59 ` [PATCH manager 2/4] migrations: cephx: judge cipher support by the version a daemon runs Kefu Chai
` (2 subsequent siblings)
3 siblings, 0 replies; 5+ messages in thread
From: Kefu Chai @ 2026-09-27 0:59 UTC (permalink / raw)
To: pve-devel; +Cc: Thomas Lamprecht
The helper asks a monitor three things about itself: which auth methods
it requires, whether it can hold two valid client keys, and which
clients have a session with it. All three go through the admin socket,
which needs no cephx and works even while auth is broken, but does need
a shell on the monitor's host.
A daemon can run on a host that is not a node of this Proxmox VE
cluster, such as the tiebreaker monitor of a stretch cluster. Nothing of
ours reaches such a host: no SSH trust, no pmxcfs, no pvestatd. The
probes fail, so the helper treats the monitor as unreachable, refuses to
stage a client key, and reports an incomplete session view.
Record whether a daemon's node is one we manage, and route the three
probes through that. 'ceph tell' reaches the same admin socket over the
network, but needs cephx and a quorum, so it is the fallback, not the
default. It runs as 'mon.' so it keeps working during a 'client.admin'
rotation, and carries a connect timeout so a monitor across a WAN link
cannot stall a run.
Signed-off-by: Kefu Chai <k.chai@proxmox.com>
---
PVE/Ceph/KeyMigration.pm | 19 +++----
bin/pve-cephx-rotate-service-keys | 78 ++++++++++++++++-------------
test/CephKeyMigrationScript_test.pl | 62 +++++++++++++++++++++--
3 files changed, 110 insertions(+), 49 deletions(-)
diff --git a/PVE/Ceph/KeyMigration.pm b/PVE/Ceph/KeyMigration.pm
index d0613e8b..1ad73826 100644
--- a/PVE/Ceph/KeyMigration.pm
+++ b/PVE/Ceph/KeyMigration.pm
@@ -1101,15 +1101,16 @@ sub merge_configured_daemons($daemons, $type, $configured, $existing = undef) {
push @$ghosts, { type => $type, id => "$id", node => $configured->{$id} };
next;
}
- push @$daemons,
- {
- type => $type,
- id => "$id",
- entity => $type eq 'mon' ? 'mon.' : "$type.$id",
- node => $configured->{$id},
- down => 1,
- $type eq 'osd' && $existing ? ('osd-uuid' => $existing->{$id}) : (),
- };
+ push @$daemons, {
+ type => $type,
+ id => "$id",
+ entity => $type eq 'mon' ? 'mon.' : "$type.$id",
+ node => $configured->{$id},
+ # pvestatd publishes this inventory, so the node is one of this cluster's own
+ managed => 1,
+ down => 1,
+ $type eq 'osd' && $existing ? ('osd-uuid' => $existing->{$id}) : (),
+ };
}
return wantarray ? ($daemons, $ghosts) : $daemons;
diff --git a/bin/pve-cephx-rotate-service-keys b/bin/pve-cephx-rotate-service-keys
index b60f3619..95d10266 100755
--- a/bin/pve-cephx-rotate-service-keys
+++ b/bin/pve-cephx-rotate-service-keys
@@ -175,11 +175,20 @@ my sub cluster_nodes() {
return { map { $_ => 1 } PVE::Cluster::get_nodelist()->@* };
}
+# Where a daemon runs decides what this cluster can do with it. 'managed' means its host is a node
+# of this Proxmox VE cluster, so the cluster's SSH trust, pmxcfs and pvestatd cover it: its files,
+# its unit and its installed packages. A daemon without it is still part of the Ceph cluster and
+# still answers over Ceph, the tiebreaker monitor of a stretch cluster for example.
my sub daemon_location($type, $id, $host) {
- return { node => undef } if !defined($host) || (!ref($host) && !length($host));
+ return { node => undef, managed => 0 } if !defined($host) || (!ref($host) && !length($host));
die "Cannot locate '$type.$id': no valid host name was reported\n" if ref($host);
- my $node = PVE::Ceph::Services::metadata_host_node($host, cluster_nodes());
- return { node => $node, $node ne $host ? ('metadata-host' => $host) : () };
+ my $nodes = cluster_nodes();
+ my $node = PVE::Ceph::Services::metadata_host_node($host, $nodes);
+ return {
+ node => $node,
+ managed => $nodes->{$node} ? 1 : 0,
+ $node ne $host ? ('metadata-host' => $host) : (),
+ };
}
my sub assert_member($node, $entity) {
@@ -559,13 +568,14 @@ my sub auth_entry($rados, $entity) {
}
# the monitor identity keeps working while a client.admin rotation invalidates the default keyring
-my sub monitor_command($args) {
+my sub monitor_command($args, $timeout = undef) {
return node_run(
$nodename,
[
'ceph', '--cluster', $ccname, '--name', 'mon.', '--keyring', $pve_mon_keyring,
@$args,
],
+ defined($timeout) ? (timeout => $timeout) : (),
);
}
@@ -579,19 +589,31 @@ my sub monitor_auth_entry($entity) {
return $res->[0];
}
+# Everything a monitor is asked about itself goes through its own admin socket, which needs no
+# cephx and answers while the authentication layer is in trouble. A monitor on a host outside this
+# cluster has no socket we can reach, and 'ceph tell' hands the command to that same socket over
+# the network, so the answer is identical. That route needs cephx and a quorum, which is why it is
+# the fallback and not the rule. It runs as 'mon.', which keeps answering while a 'client.admin'
+# rotation is in flight.
+my sub mon_admin_run($node, $id, $words, $run_node = undef) {
+ if (defined($node) && cluster_nodes()->{$node}) {
+ my $cmd = ['ceph', 'daemon', "mon.$id", @$words];
+ return $run_node ? $run_node->($node, $cmd) : node_run($node, $cmd, entity => "mon.$id");
+ }
+
+ # a monitor across a WAN link answers in well under this, while one that cannot be reached at
+ # all must not hold the run the way the 300 second default would
+ return monitor_command(['--connect-timeout', '10', 'tell', "mon.$id", @$words], 20);
+}
+
my sub poll_monitor_authentication($monitors, $nodes, $run_node = undef) {
$run_node //= sub($node, $command) { node_run($node, $command, entity => $command->[2]) };
my ($reports, $errors) = ({}, {});
for my $mon (@$monitors) {
my $option = 'auth_service_required';
my $reply = eval {
- assert_member($nodes->{$mon}, "mon.$mon");
- decode_json(
- $run_node->(
- $nodes->{$mon},
- ['ceph', 'daemon', "mon.$mon", 'config', 'get', $option],
- ),
- );
+ decode_json(mon_admin_run(
+ $nodes->{$mon}, $mon, ['config', 'get', $option], $run_node));
};
my $error = $@;
$reports->{$mon}->{$option} = ref($reply) eq 'HASH' ? $reply->{$option} : undef;
@@ -789,20 +811,12 @@ for my $fsid (@fsids) {
}
PERL
-# Every monitor is asked over its admin socket, which needs no cephx and answers even while the
-# authentication layer itself is in trouble. One that does not answer marks the result incomplete,
-# so the checks built on it stay honest.
+# A monitor that does not answer marks the result incomplete, so the checks built on it stay honest.
my sub poll_client_sessions($info) {
my $per_mon = [];
for my $mon ($info->{daemons}->{mon}->@*) {
next if $mon->{down};
- my $sessions = eval {
- decode_json(node_run(
- $mon->{node},
- ['ceph', 'daemon', "mon.$mon->{id}", 'sessions'],
- entity => "mon.$mon->{id}",
- ));
- };
+ my $sessions = eval { decode_json(mon_admin_run($mon->{node}, $mon->{id}, ['sessions'])) };
push @$per_mon, { mon => $mon->{id}, sessions => $sessions };
}
return summarize_sessions($per_mon, $info->{monmap_mons});
@@ -811,15 +825,14 @@ my sub poll_client_sessions($info) {
# Whether a monitor can keep two client keys valid is told by its admin socket, which needs no
# cephx. One that answers but does not report the option is old; one that does not answer at all
# cannot be told apart from a broken connection, and rules staging out just the same.
-my sub probe_manual_promotion($node, $id, $run_node) {
- my $out =
- eval { $run_node->($node, ['ceph', 'daemon', "mon.$id", 'config', 'get', $GRACE_OPTION]) };
+my sub probe_manual_promotion($node, $id, $run_node = undef) {
+ my $out = eval { mon_admin_run($node, $id, ['config', 'get', $GRACE_OPTION], $run_node) };
if (!$@) {
my $res = eval { decode_json($out) };
my $value = ref($res) eq 'HASH' ? $res->{$GRACE_OPTION} : undef;
return { reached => 1, value => defined($value) && !ref($value) ? "$value" : undef };
}
- my $reached = eval { $run_node->($node, ['ceph', 'daemon', "mon.$id", 'version']); 1 } ? 1 : 0;
+ my $reached = eval { mon_admin_run($node, $id, ['version'], $run_node); 1 } ? 1 : 0;
return { reached => $reached, value => undef };
}
@@ -827,11 +840,7 @@ my sub poll_manual_promotion($info) {
my $reports = {};
for my $mon ($info->{daemons}->{mon}->@*) {
next if $mon->{down};
- $reports->{ $mon->{id} } = probe_manual_promotion(
- $mon->{node},
- $mon->{id},
- sub($node, $cmd) { node_run($node, $cmd, entity => "mon.$mon->{id}") },
- );
+ $reports->{ $mon->{id} } = probe_manual_promotion($mon->{node}, $mon->{id});
}
return manual_promotion_support($reports, $info->{monmap_mons});
}
@@ -936,13 +945,9 @@ my sub collect_current_monitor_state($rados, $run_node = undef) {
$errors->{$id} = 'not in quorum';
} elsif (!$nodes->{$id}) {
$errors->{$id} = 'no hostname in monitor metadata';
- } elsif (!eval { assert_member($nodes->{$id}, "mon.$id"); 1 }) {
- $errors->{$id} = $@;
} else {
- $sessions = eval {
- decode_json($run_node->(
- $nodes->{$id}, ['ceph', 'daemon', "mon.$id", 'sessions']));
- };
+ $sessions =
+ eval { decode_json(mon_admin_run($nodes->{$id}, $id, ['sessions'], $run_node)) };
my $error = $@;
if ($error) {
$error =~ s/\s+/ /g;
@@ -6258,6 +6263,7 @@ sub key_migration_test_hooks {
client_key_files => \&client_key_files,
check_client_kernels => \&check_client_kernels,
manual_promotion_with_retries => \&manual_promotion_with_retries,
+ mon_admin_run => \&mon_admin_run,
describe_live => \&describe_live,
};
}
diff --git a/test/CephKeyMigrationScript_test.pl b/test/CephKeyMigrationScript_test.pl
index a53b1430..fa1cee4e 100755
--- a/test/CephKeyMigrationScript_test.pl
+++ b/test/CephKeyMigrationScript_test.pl
@@ -6595,7 +6595,16 @@ for my $case ([0, 0], [1, 0], [0, 1], [1, 1]) {
$rados->{host} = 'foreign.invalid';
@commands = ();
$info = $HOOKS->{collect_cluster_info}->($rados, { 'rotate-lockbox-keys' => 1 }, {});
- is_deeply(\@commands, [], 'non-member gets no commands during pre-preflight collection');
+ is(
+ scalar(grep { /foreign\.invalid/ } map { join(' ', @$_) } @commands),
+ 0,
+ 'pre-preflight collection sends nothing to the non-member host',
+ );
+ is(
+ scalar(grep { /tell mon\.a config get/ } map { join(' ', @$_) } @commands),
+ 1,
+ 'the monitor outside the cluster is asked with tell instead of its socket',
+ );
like(
$info->{lockbox}->{'client.osd-lockbox.osd-uuid'}->{missing},
qr/osd\.7.*foreign\.invalid.*not a node/,
@@ -6604,10 +6613,14 @@ for my $case ([0, 0], [1, 0], [0, 1], [1, 1]) {
@targets = ();
$monitor = $HOOKS->{collect_monitor_state}->($rados, sub { push @targets, $_[0]; return '[]' });
is_deeply(\@targets, [], 'fresh monitor collection never queries a non-member');
- like(
+ is(
$monitor->{sessions}->{errors}->{a},
- qr/mon\.a.*not a node/,
- 'monitor diagnostic names membership',
+ undef,
+ 'and its sessions come over the network, so nothing is left unanswered',
+ );
+ ok(
+ $monitor->{manual_promotion}->{supported},
+ 'so a client key can still be staged with a monitor outside the cluster',
);
for my $version (1, 2) {
@@ -7837,6 +7850,47 @@ for my $case ([0, 0], [1, 0], [0, 1], [1, 1]) {
);
}
+# Whatever a monitor is asked about itself, the route depends on whether this cluster has a shell
+# on its host. 'ceph tell' reaches the same admin socket over the network for one that it does not.
+{
+ no warnings qw(once redefine);
+ local *PVE::SSHInfo::get_ssh_info = sub { return { node => $_[0] } };
+ local *PVE::SSHInfo::ssh_info_to_command = sub { return ['ssh', $_[0]->{node}, '--'] };
+
+ my @commands;
+ local *main::run_command = sub {
+ my ($cmd, %args) = @_;
+ push @commands, join(' ', @$cmd);
+ $args{outfunc}->('answer');
+ };
+
+ my $run = sub {
+ @commands = ();
+ return $HOOKS->{mon_admin_run}->(@_);
+ };
+
+ is($run->('node-a', 'a', ['sessions']), "answer\n", 'a member answers over its admin socket');
+ is($commands[0], 'ssh node-a -- ceph daemon mon.a sessions', 'which is asked through ssh');
+
+ is($run->('witness.invalid', 'tiebreaker', ['sessions']), "answer\n", 'so does an outsider');
+ like(
+ $commands[0],
+ qr/^ceph --cluster \S+ --name mon\. --keyring \S+ --connect-timeout 10 tell mon\.tiebreaker sessions$/,
+ 'asked locally with tell, as mon., and with a connect timeout instead of a long default',
+ );
+
+ $run->(undef, 'nameless', ['version']);
+ like($commands[0], qr/tell mon\.nameless version/, 'a monitor without a known host too');
+
+ my @routed;
+ $run->('node-a', 'a', ['config', 'get', 'opt'], sub { push @routed, $_[1]; return '{}' });
+ is_deeply(
+ \@routed,
+ [['ceph', 'daemon', 'mon.a', 'config', 'get', 'opt']],
+ 'a caller-supplied runner keeps the member route, which the tests rely on',
+ );
+}
+
{
my $out = '';
{
--
2.47.3
^ permalink raw reply related [flat|nested] 5+ messages in thread
* [PATCH manager 2/4] migrations: cephx: judge cipher support by the version a daemon runs
2026-09-27 0:59 [PATCH manager 0/4] cephx: migrate keys with a monitor this cluster does not manage Kefu Chai
2026-09-27 0:59 ` [PATCH manager 1/4] migrations: cephx: ask a monitor about itself through one route Kefu Chai
@ 2026-09-27 0:59 ` Kefu Chai
2026-09-27 0:59 ` [PATCH manager 3/4] migrations: cephx: rotate the monitor key with monitors we do not manage Kefu Chai
2026-09-27 0:59 ` [PATCH manager 4/4] migrations: cephx: name the hosts keeping key copies we cannot write Kefu Chai
3 siblings, 0 replies; 5+ messages in thread
From: Kefu Chai @ 2026-09-27 0:59 UTC (permalink / raw)
To: pve-devel; +Cc: Thomas Lamprecht
A daemon's running version and its installed version answer different
questions: the running version decides whether it can hold a key in the
new cipher now, and Ceph reports that for every daemon; the installed
version decides whether it still could after a restart, and only
pvestatd, on a node of this cluster, reports that.
The check demanded both, so a daemon on a host this cluster does not
manage had no installed version, and the run stopped unconditionally
with "Could not verify 'aes256k' support for every service daemon".
That hit every daemon whenever a run still had to switch the service
cipher, which is every run before the switch unless '--only' narrows it,
so a stretch cluster with its tiebreaker outside this cluster could not
migrate anything.
Judge such a daemon by the version it runs instead, and note that only
the running version was checked, so an operator knows to keep Ceph there
at that version or newer. The run still stops if the daemon is down,
since nothing then reports a version, or if the running version lacks
cipher support.
Signed-off-by: Kefu Chai <k.chai@proxmox.com>
---
bin/pve-cephx-rotate-service-keys | 44 +++++++++++++++----
test/CephKeyMigrationScript_test.pl | 68 +++++++++++++++++++++++++++++
2 files changed, 103 insertions(+), 9 deletions(-)
diff --git a/bin/pve-cephx-rotate-service-keys b/bin/pve-cephx-rotate-service-keys
index 95d10266..28161306 100755
--- a/bin/pve-cephx-rotate-service-keys
+++ b/bin/pve-cephx-rotate-service-keys
@@ -2554,31 +2554,50 @@ my sub preflight_nodes($info, $plan, $opts) {
(@touched, grep { defined($_->{node}) && $lockbox_nodes->{ $_->{node} } } @all);
}
- my (@outdated, @unknown_versions);
+ # Two versions answer two questions. What a daemon runs decides whether it can hold a key in
+ # the new cipher now, and Ceph reports that for every daemon of the cluster. What is installed
+ # on its node decides whether it still could after a restart, and only pvestatd on a node of
+ # this cluster publishes that. A daemon this cluster does not manage is judged by the first
+ # alone, which is sound while it runs and nothing at all once it stops.
+ my (@outdated, @unknown_versions, @running_only);
for my $daemon (@judged) {
- if (!eval { assert_member($daemon->{node}, "$daemon->{type}.$daemon->{id}"); 1 }) {
- push @unknown_versions, $@;
+ my $label = "$daemon->{type}.$daemon->{id}";
+ my $stopped = ($daemon->{recovered} || $daemon->{down}) ? 1 : 0;
+
+ if (!defined($daemon->{node})) {
+ push @unknown_versions, "Cannot manage '$label': no valid host name was reported";
+ next;
+ }
+ if (!$daemon->{managed} && $stopped) {
+ push @unknown_versions,
+ "cannot verify the Ceph version of $label: it does not run, and its host"
+ . " '$daemon->{node}' is not a node of this Proxmox VE cluster";
next;
}
- if (!$daemon->{recovered} && !$daemon->{down}) {
+
+ if (!$stopped) {
if (!defined($daemon->{version}) || $daemon->{version} eq '') {
push @unknown_versions,
- "could not verify the running Ceph version of $daemon->{entity} on node"
+ "could not verify the running Ceph version of $label on node"
. " '$daemon->{node}'";
} elsif (!version_has_cipher($daemon->{version})) {
push @outdated,
- "$daemon->{entity} on node '$daemon->{node}' runs "
- . short_version($daemon->{version});
+ "$label on node '$daemon->{node}' runs " . short_version($daemon->{version});
+ } elsif (!$daemon->{managed}) {
+ push @running_only,
+ "$label on '$daemon->{node}' runs " . short_version($daemon->{version});
}
}
+ next if !$daemon->{managed};
+
if (!defined($daemon->{binary}) || $daemon->{binary} eq '') {
push @unknown_versions,
- "could not verify the installed Ceph version of $daemon->{entity} on node"
+ "could not verify the installed Ceph version of $label on node"
. " '$daemon->{node}'; check that 'pvestatd' runs there";
} elsif (!version_has_cipher($daemon->{binary})) {
push @outdated,
- "$daemon->{entity} on node '$daemon->{node}' would restart into "
+ "$label on node '$daemon->{node}' would restart into "
. short_version($daemon->{binary});
}
}
@@ -2645,6 +2664,13 @@ my sub preflight_nodes($info, $plan, $opts) {
}
}
+ if (@running_only) {
+ log_warn("These daemons run on hosts this cluster does not manage, so only the version they"
+ . " run was checked, not the one installed there. Keep Ceph on those hosts at this"
+ . " version or newer, or they lose the new cipher on their next restart:");
+ log_steps(\@running_only);
+ }
+
if (@unknown_versions) {
log_fail("Could not verify '$CIPHER' support for every service daemon:");
log_steps(\@unknown_versions);
diff --git a/test/CephKeyMigrationScript_test.pl b/test/CephKeyMigrationScript_test.pl
index fa1cee4e..a0b2fe95 100755
--- a/test/CephKeyMigrationScript_test.pl
+++ b/test/CephKeyMigrationScript_test.pl
@@ -6166,6 +6166,7 @@ for my $case ([0, 0], [1, 0], [0, 1], [1, 1]) {
id => "$_",
entity => "mon.$_",
node => "node$_",
+ managed => 1,
store => 'file',
version => '20.2.4-pve4',
binary => '20.2.4-pve4',
@@ -8026,4 +8027,71 @@ for my $case ([0, 0], [1, 0], [0, 1], [1, 1]) {
);
}
+# only cluster nodes publish their installed Ceph version, so a daemon on a host outside the
+# cluster can only be judged by the version it runs
+{
+ my $info = sub {
+ my (%tiebreaker) = @_;
+ my $mons = [
+ (
+ map { {
+ type => 'mon',
+ id => "$_",
+ entity => "mon.$_",
+ node => "node$_",
+ managed => 1,
+ store => 'file',
+ version => '20.2.4-pve4',
+ binary => '20.2.4-pve4',
+ } } 1 .. 2
+ ),
+ {
+ type => 'mon',
+ id => 'tiebreaker',
+ # every monitor shares the one 'mon.' key, and that is the entity the inventory
+ # records for each of them
+ entity => 'mon.',
+ node => 'tiebreaker.invalid',
+ managed => 0,
+ store => 'file',
+ version => '20.2.4-pve4',
+ %tiebreaker,
+ },
+ ];
+ return {
+ daemons => { mon => $mons, mgr => [], mds => [], osd => [] },
+ monmap_mons => [map { $_->{id} } @$mons],
+ rados => MonCountRados->new(),
+ };
+ };
+ my $preflight = sub {
+ my ($info) = @_;
+ my ($output, $verdict) = ('');
+ {
+ local *STDOUT;
+ open(STDOUT, '>', \$output) or die $!;
+ $verdict = $HOOKS->{preflight_nodes}
+ ->($info, { daemons => [], lockbox_keys => [], service_cipher => 1 }, {});
+ }
+ return ($verdict, $output // '');
+ };
+
+ my ($verdict, $output) = $preflight->($info->());
+ is($verdict, 1, 'a capable monitor outside the cluster does not block the cipher switch');
+ like(
+ $output,
+ qr/does not manage.*mon\.tiebreaker on 'tiebreaker\.invalid' runs 20\.2\.4/s,
+ 'naming it with the version it was judged by',
+ );
+
+ ($verdict, $output) = $preflight->($info->(version => '19.2.5'));
+ is($verdict, -1, 'an outdated monitor outside the cluster still blocks');
+ like($output, qr/mon\.tiebreaker on node 'tiebreaker\.invalid' runs 19\.2\.5/, 'as outdated');
+
+ ($verdict, $output) = $preflight->($info->(down => 1));
+ is($verdict, -1, 'a stopped monitor outside the cluster cannot be judged');
+ like($output, qr/mon\.tiebreaker.*not a node/, 'which names the membership');
+
+}
+
done_testing();
--
2.47.3
^ permalink raw reply related [flat|nested] 5+ messages in thread
* [PATCH manager 3/4] migrations: cephx: rotate the monitor key with monitors we do not manage
2026-09-27 0:59 [PATCH manager 0/4] cephx: migrate keys with a monitor this cluster does not manage Kefu Chai
2026-09-27 0:59 ` [PATCH manager 1/4] migrations: cephx: ask a monitor about itself through one route Kefu Chai
2026-09-27 0:59 ` [PATCH manager 2/4] migrations: cephx: judge cipher support by the version a daemon runs Kefu Chai
@ 2026-09-27 0:59 ` Kefu Chai
2026-09-27 0:59 ` [PATCH manager 4/4] migrations: cephx: name the hosts keeping key copies we cannot write Kefu Chai
3 siblings, 0 replies; 5+ messages in thread
From: Kefu Chai @ 2026-09-27 0:59 UTC (permalink / raw)
To: pve-devel; +Cc: Thomas Lamprecht
Rotating the shared 'mon.' key writes a new keyring to every monitor
and restarts them one at a time, both of which need a shell on the
monitor's host. A monitor this cluster does not manage stopped the run
with the generic "Cannot manage ... which is not a node of this Proxmox
VE cluster", so the key could never be rotated.
The key itself needs no shell. 'auth rotate' writes it to the auth
database, paxos carries it to every monitor in the quorum, and each
monitor reads its 'mon.' secret from there, both while running and
after a restart. So an unmanaged monitor migrates along with the rest.
What stays behind is the keyring in its data directory, the emergency
key Ceph falls back on if a monitor missed the update.
Split the rotation by what this cluster can reach: rotate the key,
update the keyring Proxmox VE keeps for the next monitor, then write
and restart the monitors it manages. For the rest, name the monitors
whose emergency keyring now holds the old key, print the refresh
command for their admin, and confirm they are still in the quorum.
The preflight follows the same split: neither the node probe nor the
check that a key lands where a daemon reads it applies to such a
monitor, since there is nothing of ours in its data directory, though
both still apply to every other daemon type. The plan still counts the
daemon as changed, so its version is judged like any other monitor's.
Signed-off-by: Kefu Chai <k.chai@proxmox.com>
---
bin/pve-cephx-rotate-service-keys | 82 ++++++++++++++--
test/CephKeyMigrationScript_test.pl | 145 +++++++++++++++++++++++++++-
test/CephKeyMigration_test.pl | 17 ++++
3 files changed, 235 insertions(+), 9 deletions(-)
diff --git a/bin/pve-cephx-rotate-service-keys b/bin/pve-cephx-rotate-service-keys
index 28161306..0b8a2c0c 100755
--- a/bin/pve-cephx-rotate-service-keys
+++ b/bin/pve-cephx-rotate-service-keys
@@ -191,6 +191,14 @@ my sub daemon_location($type, $id, $host) {
};
}
+my sub mon_keyring_path($id) {
+ return "/var/lib/ceph/mon/$ccname-$id/keyring";
+}
+
+my sub unmanaged_monitors($info) {
+ return grep { !$_->{managed} } $info->{daemons}->{mon}->@*;
+}
+
my sub assert_member($node, $entity) {
die "Cannot manage '$entity': no valid host name was reported\n"
if !defined($node) || ref($node) || !length($node);
@@ -1594,7 +1602,12 @@ my sub probe_nodes($info, $plan, $opts = {}) {
my $installed = installed_versions();
my $specs = {};
for my $daemon (touched_daemons($info, $plan)) {
- assert_member($daemon->{node}, "$daemon->{type}.$daemon->{id}");
+ my $label = "$daemon->{type}.$daemon->{id}";
+ die "Cannot manage '$label': no valid host name was reported\n"
+ if !defined($daemon->{node});
+ # a host this cluster does not manage holds no keyring to read, and the daemon there is
+ # judged by the version Ceph reports it running instead
+ next if !$daemon->{managed};
push $specs->{ $daemon->{node} }->@*, "$daemon->{type}:$daemon->{id}";
}
@@ -1617,6 +1630,7 @@ my sub probe_nodes($info, $plan, $opts = {}) {
for $info->{daemons}->{$type}->@*;
}
for my $daemon (touched_daemons($info, $plan)) {
+ next if !$daemon->{managed};
my $probe = $by_node->{ $daemon->{node} }->{"$daemon->{type}:$daemon->{id}"} // {};
$daemon->{store} = $probe->{store};
$daemon->{error} = $probe->{error};
@@ -2604,6 +2618,11 @@ my sub preflight_nodes($info, $plan, $opts) {
my @unusable;
for my $daemon (@touched) {
+ # the shared 'mon.' key reaches a monitor through the auth database, so a monitor we do not
+ # manage needs nothing written where it runs. Every other daemon reads its key from a file
+ # there, which is why only the monitors are exempt
+ next if !$daemon->{managed} && $daemon->{type} eq 'mon';
+
eval { assert_daemon_location($info->{rados}, $daemon, $info->{fsid}, $daemon) };
push @unusable, $@ if $@;
my $type = $daemon->{type};
@@ -2899,14 +2918,25 @@ my sub print_plan($info, $plan, $state, $opts, $storage_entities) {
$step++;
log_text("Step $step: rotate the shared 'mon.' key. Every keyring is written first, then"
. " the monitors restart one at a time.");
+ my @unmanaged = unmanaged_monitors($info);
+ log_step("Not managed by this cluster, so their emergency keyring keeps the old key: "
+ . join(', ', map { "mon.$_->{id} on '$_->{node}'" } @unmanaged)
+ . ". The commands for their admin follow at the end of the run.")
+ if @unmanaged;
if ($opts->{verbose}) {
log_step(
$info->{mon_key_in_auth_db}
? "Ceph lists the key in its health checks."
: "Ceph's health checks cannot see the key."
);
- log_step("monitors, restarted one at a time: "
- . join(', ', map { "$_->{id} (node $_->{node})" } $info->{daemons}->{mon}->@*));
+ log_step(
+ "monitors, restarted one at a time: "
+ . join(
+ ', ',
+ map { "$_->{id} (node $_->{node})" }
+ grep { $_->{managed} } $info->{daemons}->{mon}->@*,
+ ),
+ );
}
}
@@ -3259,6 +3289,39 @@ my sub merge_pve_mon_keyring($entry) {
return;
}
+# A monitor in the quorum takes the rotated key from the auth database while it runs, and reads it
+# from its own store after a restart, so nothing here is pending work. The keyring in its data
+# directory is the emergency copy Ceph falls back on when a monitor missed the update, and that
+# file sits on a host only its own admin reaches.
+my sub report_unmanaged_monitors($rados, $unmanaged) {
+ log_warn("These monitors run on hosts outside this Proxmox VE cluster. They use the new key"
+ . " already, while their emergency keyring still holds the old one. To bring that copy up"
+ . " to date, their admin can run:");
+ for my $mon (@$unmanaged) {
+ my $path = mon_keyring_path($mon->{id});
+ log_step("on '$mon->{node}', as root:");
+ log_step(" ceph auth get mon. -o $path");
+ log_step(" chown ceph:ceph $path && chmod 600 $path");
+ log_step(" systemctl restart ceph-mon\@$mon->{id}");
+ }
+ log_text("The first command needs an admin keyring on that host. Without one, run it on a node"
+ . " of this cluster and copy the file over. Nothing is broken until then, but such a"
+ . " monitor cannot authenticate from its keyring alone, which is what it falls back on"
+ . " after losing its store.");
+
+ my @lost = map { "mon.$_->{id} on '$_->{node}'" }
+ grep { !daemon_is_running($rados, 'mon', $_->{id}) } @$unmanaged;
+ if (@lost) {
+ log_warn("These are not in the quorum right now, so check them before anything else: "
+ . join(', ', @lost));
+ return;
+ }
+
+ log_pass("every monitor outside this cluster is still in the quorum");
+
+ return;
+}
+
my sub migrate_mon_key($rados, $state, $info, $opts, $plan) {
log_heading(
$plan->{mon_repair_only}
@@ -3269,8 +3332,9 @@ my sub migrate_mon_key($rados, $state, $info, $opts, $plan) {
# a stale-copy repair is not gated behind the opt-in, so it must not rotate and restart the
# quorum unasked. Finishing a started rotation is the exception
my $rotate = !$plan->{mon_repair_only};
- assert_daemon_location($rados, $_, $info->{fsid})
- for $rotate ? $info->{daemons}->{mon}->@* : ();
+ my @managed = grep { $_->{managed} } $info->{daemons}->{mon}->@*;
+ my @unmanaged = unmanaged_monitors($info);
+ assert_daemon_location($rados, $_, $info->{fsid}) for $rotate ? @managed : ();
my $entry = $rotate ? rotate_entity($rados, $state, 'mon.') : auth_entry($rados, 'mon.');
my $keyring = keyring_text($entry);
my $target = key_fingerprint($entry->{key});
@@ -3278,11 +3342,11 @@ my sub migrate_mon_key($rados, $state, $info, $opts, $plan) {
merge_pve_mon_keyring($entry) if ($info->{pve_mon_key} // '') ne $entry->{key};
# all keyrings first, so a monitor going down in between still finds the new key locally
- for my $mon ($info->{daemons}->{mon}->@*) {
+ for my $mon (@managed) {
next if $plan->{mon_repair_only};
next if ($state->{mon_keyring}->{ $mon->{id} } // '') eq $target;
- my $path = "/var/lib/ceph/mon/$ccname-$mon->{id}/keyring";
+ my $path = mon_keyring_path($mon->{id});
log_info("writing the new key to '$path' on node '$mon->{node}'");
write_node_file($mon->{node}, $path, $keyring);
@@ -3291,7 +3355,7 @@ my sub migrate_mon_key($rados, $state, $info, $opts, $plan) {
}
# only monitors holding the superseded key restart, so a keyring repair leaves the quorum alone
- for my $mon ($info->{daemons}->{mon}->@*) {
+ for my $mon (@managed) {
next if $plan->{mon_repair_only};
next if ($state->{mon_restarted}->{ $mon->{id} } // '') eq $target;
@@ -3327,6 +3391,8 @@ my sub migrate_mon_key($rados, $state, $info, $opts, $plan) {
log_pass("the shared monitor key now uses the '$CIPHER' cipher");
+ report_unmanaged_monitors($rados, \@unmanaged) if @unmanaged;
+
return;
}
diff --git a/test/CephKeyMigrationScript_test.pl b/test/CephKeyMigrationScript_test.pl
index a0b2fe95..62ccbdbd 100755
--- a/test/CephKeyMigrationScript_test.pl
+++ b/test/CephKeyMigrationScript_test.pl
@@ -6188,6 +6188,50 @@ for my $case ([0, 0], [1, 0], [0, 1], [1, 1]) {
return ($verdict, $output // '');
};
+ # the shared key reaches a monitor we do not manage through the auth database, so the run must
+ # not demand a keyring where that monitor runs
+ {
+ my $mons = [
+ (
+ map { {
+ type => 'mon',
+ id => "$_",
+ entity => 'mon.',
+ node => "node$_",
+ managed => 1,
+ store => 'file',
+ version => '20.2.4',
+ binary => '20.2.4',
+ } } 1 .. 2
+ ),
+ {
+ type => 'mon',
+ id => 'tiebreaker',
+ entity => 'mon.',
+ node => 'witness.example',
+ managed => 0,
+ version => '20.2.4',
+ },
+ ];
+ my $info = {
+ daemons => { mon => $mons, mgr => [], mds => [], osd => [] },
+ monmap_mons => [map { $_->{id} } @$mons],
+ fsid => 'cluster-fsid',
+ rados => MonCountRados->new(),
+ };
+ my ($output, $verdict) = ('');
+ {
+ local *STDOUT;
+ open(STDOUT, '>', \$output) or die $!;
+ $verdict = $HOOKS->{preflight_nodes}->(
+ $info, { daemons => [], lockbox_keys => [], mon_key => 1 }, {},
+ );
+ }
+ is($verdict, 1, 'a monitor we do not manage does not block the monitor key rotation');
+ unlike($output, qr/neither a keyring file nor a bluestore device/, 'nothing is read there');
+ unlike($output, qr/Cannot manage 'mon\.tiebreaker'/, 'and nothing is verified there');
+ }
+
my ($verdict, $output) = $preflight->(2, { mon_key => 1 });
is($verdict, -1, 'a two-monitor cluster refuses the shared monitor key rotation');
like($output, qr/monitor map holds 2 monitors/, 'naming what it found');
@@ -7822,7 +7866,10 @@ for my $case ([0, 0], [1, 0], [0, 1], [1, 1]) {
fsid => 'cluster-fsid',
pve_mon_key => $OLD,
exported => { 'mon.' => { key => $OLD } },
- daemons => { mon => [map { { type => 'mon', id => $_, node => "node-$_" } } qw(a b)] },
+ daemons => {
+ mon =>
+ [map { { type => 'mon', id => $_, node => "node-$_", managed => 1 } } qw(a b)],
+ },
};
my $plan = { mon_key => 1 };
my $state = {};
@@ -7892,6 +7939,102 @@ for my $case ([0, 0], [1, 0], [0, 1], [1, 1]) {
);
}
+# A monitor on a host outside this cluster, the tiebreaker of a stretch cluster for example,
+# takes the rotated key from the auth database. Only its local keyring copy and its restart are
+# left, and both need a shell the cluster does not have.
+{
+ no warnings qw(once redefine);
+ local *PVE::Ceph::Services::get_blocking_health_errors = sub { return [] };
+ local *PVE::Ceph::Services::wait_for_safe_to_stop = sub { return (1, '') };
+ local *PVE::Ceph::Services::wait_for_daemon_up = sub { };
+ local *PVE::SSHInfo::get_ssh_info = sub { return { node => $_[0] } };
+ local *PVE::SSHInfo::ssh_info_to_command = sub { return ['ssh', $_[0]->{node}, '--'] };
+ local *main::file_set_contents = sub { };
+
+ my @commands;
+ local *main::run_command = sub { push @commands, join(' ', $_[0]->@*) };
+
+ my $run = sub {
+ my ($quorum) = @_;
+ my $rados = ClientRotationRados->new($OLD);
+ my $mon_command = ClientRotationRados->can('mon_command');
+ local *ClientRotationRados::mon_command = sub {
+ my ($self, $args) = @_;
+ die "the cluster cannot be read\n"
+ if !defined($quorum) && $args->{prefix} =~ m/^(?:quorum_status|health)$/;
+ return $quorum if $args->{prefix} eq 'quorum_status';
+ return {} if $args->{prefix} eq 'health';
+ return $mon_command->(@_);
+ };
+ my $info = {
+ fsid => 'cluster-fsid',
+ pve_mon_key => $OLD,
+ exported => { 'mon.' => { key => $OLD } },
+ daemons => {
+ mon => [
+ (
+ map { { type => 'mon', id => $_, node => "node-$_", managed => 1 } }
+ qw(a b)
+ ),
+ {
+ type => 'mon',
+ id => 'tiebreaker',
+ node => 'witness.example',
+ managed => 0,
+ },
+ ],
+ },
+ };
+ my $state = {};
+ my $out = '';
+ {
+ open(my $stdout, '>', \$out) or die $!;
+ local *STDOUT = $stdout;
+ eval {
+ $HOOKS->{migrate_mon_key}
+ ->($rados, $state, $info, { timeout => 30 }, { mon_key => 1 });
+ };
+ main::log_fail($@) if $@;
+ }
+ return ($state, $out, $rados);
+ };
+
+ @commands = ();
+ my ($state, $out, $rados) = $run->({ quorum_names => [qw(a b tiebreaker)] });
+ is($rados->{key}, $NEW, 'the shared key is rotated, not refused over the outside monitor');
+ is_deeply(
+ $state->{mon_restarted},
+ { map { $_ => key_fingerprint($NEW) } qw(a b) },
+ 'the monitors this cluster manages are restarted',
+ );
+ is(
+ scalar(grep { /ssh witness\.example/ } @commands),
+ 0,
+ 'and nothing is sent to the host outside it',
+ );
+ like(
+ $out,
+ qr{ceph auth get mon\. -o /var/lib/ceph/mon/\S+-tiebreaker/keyring},
+ 'the admin is handed the keyring command for that monitor',
+ );
+ like($out, qr/systemctl restart ceph-mon\@tiebreaker/, 'and the restart to run there');
+ like($out, qr/chmod 600/, 'with the mode that keeps the key private');
+ like($out, qr/needs an admin keyring on that host/, 'and what to do without admin access');
+ like($out, qr/PASS.*outside this cluster is still in the quorum/, 'the quorum is confirmed');
+
+ (undef, $out) = $run->({ quorum_names => [qw(a b)] });
+ like(
+ $out,
+ qr/WARN.*not in the quorum right now.*mon\.tiebreaker on 'witness\.example'/,
+ 'a monitor that left the quorum is named instead',
+ );
+
+ # a cluster that cannot be read at all is not evidence against that monitor, and the run has
+ # already done its part by then
+ (undef, $out) = $run->(undef);
+ unlike($out, qr/not in the quorum right now/, 'an unreadable cluster accuses nobody');
+}
+
{
my $out = '';
{
diff --git a/test/CephKeyMigration_test.pl b/test/CephKeyMigration_test.pl
index 79fccedf..bac5b26e 100755
--- a/test/CephKeyMigration_test.pl
+++ b/test/CephKeyMigration_test.pl
@@ -566,6 +566,23 @@ is(
1,
'repairing only the stored copy writes to none of them, so none is validated either',
);
+
+ # its key is rotated under it, so its version has to be judged like any other monitor's; only
+ # the node-side work is skipped, and that is the prober's business
+ my $stretch = {
+ daemons => {
+ mon => [
+ { entity => 'mon.', id => 'a', managed => 1 },
+ { entity => 'mon.', id => 'b', managed => 1 },
+ { entity => 'mon.', id => 'tiebreaker', managed => 0 },
+ ],
+ },
+ };
+ is(
+ scalar(touched_daemons($stretch, { daemons => [$osd], mon_key => 1 })),
+ 4,
+ 'a monitor outside the cluster counts as changed as well',
+ );
}
# --- daemons ceph cannot see ------------------------------------------------------------------
--
2.47.3
^ permalink raw reply related [flat|nested] 5+ messages in thread
* [PATCH manager 4/4] migrations: cephx: name the hosts keeping key copies we cannot write
2026-09-27 0:59 [PATCH manager 0/4] cephx: migrate keys with a monitor this cluster does not manage Kefu Chai
` (2 preceding siblings ...)
2026-09-27 0:59 ` [PATCH manager 3/4] migrations: cephx: rotate the monitor key with monitors we do not manage Kefu Chai
@ 2026-09-27 0:59 ` Kefu Chai
3 siblings, 0 replies; 5+ messages in thread
From: Kefu Chai @ 2026-09-27 0:59 UTC (permalink / raw)
To: pve-devel; +Cc: Thomas Lamprecht
Proxmox VE rewrites the copies of a client key that it manages. A host
running a Ceph daemon without being a node of this cluster keeps its own
copy, typically the 'client.admin' keyring it needs for Ceph commands,
and nothing of ours writes it. Rotating the key leaves that copy stale,
and the first sign is a failing command on that host.
Name those hosts at the two moments that matter, from one place, so the
advice stays consistent. While the key is staged, the stale copy still
works, so its admin can use it to fetch the new key. Once the rotation
is confirmed, that copy is gone, and the fix needs a credential that
still works there: the monitor's own keyring on a monitor host, as long
as it still holds the current 'mon.' key, otherwise the key has to be
fetched on a node of this cluster and copied over by hand.
Both commands name the keyring explicitly, since a configuration copied
from here looks for keys under '/etc/pve', a path a host outside this
cluster does not have.
A session-based check cannot stand in for this note: a keyring file
nobody is using at that moment opens no session, so the confirmation
gate has nothing to refuse.
Signed-off-by: Kefu Chai <k.chai@proxmox.com>
---
bin/pve-cephx-rotate-service-keys | 50 +++++++++++++++-
test/CephKeyMigrationScript_test.pl | 91 ++++++++++++++++++++++++++++-
2 files changed, 139 insertions(+), 2 deletions(-)
diff --git a/bin/pve-cephx-rotate-service-keys b/bin/pve-cephx-rotate-service-keys
index 0b8a2c0c..9d9f686e 100755
--- a/bin/pve-cephx-rotate-service-keys
+++ b/bin/pve-cephx-rotate-service-keys
@@ -199,6 +199,48 @@ my sub unmanaged_monitors($info) {
return grep { !$_->{managed} } $info->{daemons}->{mon}->@*;
}
+# the hosts this cluster does not manage, so nothing of ours writes there, not even a key copy.
+# '$DAEMON_TYPES' leaves the monitors out, and a monitor is the daemon such a host usually runs
+my sub unmanaged_hosts_of($info) {
+ my $hosts = {};
+ for my $type ('mon', @$DAEMON_TYPES) {
+ $hosts->{ $_->{node} } = 1
+ for grep { !$_->{managed} && defined($_->{node}) } $info->{daemons}->{$type}->@*;
+ }
+ return sort keys %$hosts;
+}
+
+# Such a host keeps its own copies of a client key, the 'client.admin' keyring it needs to run Ceph
+# commands there most of all, and nothing of ours writes them. While the key is staged the stale
+# copy still authenticates, so its admin can fetch the new key with it. Once the previous key is
+# gone that takes a credential which still works there, and on a monitor host that is its own
+# 'mon.' keyring.
+my sub unmanaged_copy_note($info, $committed = 0) {
+ my @hosts = unmanaged_hosts_of($info);
+ return if !@hosts;
+
+ my $where = join(', ', map { "'$_'" } @hosts);
+ if (!$committed) {
+ log_step("Proxmox VE rewrites the key copies it manages. These hosts run Ceph daemons"
+ . " without being nodes of this cluster, so a keyring they keep for their own use is"
+ . " not among them: $where. Their admin can refresh one while both keys still work."
+ . " The command names that copy, as a configuration copied from here looks for keys"
+ . " under '/etc/pve', which such a host does not have:");
+ log_step(" ceph -n USER -k <keyring file> auth get USER -o <keyring file>");
+ return;
+ }
+
+ log_warn("The previous key no longer works. A keyring these hosts keep for their own use was"
+ . " not rewritten, as they are not nodes of this cluster: $where. On a monitor host, and"
+ . " while its keyring still holds the current 'mon.' key, its admin can refresh one with"
+ . " the monitor's own credential:");
+ log_step(" ceph --name mon. --keyring /var/lib/ceph/mon/$ccname-<id>/keyring"
+ . " auth get USER -o <keyring file>");
+ log_step("Otherwise fetch the key on a node of this cluster and copy the file over.");
+
+ return;
+}
+
my sub assert_member($node, $entity) {
die "Cannot manage '$entity': no valid host name was reported\n"
if !defined($node) || ref($node) || !length($node);
@@ -2037,6 +2079,7 @@ my sub preflight_cluster(
};
my $confirmation_refused = 0;
+ my $committed_copies = 0;
my $confirm_all = $opts->{'confirm-all-clients-refreshed'};
my $confirmation_seen = {};
my @confirm =
@@ -2264,7 +2307,7 @@ my sub preflight_cluster(
}
# The record remains open if the commit fails, so the confirmation is retryable.
- commit_staged_key($info->{rados}, $state, $entity) if $staged;
+ $committed_copies = 1 if $staged && commit_staged_key($info->{rados}, $state, $entity);
my $mark = $state->{client_refresh}->{$entity};
$mark->{cleared} = time();
$mark->{acknowledged} = time();
@@ -2273,6 +2316,9 @@ my sub preflight_cluster(
}
}
+ # a copy on a host we do not manage is only stale from here on, so this is where to say it
+ unmanaged_copy_note($info, 1) if $committed_copies;
+
# Confirmation can promote staged keys. Recollect before any remaining preflight decision and
# update the caller's shared cluster picture before it builds the migration plan.
if (!$confirmation_refused && scalar(@confirm)) {
@@ -3055,6 +3101,7 @@ my sub print_plan($info, $plan, $state, $opts, $storage_entities) {
log_step("For each staged Ceph user key, both the current and new keys authenticate"
. " until the new key is committed with '--confirm-clients-refreshed USER' or,"
. " once every open record is ready, '--confirm-all-clients-refreshed'.");
+ unmanaged_copy_note($info);
log_step("the monitors stop automatic pending-key promotion until every staged key is"
. " committed or aborted")
if $opts->{verbose};
@@ -6356,6 +6403,7 @@ sub key_migration_test_hooks {
check_client_kernels => \&check_client_kernels,
manual_promotion_with_retries => \&manual_promotion_with_retries,
mon_admin_run => \&mon_admin_run,
+ unmanaged_copy_note => \&unmanaged_copy_note,
describe_live => \&describe_live,
};
}
diff --git a/test/CephKeyMigrationScript_test.pl b/test/CephKeyMigrationScript_test.pl
index 62ccbdbd..6520e7a7 100755
--- a/test/CephKeyMigrationScript_test.pl
+++ b/test/CephKeyMigrationScript_test.pl
@@ -3619,7 +3619,15 @@ sub run_aggregate_confirmation {
mon => [],
mgr => [],
mds => [],
- osd => [{ type => 'osd', id => 7, entity => 'osd.7', node => 'node-daemon' }],
+ osd => [
+ {
+ type => 'osd',
+ id => 7,
+ entity => 'osd.7',
+ node => 'node-daemon',
+ managed => 1,
+ },
+ ],
},
};
my $plan = {
@@ -3717,6 +3725,42 @@ sub run_aggregate_confirmation {
qr/For each staged Ceph user key, both the current and new keys authenticate.*committed with '--confirm-clients-refreshed USER' or, once every open record is ready, '--confirm-all-clients-refreshed'/s,
'the default names both staged completion paths without offering the aggregate early',
);
+ unlike(
+ $concise,
+ qr/without being nodes of this cluster/,
+ 'a cluster that manages every daemon host is told nothing about copies elsewhere',
+ );
+
+ # a host running a Ceph daemon without being a node of this cluster keeps its own key copies,
+ # and the staging window is when its admin can still fetch the new key with the old one
+ {
+ my $stretched = { %$info, daemons => { %{ $info->{daemons} } } };
+ $stretched->{daemons}->{mon} = [
+ { type => 'mon', id => 'a', node => 'node-a', managed => 1 },
+ { type => 'mon', id => 'tiebreaker', node => 'witness.example', managed => 0 },
+ ];
+ my $output = $render->(0, undef, undef, undef, $stretched);
+ like(
+ $output,
+ qr/without being nodes of this cluster.*'witness\.example'.*while both keys still work/s,
+ 'the host is named while the old key still authenticates',
+ );
+ like(
+ $output,
+ qr/ceph -n USER -k <keyring file> auth get USER -o <keyring file>/,
+ 'with a command that names the copy, as a config from here points into /etc/pve',
+ );
+ unlike(
+ $output,
+ qr/'node-a'/,
+ 'and a managed node is not, as Proxmox VE rewrites its copies',
+ );
+ unlike(
+ $output,
+ qr/--name mon\./,
+ 'the monitor credential is not offered while the simpler one still works',
+ );
+ }
like(
$concise,
qr/For staged keys.*live-migrate affected\s+VMs, including those with kernel RBD disks/s,
@@ -7898,6 +7942,51 @@ for my $case ([0, 0], [1, 0], [0, 1], [1, 1]) {
);
}
+# Once the previous key is gone the stale copy is no longer a credential, so the note has to name
+# one that still works on that host.
+{
+ my $stretched = {
+ daemons => {
+ mon => [
+ { type => 'mon', id => 'a', node => 'node-a', managed => 1 },
+ { type => 'mon', id => 'witness', node => 'witness.example', managed => 0 },
+ ],
+ mgr => [],
+ mds => [],
+ osd => [],
+ },
+ };
+ my $note = sub {
+ my ($committed) = @_;
+ my $out = '';
+ open(my $stdout, '>', \$out) or die $!;
+ local *STDOUT = $stdout;
+ $HOOKS->{unmanaged_copy_note}->($stretched, $committed);
+ return $out;
+ };
+
+ my $committed = $note->(1);
+ like($committed, qr/WARN.*previous key no longer works/, 'the note says the old key is gone');
+ like(
+ $committed,
+ qr{ceph --name mon\. --keyring /var/lib/ceph/mon/\S+-<id>/keyring auth get USER},
+ "and gives the monitor's own credential as the way back",
+ );
+ like($committed, qr/copy the file over/, 'with the fallback for a host running no monitor');
+ like($committed, qr/'witness\.example'/, 'naming the unmanaged host');
+ unlike($committed, qr/'node-a'/, 'and not the one Proxmox VE rewrites itself');
+ like($note->(0), qr/while both keys still work/, 'before the commit it is the simpler route');
+
+ my $managed = { daemons => { map { $_ => [] } qw(mon mgr mds osd) } };
+ my $out = '';
+ {
+ open(my $stdout, '>', \$out) or die $!;
+ local *STDOUT = $stdout;
+ $HOOKS->{unmanaged_copy_note}->($managed, 1);
+ }
+ is($out, '', 'a cluster that manages every daemon host hears nothing about copies elsewhere');
+}
+
# Whatever a monitor is asked about itself, the route depends on whether this cluster has a shell
# on its host. 'ceph tell' reaches the same admin socket over the network for one that it does not.
{
--
2.47.3
^ permalink raw reply related [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-09-27 1:00 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-27 0:59 [PATCH manager 0/4] cephx: migrate keys with a monitor this cluster does not manage Kefu Chai
2026-09-27 0:59 ` [PATCH manager 1/4] migrations: cephx: ask a monitor about itself through one route Kefu Chai
2026-09-27 0:59 ` [PATCH manager 2/4] migrations: cephx: judge cipher support by the version a daemon runs Kefu Chai
2026-09-27 0:59 ` [PATCH manager 3/4] migrations: cephx: rotate the monitor key with monitors we do not manage Kefu Chai
2026-09-27 0:59 ` [PATCH manager 4/4] migrations: cephx: name the hosts keeping key copies we cannot write Kefu Chai
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox