public inbox for pve-devel@lists.proxmox.com
 help / color / mirror / Atom feed
From: Joaquin Varela <joaquinvarela@neatech.ar>
To: pve-devel@lists.proxmox.com
Subject: [PATCH storage v3 1/4] zfsnvme: add ZFS over NVMe/TCP storage plugin
Date: Sun,  4 Oct 2026 21:26:04 -0300	[thread overview]
Message-ID: <20261005002609.571-2-joaquinvarela@neatech.ar> (raw)
In-Reply-To: <20261005002609.571-1-joaquinvarela@neatech.ar>

Add the shared storage type 'zfsnvme': ZFS zvols on a remote Linux
target, exported with the kernel NVMe target (nvmet, configured through
configfs) over NVMe/TCP, and consumed by the cluster nodes with native
NVMe multipath and DH-HMAC-CHAP authentication.

The ZFS dataset is the source of truth for a namespace: the subsystem
NQN, NSID and UUID of every volume are stored in the ZFS user properties
'proxmox:nvme-subsys', 'proxmox:nvme-nsid' and 'proxmox:nvme-uuid', and
the pool keeps the highest NSID handed out in 'proxmox:nvme-last-nsid'.
The nodes find a namespace by its UUID. Everything else on the target
(nvmet_tcp module, configfs mount, subsystem, namespaces, ports, host
ACLs) is derived state that an activation rebuilds after a target
restart.

All parsing, validation and planning happen locally. The target state
is read with one 'zfs get' and one 'find' over configfs; other reads
are 'zfs get' and 'zfs list', 'test -b' while a new zvol appears, and
'cat /proc/mounts' to decide whether configfs needs to be mounted. The
target runs quoted simple commands over SSH, joined with '&&' where a
step depends on the previous one: an export starts with 'test -b' on
its zvol, and each publish link repeats two read-only checks right
before it links the subsystem to a port. There are no scripts, loops
or variables on the target. As with ZFS over iSCSI, the SSH key is
/etc/pve/priv/zfs/<server>_id_rsa. Target changes are serialized
cluster-wide by a pmxcfs domain lock per target server, and an
operation whose SSH connection broke is reported as unknown and
repaired by the next activation rather than compensated from a read
that could overtake it.

Each storage needs its own NVMe/TCP listener (address family, address
and port) on the target. nvmet answers a connect to a listener that
does not publish the subsystem yet with DNR, and the host then deletes
the controller whatever its loss timeout, so after a target restart the
storage published second on a shared listener would lose its paths.
Adding, updating and activating a storage therefore refuse a portal
whose listener another zfsnvme storage uses, disabled ones included,
and a target address that another storage reaches through a different
'server' value, since that value names the target lock. A data address
thus identifies one target for all zfsnvme storages of the cluster.

The DH-HMAC-CHAP key is a sensitive property, stored like the other
storage secrets in /etc/pve/priv/storage/<storeid>.nvme-dhchap. It
reaches the target through stdin, and the plugin restricts the
dhchap_key and dhchap_ctrl_key attributes there to 0600 whenever it
writes the key. On the nodes, controllers are created with exactly one
write to /dev/nvme-fabrics in a bounded child (run_fork_with_timeout),
so the key never appears on a command line, and the secret attributes of
a new controller are restricted before the child returns. Rescans and
disconnects are sysfs writes; nvme-cli is not run.

Keep-alive, reconnect delay, controller loss and optional fast I/O
failure timeouts are configurable, with defaults that prefer waiting for
the target over failing guest I/O (ctrl_loss_tmo 600, fast_io_fail_tmo
unset).

Outside the plugin: register the type in PVE::Storage and the Makefile,
add it to @SHARED_STORAGE in PVE::Storage::Plugin, let 'pvesm add' and
'pvesm set' read --dhchap-key from a file, and depend on nvme-cli, which
creates /etc/nvme/hostnqn and /etc/nvme/hostid and remains the
administration tool, like the client packages of the other shared types.

Signed-off-by: Joaquin Varela <joaquinvarela@neatech.ar>
---
 debian/control                   |    1 +
 src/PVE/CLI/pvesm.pm             |   21 +-
 src/PVE/Storage.pm               |    2 +
 src/PVE/Storage/Makefile         |    1 +
 src/PVE/Storage/Plugin.pm        |    2 +-
 src/PVE/Storage/ZFSNVMePlugin.pm | 3177 ++++++++++++++++++++++++++++++
 6 files changed, 3201 insertions(+), 3 deletions(-)
 create mode 100644 src/PVE/Storage/ZFSNVMePlugin.pm

diff --git a/debian/control b/debian/control
index 850cd57c..ab46e6b6 100644
--- a/debian/control
+++ b/debian/control
@@ -43,6 +43,7 @@ Depends: bzip2,
          lvm2,
          lzop,
          nfs-common,
+         nvme-cli,
          proxmox-backup-client (>= 2.1.10~),
          proxmox-backup-file-restore,
          pve-cluster (>= 5.0-32),
diff --git a/src/PVE/CLI/pvesm.pm b/src/PVE/CLI/pvesm.pm
index 1aaa9f3c..3c68ffb5 100755
--- a/src/PVE/CLI/pvesm.pm
+++ b/src/PVE/CLI/pvesm.pm
@@ -78,12 +78,29 @@ sub param_mapping {
         },
     };
 
+    my $dhchap_key_map = {
+        name => 'dhchap-key',
+        desc => 'a file containing the NVMe DH-HMAC-CHAP key',
+        func => sub {
+            my ($value) = @_;
+            # never repeat a key passed in place of the file name
+            die "dhchap-key expects the path of a file containing the key\n"
+                if $value =~ /^DHHC-1:/i || -d $value;
+            my ($key) = split(/\n/, PVE::Tools::file_get_contents($value), 2);
+            $key = PVE::Tools::trim($key // '');
+            die "DH-HMAC-CHAP key file '$value' is empty\n" if $key eq '';
+            return $key;
+        },
+    };
+
     my $mapping = {
         'cifsscan' => [$password_map],
         'cifs' => [$password_map],
         'pbs' => [$password_map],
-        'create' => [$password_map, $enc_key_map, $master_key_map, $keyring_map],
-        'update' => [$password_map, $enc_key_map, $master_key_map, $keyring_map],
+        'create' =>
+            [$password_map, $enc_key_map, $master_key_map, $keyring_map, $dhchap_key_map],
+        'update' =>
+            [$password_map, $enc_key_map, $master_key_map, $keyring_map, $dhchap_key_map],
     };
     return $mapping->{$name};
 }
diff --git a/src/PVE/Storage.pm b/src/PVE/Storage.pm
index fc3db812..2e5a8031 100755
--- a/src/PVE/Storage.pm
+++ b/src/PVE/Storage.pm
@@ -36,6 +36,7 @@ use PVE::Storage::CephFSPlugin;
 use PVE::Storage::ISCSIDirectPlugin;
 use PVE::Storage::ZFSPoolPlugin;
 use PVE::Storage::ZFSPlugin;
+use PVE::Storage::ZFSNVMePlugin;
 use PVE::Storage::PBSPlugin;
 use PVE::Storage::BTRFSPlugin;
 use PVE::Storage::ESXiPlugin;
@@ -61,6 +62,7 @@ PVE::Storage::CephFSPlugin->register();
 PVE::Storage::ISCSIDirectPlugin->register();
 PVE::Storage::ZFSPoolPlugin->register();
 PVE::Storage::ZFSPlugin->register();
+PVE::Storage::ZFSNVMePlugin->register();
 PVE::Storage::PBSPlugin->register();
 PVE::Storage::BTRFSPlugin->register();
 PVE::Storage::ESXiPlugin->register();
diff --git a/src/PVE/Storage/Makefile b/src/PVE/Storage/Makefile
index a67dc25f..d1cbfe29 100644
--- a/src/PVE/Storage/Makefile
+++ b/src/PVE/Storage/Makefile
@@ -11,6 +11,7 @@ SOURCES= \
 	ISCSIDirectPlugin.pm \
 	ZFSPoolPlugin.pm \
 	ZFSPlugin.pm \
+	ZFSNVMePlugin.pm \
 	PBSPlugin.pm \
 	BTRFSPlugin.pm \
 	LvmThinPlugin.pm \
diff --git a/src/PVE/Storage/Plugin.pm b/src/PVE/Storage/Plugin.pm
index a9e17513..a5aedc2b 100644
--- a/src/PVE/Storage/Plugin.pm
+++ b/src/PVE/Storage/Plugin.pm
@@ -35,7 +35,7 @@ our @COMMON_TAR_FLAGS = qw(
 );
 
 our @SHARED_STORAGE = (
-    'iscsi', 'nfs', 'cifs', 'rbd', 'cephfs', 'iscsidirect', 'zfs', 'drbd', 'pbs',
+    'iscsi', 'nfs', 'cifs', 'rbd', 'cephfs', 'iscsidirect', 'zfs', 'zfsnvme', 'drbd', 'pbs',
 );
 
 our $QCOW2_PREALLOCATION = {
diff --git a/src/PVE/Storage/ZFSNVMePlugin.pm b/src/PVE/Storage/ZFSNVMePlugin.pm
new file mode 100644
index 00000000..127680fb
--- /dev/null
+++ b/src/PVE/Storage/ZFSNVMePlugin.pm
@@ -0,0 +1,3177 @@
+package PVE::Storage::ZFSNVMePlugin;
+
+use v5.36;
+
+use Compress::Zlib qw(crc32);
+use Digest::SHA qw(sha256_hex);
+use Errno qw(ENOENT);
+use Fcntl qw(O_NOFOLLOW O_RDONLY O_RDWR S_ISCHR);
+use File::Path qw(make_path);
+use List::Util qw(max);
+use MIME::Base64 qw(decode_base64);
+use POSIX qw(SIG_BLOCK SIG_SETMASK SIGHUP SIGINT SIGQUIT SIGTERM sigprocmask);
+use Socket qw(AF_INET6 inet_ntop inet_pton);
+use Time::HiRes;
+
+use PVE::Cluster;
+use PVE::File;
+use PVE::JSONSchema;
+use PVE::Network;
+use PVE::RESTEnvironment qw(log_warn);
+use PVE::RPCEnvironment;
+use PVE::SysFSTools;
+use PVE::Tools qw(file_read_firstline file_set_contents run_command trim);
+
+use base qw(PVE::Storage::Plugin);
+
+# ZFS zvols on a remote Linux target, exported through the kernel NVMe target
+# (nvmet, configured through configfs) over NVMe/TCP and consumed with native
+# NVMe multipath.
+#
+# The ZFS dataset is the source of truth for a namespace: its owner NQN, NSID
+# and UUID are ZFS user properties. Everything else on the target is derived
+# state that the plugin rebuilds after a target restart.
+#
+# Storage parsing, validation and planning happen locally. SSH carries quoted
+# commands, sometimes joined with `&&`, to read or change ZFS/configfs state.
+# A pmxcfs domain lock per target serializes the changes within the cluster,
+# held like the storage lock of the other shared storage types.
+
+# ---------------------------------------------------------------------------
+# Constants and regular expressions
+# ---------------------------------------------------------------------------
+
+# LogLevel=ERROR keeps a login banner out of the first stderr line, which is
+# the error text of a failed call.
+my @ssh_cmd = (
+    '/usr/bin/ssh', '-o', 'BatchMode=yes', '-o', 'ConnectTimeout=10', '-o', 'LogLevel=ERROR',
+);
+my $id_rsa_path = '/etc/pve/priv/zfs';
+my $secret_dir = '/etc/pve/priv/storage';
+my $max_paths = 16;
+my $max_hosts = 64;
+my $slow_path_backoff = 60;
+
+my $nvmet_root = '/sys/kernel/config/nvmet';
+my $nvmet_model = 'Proxmox ZFS NVMe';
+my $nvmet_marker = 'ZFSNVME-CONFIGFS';
+my $nvmet_max_nsid = 0xfffffffe;
+my $nvmet_max_command = 65536;
+# Seconds to wait for the target lock while another node changes the target;
+# a destroy or rollback can hold it for several seconds.
+my $nvmet_lock_wait = 30;
+# On the pool dataset: the highest NSID handed out for a volume of the pool.
+my $nvmet_last_nsid = 'proxmox:nvme-last-nsid';
+
+my @nvmet_zfs_props =
+    ('type', 'proxmox:nvme-subsys', 'proxmox:nvme-nsid', 'proxmox:nvme-uuid', $nvmet_last_nsid);
+my %nvmet_zfs_keys = (
+    type => 'type',
+    'proxmox:nvme-subsys' => 'nqn',
+    'proxmox:nvme-nsid' => 'nsid',
+    'proxmox:nvme-uuid' => 'uuid',
+    $nvmet_last_nsid => 'last_nsid',
+);
+my @nvmet_cfs_attrs = qw(
+    enable device_path device_uuid buffered_io
+    attr_model attr_serial attr_allow_any_host
+    addr_trtype addr_adrfam addr_traddr addr_trsvcid
+);
+# Every command the plugin runs on the target, besides the `dd | tee` pipe of
+# a key write. The configfs read runs find through env to set its locale.
+my %nvmet_commands = map { $_ => 1 } qw(
+    cat chmod env grep ln mkdir modprobe mount printf rm rmdir test zfs
+);
+
+my $RE_NQN = qr{
+    \A
+    nqn \.
+    [A-Za-z0-9] [A-Za-z0-9.-]*
+    :
+    [A-Za-z0-9] [A-Za-z0-9._:-]*
+    \z
+}nxx;
+my $RE_IPV4_PORTAL = qr{
+    \A
+    (?<address> [^:]+)
+    (?: : (?<port> [0-9]+))?
+    \z
+}nxx;
+my $RE_IPV6_PORTAL = qr{
+    \A
+    \[ (?<address> [^\]]+) \]
+    (?: : (?<port> [0-9]+))?
+    \z
+}nxx;
+# A canonical IPv4-mapped IPv6 address.
+my $RE_IPV4_MAPPED = qr{\A ::ffff: (?<address> [0-9.]+) \z}nxx;
+my $RE_HOST_IFACE = qr{\A [A-Za-z0-9_.-]+ \z}nxx;
+my $RE_DHCHAP_KEY = qr{
+    \A DHHC-1 : (?<hash> 0[0-3]) : (?<secret> [A-Za-z0-9+/]+ ={0,2}) : \z
+}nxx;
+# A stopped PVE task: the worker's signal handler dies with the first text,
+# run_command adds the command context, and run_fork_with_timeout reports
+# the signal in its own words.
+my $RE_TASK_INTERRUPT = qr{
+    \A (?: (?: command \x20 ' .* ' \x20 failed: \x20 )? received \x20 interrupt
+    | interrupted \x20 by \x20 unexpected \x20 signal ) \n \z
+}nsxx;
+my $RE_COMMAND_TIMEOUT = qr{
+    \A (?: command \x20 ' .* ' \x20 failed: \x20 )? got \x20 timeout \n \z
+}nsxx;
+my $RE_COMMAND_EXIT =
+    qr{\A command \x20 ' .* ' \x20 failed: \x20 exit \x20 code \x20 (?<code> [0-9]+) \n \z}nsxx;
+my $RE_FABRICS_VALUE = qr{\A [^\s,\x00]+ \z}nxx;
+my $RE_FABRICS_RESULT = qr{\A instance=(?<instance> [0-9]+) ,cntlid= [0-9]+ \n \z}nxx;
+my $RE_ZERO_UUID = qr{\A 0{8} (?: -0{4}){3} -0{12} \z}nxx;
+my $RE_CONFIG_INT = qr{\A -?[0-9]+ \z}nxx;
+my $RE_NVME_CONTROLLER = qr{\A nvme [0-9]+ \z}nxx;
+my $RE_NVME_SUBSYSTEM = qr{\A nvme-subsys [0-9]+ \z}nxx;
+my $RE_NVME_NAMESPACE = qr{\A nvme [0-9]+ n [0-9]+ \z}nxx;
+my $RE_NVME_PARTITION = qr{\A nvme [0-9]+ n [0-9]+ p [0-9]+ \z}nxx;
+my $RE_TRADDR = qr{(?: \A | ,) traddr=(?<value>[^,]+)}nxx;
+my $RE_TRSVCID = qr{(?: \A | ,) trsvcid=(?<value>[^,]+)}nxx;
+my $RE_HOST_IFACE_ADDRESS = qr{(?: \A | ,) host_iface=(?<value>[^,]+)}nxx;
+my $RE_ZVOL_OWNER = qr{^ (?:vm|base|subvol|basevol)- (?<owner>\d+) - \S+ $}nxx;
+my $RE_VOLNAME = qr{
+    ^
+    (?: (?<base> (?:base|basevol)- (?<base_vmid>\d+) - \S+) /)?
+    (?<name> (?<type>base|basevol|vm|subvol)- (?<vmid>\d+) - \S+)
+    $
+}nxx;
+my $RE_BASE_SNAPSHOT = qr{^ (?<base>\S+) \@__base__ $}nxx;
+my $RE_UNSIGNED_INTEGER = qr{^ (?<value>\d+) $}nxx;
+
+my $RE_NVMET_UUID = qr{\A [0-9a-fA-F]{8} (?: - [0-9a-fA-F]{4}){3} - [0-9a-fA-F]{12} \z}nxx;
+my $RE_NVMET_NSID = qr{\A [1-9] [0-9]* \z}nxx;
+my $RE_NVMET_POOL = qr{\A [A-Za-z0-9] [A-Za-z0-9_.:/-]* \z}nxx;
+my $RE_NVMET_DATASET_NAME = qr{\A [A-Za-z0-9] [A-Za-z0-9_.:-]* \z}nxx;
+my $RE_NVMET_SNAPSHOT = qr{\A [A-Za-z0-9_.:-]+ \z}nxx;
+my $RE_NVMET_TEMPLATE_ZVOL = qr{/ (?<type> vm | base) - (?<suffix> [^/]+) \z}nxx;
+my $RE_NVMET_INHERITED = qr{\A inherited \x20 from \x20 (?<source> .+) \z}nxx;
+# Errors after which the target state is not known: an interrupted task, or a
+# change whose command may still be running on the target.
+my $RE_NVMET_ABANDONED = qr{
+    (?: \A received \x20 interrupt | ; \x20 the \x20 target \x20 state \x20 is \x20 unknown) \n \z
+}nxx;
+
+# Lines of the configfs read (nvmet_configfs_step below): `D <dir>`,
+# `L <link>`, `<file>:<value>` from grep, and `<sha256>  <file>` from
+# sha256sum. Attribute names anchor the split, so an NQN containing ':' is
+# never ambiguous.
+my $cfs_root = quotemeta($nvmet_root);
+my $RE_CFS_TOP = qr{\A D \x20 $cfs_root (?: / (?<top> hosts | ports | subsystems))? \z}nxx;
+my $RE_CFS_SUBSYS_DIR = qr{\A D \x20 $cfs_root /subsystems/ (?<nqn> [^/]+) \z}nxx;
+my $RE_CFS_SUBSYS_GROUP = qr{
+    \A D \x20 $cfs_root /subsystems/ [^/]+ / (?: namespaces | allowed_hosts) \z
+}nxx;
+my $RE_CFS_NS_DIR = qr{
+    \A D \x20 $cfs_root /subsystems/ (?<nqn> [^/]+) /namespaces/ (?<nsid> [0-9]+) \z
+}nxx;
+my $RE_CFS_PORT_DIR = qr{\A D \x20 $cfs_root /ports/ (?<port> [0-9]+) \z}nxx;
+my $RE_CFS_PORT_GROUP = qr{\A D \x20 $cfs_root /ports/ [0-9]+ /subsystems \z}nxx;
+my $RE_CFS_HOST_DIR = qr{\A D \x20 $cfs_root /hosts/ (?<host> [^/]+) \z}nxx;
+my $RE_CFS_PORT_LINK = qr{
+    \A L \x20 $cfs_root /ports/ (?<port> [0-9]+) /subsystems/ (?<nqn> [^/]+) \z
+}nxx;
+my $RE_CFS_ACL_LINK = qr{
+    \A L \x20 $cfs_root /subsystems/ (?<nqn> [^/]+) /allowed_hosts/ (?<host> [^/]+) \z
+}nxx;
+my $RE_CFS_OTHER_ENTRY = qr{\A [DL] \x20 $cfs_root /}nxx;
+my $RE_CFS_NS_ATTR = qr{
+    \A $cfs_root /subsystems/ (?<nqn> [^/]+) /namespaces/ (?<nsid> [0-9]+)
+    / (?<attr> enable | device_path | device_uuid | buffered_io) : (?<value> .*) \z
+}nxx;
+my $RE_CFS_SUBSYS_ATTR = qr{
+    \A $cfs_root /subsystems/ (?<nqn> [^/]+)
+    / (?<attr> attr_model | attr_serial | attr_allow_any_host) : (?<value> .*) \z
+}nxx;
+my $RE_CFS_PORT_ATTR = qr{
+    \A $cfs_root /ports/ (?<port> [0-9]+)
+    / (?<attr> addr_trtype | addr_adrfam | addr_traddr | addr_trsvcid) : (?<value> .*) \z
+}nxx;
+my $RE_CFS_KEY_DIGEST = qr{
+    \A (?<sha> [0-9a-f]{64}) \x20 [\x20*] $cfs_root /hosts/ (?<host> [^/]+) /dhchap_key \z
+}nxx;
+my $RE_CFS_OTHER_ATTR = qr{\A $cfs_root /}nxx;
+
+# ---------------------------------------------------------------------------
+# Target addressing and validated names
+# ---------------------------------------------------------------------------
+
+my sub nvmet_server($scfg) {
+    return $scfg->{server};
+}
+
+my sub nvmet_ssh_key($scfg) {
+    return "$id_rsa_path/" . nvmet_server($scfg) . '_id_rsa';
+}
+
+my sub nvmet_lock_id($scfg) {
+    my $server = nvmet_server($scfg) // '';
+    return 'zfsnvme-' . ($server =~ s/[^A-Za-z0-9.-]/_/gr);
+}
+
+my sub secret_path($storeid) {
+    return "$secret_dir/$storeid.nvme-dhchap";
+}
+
+my sub nvmet_nqn($scfg) {
+    my $nqn = $scfg->{subsysnqn};
+    die "invalid NVMe subsystem NQN\n" if !defined($nqn) || !verify_nvme_nqn($nqn, 1);
+    return $nqn;
+}
+
+my sub nvmet_pool($scfg) {
+    my $pool = $scfg->{pool};
+    die "invalid ZFS pool name\n" if !defined($pool) || $pool !~ $RE_NVMET_POOL;
+    return $pool;
+}
+
+my sub nvmet_dataset($scfg, $name) {
+    die "invalid ZFS volume name\n" if !defined($name) || $name !~ $RE_NVMET_DATASET_NAME;
+    return nvmet_pool($scfg) . "/$name";
+}
+
+# zfs destroy reads '%' and ',' in a snapshot name as a range and a list.
+my sub nvmet_snapshot($scfg, $name, $snap) {
+    die "invalid snapshot name\n" if !defined($snap) || $snap !~ $RE_NVMET_SNAPSHOT;
+    return nvmet_dataset($scfg, $name) . "\@$snap";
+}
+
+my sub nvmet_valid_nsid($value) {
+    return defined($value) && $value =~ $RE_NVMET_NSID && $value <= $nvmet_max_nsid;
+}
+
+# IPv6 addresses have many spellings; compare and store only the canonical one.
+my sub nvmet_canonical_address($address) {
+    return $address if !defined($address) || index($address, ':') < 0;
+    my $packed = inet_pton(AF_INET6, $address) // return $address;
+    return inet_ntop(AF_INET6, $packed);
+}
+
+# The address family and address that a portal of parse_nvme_portals reaches:
+# an IPv4-mapped IPv6 address reaches the IPv4 address.
+my sub nvmet_listener_address($portal) {
+    return ('ipv4', $+{address})
+        if $portal->{family} eq 'ipv6' && $portal->{address} =~ $RE_IPV4_MAPPED;
+    return $portal->@{qw(family address)};
+}
+
+# Target-provided names only appear in messages when they are well-formed.
+my sub nvmet_label($name) {
+    return $name =~ $RE_NQN ? " '$name'" : '';
+}
+
+# The other name of a zvol that a template conversion renames: vm-* and base-*.
+my sub nvmet_template_twin($name) {
+    return '' if $name !~ $RE_NVMET_TEMPLATE_ZVOL;
+    my $twin = ($+{type} eq 'vm' ? 'base' : 'vm') . "-$+{suffix}";
+    return substr($name, 0, $-[0]) . "/$twin";
+}
+
+# The timeout of a call that changes ZFS: ZFS waits for transaction groups,
+# which a busy pool can take minutes to sync.
+my sub nvmet_long_timeout() {
+    return PVE::RPCEnvironment->is_worker() ? 60 * 60 : 60;
+}
+
+# ---------------------------------------------------------------------------
+# Remote steps
+#
+# A step is an argv array, a write `{ write => $path, value => $value }` or a
+# key write `{ key => [$path, ...] }` fed from stdin. A unit is a list of steps
+# that must stay in one call.
+# ---------------------------------------------------------------------------
+
+my sub nvmet_subsys_path($nqn) {
+    return "$nvmet_root/subsystems/$nqn";
+}
+
+my sub nvmet_ns_path($nqn, $nsid) {
+    return "$nvmet_root/subsystems/$nqn/namespaces/$nsid";
+}
+
+my sub nvmet_port_path($id) {
+    return "$nvmet_root/ports/$id";
+}
+
+my sub nvmet_host_path($hostnqn) {
+    return "$nvmet_root/hosts/$hostnqn";
+}
+
+my sub nvmet_write($path, $value) {
+    return { write => $path, value => $value };
+}
+
+# Parts of a parsed configfs state; missing parts read as empty.
+my sub nvmet_namespaces($cfs, $nqn) {
+    my $subsys = $cfs->{subsystems}->{$nqn};
+    return ($subsys ? $subsys->{namespaces} : undef) // {};
+}
+
+my sub nvmet_port_linked($cfs, $id, $nqn) {
+    my $port = $cfs->{ports}->{$id};
+    return $port && ($port->{links} // {})->{$nqn};
+}
+
+my sub nvmet_linked_hosts($cfs) {
+    my %linked;
+    for my $subsys (values $cfs->{subsystems}->%*) {
+        $linked{$_} = 1 for keys(($subsys->{acl} // {})->%*);
+    }
+    return \%linked;
+}
+
+# A subsystem whose creation by this plugin stopped after its model was
+# written, before its serial: nobody else writes our model, and it exports,
+# allows and publishes nothing, so completing it takes nothing from anybody.
+# A subsystem with the kernel's default model may belong to another tool.
+my sub nvmet_unfinished_subsystem($cfs, $nqn) {
+    my $subsys = $cfs->{subsystems}->{$nqn};
+    return 0 if !$subsys;
+    return
+        $subsys->{attr_model} eq $nvmet_model
+        && !($subsys->{namespaces} // {})->%*
+        && !($subsys->{acl} // {})->%*
+        && !grep { nvmet_port_linked($cfs, $_, $nqn) } keys $cfs->{ports}->%*;
+}
+
+# Same write order as the kernel requires: the device before enabling it.
+my sub nvmet_build_unit($nqn, $nsid, $uuid, $dev) {
+    my $ns = nvmet_ns_path($nqn, $nsid);
+    return [
+        ['test', '-b', $dev],
+        ['mkdir', $ns],
+        nvmet_write("$ns/device_path", $dev),
+        nvmet_write("$ns/device_uuid", $uuid),
+        nvmet_write("$ns/buffered_io", 0),
+        nvmet_write("$ns/enable", 1),
+    ];
+}
+
+my sub nvmet_enable_unit($nqn, $nsid, $dev) {
+    return [['test', '-b', $dev], nvmet_write(nvmet_ns_path($nqn, $nsid) . '/enable', 1)];
+}
+
+my sub nvmet_unexport_unit($nqn, $nsid, $enabled) {
+    my $ns = nvmet_ns_path($nqn, $nsid);
+    return [($enabled ? (nvmet_write("$ns/enable", 0)) : ()), ['rmdir', $ns]];
+}
+
+my sub nvmet_identity($nqn, $nsid, $uuid) {
+    return ("proxmox:nvme-subsys=$nqn", "proxmox:nvme-nsid=$nsid", "proxmox:nvme-uuid=$uuid");
+}
+
+# The pool and its direct children with their identity properties.
+my sub nvmet_zfs_inventory_step($pool) {
+    return [
+        'zfs',
+        'get',
+        '-H',
+        '-p',
+        '-d',
+        '1',
+        '-t',
+        'filesystem,volume',
+        '-o',
+        'name,property,value,source',
+        join(',', @nvmet_zfs_props),
+        $pool,
+    ];
+}
+
+# Every configfs object the plugin uses, in one traversal. Keys are never
+# read back; with host NQNs, only their sha256 digests are returned. Only the
+# kernel's own groups of objects the plugin does not use are skipped, so an
+# object of another tool with one of their names is still listed.
+my sub nvmet_configfs_step($hostnqns) {
+    my @names = map { ('-o', '-name', $_) } @nvmet_cfs_attrs;
+    shift @names;
+    my @keys = map { ('-o', '-path', nvmet_host_path($_) . '/dhchap_key') } $hostnqns->@*;
+    shift @keys;
+    return [
+        'env',
+        'LC_ALL=C',
+        'find',
+        $nvmet_root,
+        '(',
+        '-path',
+        "$nvmet_root/subsystems/*/passthru",
+        '-o',
+        '-path',
+        "$nvmet_root/ports/*/ana_groups",
+        '-o',
+        '-path',
+        "$nvmet_root/ports/*/referrals",
+        ')',
+        '-type',
+        'd',
+        '-prune',
+        '-o',
+        '-type',
+        'd',
+        '-exec',
+        'printf',
+        'D %s\n',
+        '{}',
+        '+',
+        '-o',
+        '-type',
+        'l',
+        '-exec',
+        'printf',
+        'L %s\n',
+        '{}',
+        '+',
+        '-o',
+        '-type',
+        'f',
+        '(',
+        @names,
+        ')',
+        '-exec',
+        'grep',
+        '',
+        '/dev/null',
+        '{}',
+        '+',
+        (@keys ? ('-o', '-type', 'f', '(', @keys, ')', '-exec', 'sha256sum', '{}', '+') : ()),
+    ];
+}
+
+# The whole target state: the ZFS inventory, a marker and configfs. Only a
+# locked read loads nvmet first (modprobe).
+my sub nvmet_state_steps($pool, $hostnqns, $modprobe) {
+    return [
+        ($modprobe ? (['modprobe', 'nvmet_tcp']) : ()),
+        nvmet_zfs_inventory_step($pool),
+        ['printf', '%s\n', $nvmet_marker],
+        nvmet_configfs_step($hostnqns),
+    ];
+}
+
+# ---------------------------------------------------------------------------
+# Renderer and runner
+# ---------------------------------------------------------------------------
+
+my sub nvmet_valid_word($word) {
+    return defined($word) && !ref($word) && $word !~ /[\0\r\n]/;
+}
+
+my sub nvmet_valid_path($path) {
+    return nvmet_valid_word($path) && $path =~ m{\A/};
+}
+
+my sub nvmet_quote($word) {
+    return PVE::Tools::shellquote($word);
+}
+
+# Renders steps into one POSIX shell command line: simple commands joined with
+# `&&`, plus `>` for configfs writes and one `dd | tee` pipe for a key read from
+# stdin. There are no loops, variables, conditionals or substitutions. Every
+# word is quoted; all operands are built from validated configuration.
+sub _nvmet_render($steps) {
+    my $invalid = "internal error: invalid NVMe target step\n";
+    die $invalid if ref($steps) ne 'ARRAY' || !$steps->@*;
+
+    my $keys = 0;
+    my @commands;
+    for my $step ($steps->@*) {
+        if (ref($step) eq 'ARRAY') {
+            die $invalid if !$step->@* || grep { !nvmet_valid_word($_) } $step->@*;
+            die $invalid if !$nvmet_commands{ $step->[0] };
+            push @commands, join(' ', map { nvmet_quote($_) } $step->@*);
+        } elsif (ref($step) eq 'HASH' && exists($step->{write})) {
+            die $invalid
+                if keys($step->%*) != 2
+                || !nvmet_valid_path($step->{write})
+                || !nvmet_valid_word($step->{value});
+            push @commands,
+                q{printf '%s\n' }
+                . nvmet_quote($step->{value}) . ' > '
+                . nvmet_quote($step->{write});
+        } elsif (ref($step) eq 'HASH' && exists($step->{key})) {
+            my $paths = $step->{key};
+            die $invalid
+                if keys($step->%*) != 1
+                || ref($paths) ne 'ARRAY'
+                || !$paths->@*
+                || grep { !nvmet_valid_path($_) } $paths->@*;
+            $keys++;
+            die $invalid if $keys > 1;
+            # configfs needs the whole key in one write(2). dd collects its
+            # input into one output block: GNU dd without bs=, busybox dd only
+            # when ibs and obs differ. tee writes it to every host object.
+            push @commands,
+                'dd ibs=4096 obs=8192 2>/dev/null | tee '
+                . join(' ', map { nvmet_quote($_) } $paths->@*)
+                . ' >/dev/null';
+        } else {
+            die $invalid;
+        }
+    }
+    return join(' && ', @commands);
+}
+
+# Packs whole units into calls of at most $max bytes of rendered command, the
+# single argument sshd runs, below its limit. A rendered command is ASCII
+# (every operand is validated), so its length is its size in bytes.
+sub _nvmet_chunk($units, $max = $nvmet_max_command) {
+    my (@chunks, @current);
+    for my $unit ($units->@*) {
+        my @next = (@current, $unit->@*);
+        if (length(_nvmet_render(\@next)) > $max) {
+            die "internal error: NVMe target command too long\n"
+                if !@current || length(_nvmet_render($unit)) > $max;
+            push @chunks, [@current];
+            @next = $unit->@*;
+        }
+        @current = @next;
+    }
+    push @chunks, [@current] if @current;
+    return \@chunks;
+}
+
+# Never propagate the context run_command adds to the task marker: it quotes
+# the command line.
+my sub rethrow_task_interrupt($error) {
+    die "received interrupt\n" if $error =~ $RE_TASK_INTERRUPT;
+}
+
+# The only code that runs anything on the target. Returns
+# { rc => exit code, or -1 when ssh did not run to its end, out => [lines],
+# err => text } and never dies on a failing command; a stopped task dies with
+# the task marker. %opts: op (label, required), timeout and input (stdin,
+# used for keys).
+sub _nvmet_run($scfg, $steps, %opts) {
+    die "internal error: NVMe target call without label\n" if !defined($opts{op});
+    my $command = _nvmet_render($steps);
+    die "internal error: NVMe target command too long\n" if length($command) > $nvmet_max_command;
+    my $cmd = [@ssh_cmd, '-i', nvmet_ssh_key($scfg), 'root@' . nvmet_server($scfg), $command];
+    my (@out, $err);
+    my $rc = 0;
+    # Not noerr: run_command would then also swallow the interrupt of a stopped
+    # task, which must end the call. The exit code is taken from the exception.
+    eval {
+        run_command(
+            $cmd,
+            timeout => $opts{timeout} // 15,
+            outfunc => sub($line) { push @out, $line },
+            errfunc => sub($line) { $err = $line if !defined($err) && $line ne '' },
+            (defined($opts{input}) ? (input => $opts{input}) : ()),
+        );
+    };
+    if (my $error = $@) {
+        rethrow_task_interrupt($error);
+        return { rc => -1, out => [], err => 'timeout' } if $error =~ $RE_COMMAND_TIMEOUT;
+        return { rc => -1, out => [], err => 'ssh failed' } if $error !~ $RE_COMMAND_EXIT;
+        $rc = $+{code};
+    }
+    return { rc => $rc, out => \@out, err => $err // ($rc ? "exit code $rc" : '') };
+}
+
+# Monotonic clock for the activation backoff and local waits (a test seam).
+sub _now() {
+    return Time::HiRes::clock_gettime(Time::HiRes::CLOCK_MONOTONIC());
+}
+
+# Every local wait of the plugin (a test seam).
+sub _sleep($seconds) {
+    Time::HiRes::sleep($seconds);
+    return;
+}
+
+# ---------------------------------------------------------------------------
+# Parsers (pure)
+# ---------------------------------------------------------------------------
+
+sub _nvmet_split_state($text) {
+    my @parts = split /^\Q$nvmet_marker\E\n/m, $text, -1;
+    die "malformed NVMe target state\n" if @parts != 2;
+    return @parts;
+}
+
+sub _nvmet_property_value($value, $source) {
+    die "malformed ZFS property value\n"
+        if !defined($value)
+        || !defined($source)
+        || $value eq ''
+        || $source eq ''
+        || $value =~ /[\r\n\t]/
+        || $source =~ /[\r\n\t]/;
+    return $value if $source eq 'local' || $source eq 'received';
+    return '-' if $value eq '-' && $source eq '-';
+    if ($source =~ $RE_NVMET_INHERITED) {
+        die "invalid ZFS pool name\n" if $+{source} !~ $RE_NVMET_POOL;
+        return '-';
+    }
+    die "invalid ZFS property source\n";
+}
+
+# Parses the ZFS inventory into { types => { dataset => type }, volumes =>
+# { dataset => { nqn, nsid, uuid } }, last_nsid => highest NSID handed out }.
+# The inventory must be complete: the pool rows and all rows of every
+# dataset, or the read fails. Children whose names PVE never uses (for
+# example with a space) are kept as datasets but never count as owned.
+sub _nvmet_parse_zfs_inventory($text, $pool) {
+    die "invalid ZFS pool name\n" if $pool !~ $RE_NVMET_POOL;
+
+    my (%rows, %foreign);
+    for my $line (split /\n/, $text) {
+        my @fields = split /\t/, $line, -1;
+        die "malformed ZFS inventory row\n" if @fields != 4 || $line =~ /\r/;
+        my ($dataset, $property, $value, $source) = @fields;
+        if ($dataset ne $pool) {
+            die "ZFS dataset is outside configured pool\n" if index($dataset, "$pool/") != 0;
+            my $child = substr($dataset, length($pool) + 1);
+            die "ZFS dataset is outside configured pool\n"
+                if $child eq '' || index($child, '/') >= 0;
+            $foreign{$dataset} = 1 if $child !~ $RE_NVMET_DATASET_NAME;
+        }
+        my $on = $foreign{$dataset} ? '' : " on '$dataset'";
+        my $key = $nvmet_zfs_keys{$property} // die "unexpected ZFS property$on\n";
+        die "duplicate ZFS property '$property'$on\n" if exists($rows{$dataset}->{$key});
+        if ($key eq 'type') {
+            die "invalid ZFS dataset type$on\n"
+                if ($value ne 'filesystem' && $value ne 'volume') || $source ne '-';
+            $rows{$dataset}->{type} = $value;
+        } else {
+            $rows{$dataset}->{$key} = _nvmet_property_value($value, $source);
+        }
+    }
+    die "ZFS dataset inventory is missing the configured pool\n" if !$rows{$pool};
+
+    my (%types, %volumes);
+    for my $dataset (keys %rows) {
+        my $row = $rows{$dataset};
+        die "incomplete ZFS properties" . ($foreign{$dataset} ? '' : " on '$dataset'") . "\n"
+            if keys($row->%*) != scalar(@nvmet_zfs_props);
+        $types{$dataset} = $row->{type};
+        next if $row->{type} ne 'volume';
+        # nvmet shows device_uuid and udev names the by-id link in lower case,
+        # so identities are compared in lower case.
+        $volumes{$dataset} =
+            $foreign{$dataset}
+            ? { nqn => '-', nsid => '-', uuid => '-' }
+            : { nqn => $row->{nqn}, nsid => $row->{nsid}, uuid => lc($row->{uuid}) };
+    }
+    my $last_nsid = $rows{$pool}->{last_nsid};
+    return {
+        types => \%types,
+        volumes => \%volumes,
+        last_nsid => nvmet_valid_nsid($last_nsid) ? $last_nsid : 0,
+    };
+}
+
+# Parses the configfs read into
+# { subsystems => { nqn => { attr_model, attr_serial, attr_allow_any_host,
+#       acl => { hostnqn => 1 }, namespaces => { nsid => { enable, device_path,
+#       device_uuid, buffered_io } } } },
+#   ports => { id => { addr_trtype, addr_adrfam, addr_traddr, addr_trsvcid,
+#       links => { nqn => 1 } } },
+#   hosts => { hostnqn => { key_sha256 => hex or undef } } }
+# Every object must be complete; unknown objects of newer kernels are ignored.
+sub _nvmet_parse_configfs($text) {
+    my $state = { subsystems => {}, ports => {}, hosts => {} };
+    my (%top, %dir);
+
+    my $subsystem = sub($nqn) {
+        return $state->{subsystems}->{$nqn} //= { acl => {}, namespaces => {} };
+    };
+    my $namespace = sub($nqn, $nsid) {
+        return $subsystem->($nqn)->{namespaces}->{$nsid} //= {};
+    };
+    my $port = sub($id) { return $state->{ports}->{$id} //= { links => {} } };
+    my $host = sub($hostnqn) { return $state->{hosts}->{$hostnqn} //= { key_sha256 => undef } };
+    my $set = sub($object, $attr, $value) {
+        die "duplicate NVMe target configfs attribute\n" if exists($object->{$attr});
+        $object->{$attr} = $value;
+    };
+
+    for my $line (split /\n/, $text) {
+        if ($line =~ $RE_CFS_TOP) {
+            $top{ $+{top} // 'root' } = 1;
+        } elsif ($line =~ $RE_CFS_SUBSYS_DIR) {
+            my $nqn = $+{nqn};
+            $subsystem->($nqn);
+            $dir{"subsystem $nqn"} = 1;
+        } elsif ($line =~ $RE_CFS_SUBSYS_GROUP || $line =~ $RE_CFS_PORT_GROUP) {
+            # default groups, always present with their parent
+        } elsif ($line =~ $RE_CFS_NS_DIR) {
+            my ($nqn, $nsid) = @+{qw(nqn nsid)};
+            $namespace->($nqn, $nsid);
+            $dir{"namespace $nqn/$nsid"} = 1;
+        } elsif ($line =~ $RE_CFS_PORT_DIR) {
+            my $id = $+{port};
+            $port->($id);
+            $dir{"port $id"} = 1;
+        } elsif ($line =~ $RE_CFS_HOST_DIR) {
+            my $hostnqn = $+{host};
+            $host->($hostnqn);
+            $dir{"host $hostnqn"} = 1;
+        } elsif ($line =~ $RE_CFS_PORT_LINK) {
+            my ($id, $nqn) = @+{qw(port nqn)};
+            $port->($id)->{links}->{$nqn} = 1;
+        } elsif ($line =~ $RE_CFS_ACL_LINK) {
+            my ($nqn, $hostnqn) = @+{qw(nqn host)};
+            $subsystem->($nqn)->{acl}->{$hostnqn} = 1;
+        } elsif ($line =~ $RE_CFS_OTHER_ENTRY) {
+            # objects of a shape the plugin does not use
+        } elsif ($line =~ $RE_CFS_NS_ATTR) {
+            my ($nqn, $nsid, $attr, $value) = @+{qw(nqn nsid attr value)};
+            $set->($namespace->($nqn, $nsid), $attr, $value);
+        } elsif ($line =~ $RE_CFS_SUBSYS_ATTR) {
+            my ($nqn, $attr, $value) = @+{qw(nqn attr value)};
+            $set->($subsystem->($nqn), $attr, $value);
+        } elsif ($line =~ $RE_CFS_PORT_ATTR) {
+            my ($id, $attr, $value) = @+{qw(port attr value)};
+            $set->($port->($id), $attr, $value);
+        } elsif ($line =~ $RE_CFS_KEY_DIGEST) {
+            my ($sha, $hostnqn) = @+{qw(sha host)};
+            my $object = $host->($hostnqn);
+            die "duplicate DH-HMAC-CHAP key digest" . nvmet_label($hostnqn) . "\n"
+                if defined($object->{key_sha256});
+            $object->{key_sha256} = $sha;
+        } elsif ($line =~ $RE_CFS_OTHER_ATTR) {
+            # attribute of an object shape the plugin does not use
+        } else {
+            die "unexpected NVMe target configfs line\n";
+        }
+    }
+
+    for my $name (qw(root hosts ports subsystems)) {
+        die "NVMe target configfs is unavailable\n" if !$top{$name};
+    }
+    for my $nqn (keys $state->{subsystems}->%*) {
+        my $subsys = $state->{subsystems}->{$nqn};
+        die "incomplete NVMe subsystem" . nvmet_label($nqn) . "\n"
+            if !$dir{"subsystem $nqn"}
+            || grep { !defined($subsys->{$_}) } qw(attr_model attr_serial attr_allow_any_host);
+        for my $nsid (keys $subsys->{namespaces}->%*) {
+            my $ns = $subsys->{namespaces}->{$nsid};
+            die "incomplete NVMe namespace '$nsid' of subsystem" . nvmet_label($nqn) . "\n"
+                if !$dir{"namespace $nqn/$nsid"}
+                || grep { !defined($ns->{$_}) } qw(enable device_path device_uuid buffered_io);
+        }
+    }
+    for my $id (keys $state->{ports}->%*) {
+        my $object = $state->{ports}->{$id};
+        my @attrs = qw(addr_trtype addr_adrfam addr_traddr addr_trsvcid);
+        die "incomplete NVMe port '$id'\n"
+            if !$dir{"port $id"} || grep { !defined($object->{$_}) } @attrs;
+    }
+    for my $hostnqn (keys $state->{hosts}->%*) {
+        die "DH-HMAC-CHAP key digest for unknown NVMe host" . nvmet_label($hostnqn) . "\n"
+            if !$dir{"host $hostnqn"};
+    }
+    return $state;
+}
+
+sub _nvmet_parse_mounts($text) {
+    for my $line (split /\n/, $text) {
+        my (undef, $mountpoint, $type) = split / /, $line;
+        return 1 if ($mountpoint // '') eq '/sys/kernel/config' && ($type // '') eq 'configfs';
+    }
+    return 0;
+}
+
+# ---------------------------------------------------------------------------
+# Planners (pure)
+# ---------------------------------------------------------------------------
+
+sub _nvmet_serial($nqn) {
+    return 'PVEZFS' . substr(sha256_hex($nqn), 0, 14);
+}
+
+sub _nvmet_template_name($dataset) {
+    die "only VM zvols can become templates\n"
+        if $dataset !~ $RE_NVMET_TEMPLATE_ZVOL || $+{type} ne 'vm';
+    return nvmet_template_twin($dataset);
+}
+
+# The durable identity (NSID, UUID) of a volume owned by $nqn. Refuses a
+# volume that another owned volume duplicates anywhere in the pool.
+sub _nvmet_owned_identity($inv, $nqn, $dataset) {
+    my $type = $inv->{types}->{$dataset} // die "ZFS volume '$dataset' does not exist\n";
+    die "'$dataset' is not a ZFS volume\n" if $type ne 'volume';
+    my $row = $inv->{volumes}->{$dataset};
+    die "ZFS volume '$dataset' is not owned by NVMe subsystem '$nqn'\n" if $row->{nqn} ne $nqn;
+    die "invalid NSID on '$dataset'\n" if !nvmet_valid_nsid($row->{nsid});
+    die "invalid namespace UUID on '$dataset'\n" if $row->{uuid} !~ $RE_NVMET_UUID;
+    for my $other (keys $inv->{volumes}->%*) {
+        next if $other eq $dataset;
+        my $identity = $inv->{volumes}->{$other};
+        next if $identity->{nqn} ne $nqn;
+        die "duplicate NVMe identity on '$dataset'\n"
+            if $identity->{nsid} eq $row->{nsid} || $identity->{uuid} eq $row->{uuid};
+    }
+    return ($row->{nsid}, $row->{uuid});
+}
+
+# { nsid => [uuid, device] } for every volume owned by $nqn, after validating
+# all of them, so a bad identity refuses the whole plan before any change.
+sub _nvmet_desired_namespaces($inv, $nqn) {
+    my (%desired, %uuids);
+    for my $dataset (sort keys $inv->{volumes}->%*) {
+        my $row = $inv->{volumes}->{$dataset};
+        next if $row->{nqn} ne $nqn;
+        die "invalid NSID on '$dataset'\n" if !nvmet_valid_nsid($row->{nsid});
+        die "invalid namespace UUID on '$dataset'\n" if $row->{uuid} !~ $RE_NVMET_UUID;
+        die "duplicate NSID '$row->{nsid}'\n" if $desired{ $row->{nsid} };
+        die "duplicate namespace UUID '$row->{uuid}'\n" if $uuids{ lc($row->{uuid}) }++;
+        $desired{ $row->{nsid} } = [$row->{uuid}, "/dev/zvol/$dataset"];
+    }
+    return \%desired;
+}
+
+# NSIDs grow monotonically, also past the last NSID handed out in the pool,
+# so a host that missed a namespace removal never sees a new volume at an old
+# NSID. Only after the last NSID is used does the allocation fall back to the
+# lowest free one.
+sub _nvmet_allocate_nsid($inv, $cfs, $nqn) {
+    my %used;
+    for my $row (values $inv->{volumes}->%*) {
+        $used{ $row->{nsid} } = 1 if $row->{nqn} eq $nqn && nvmet_valid_nsid($row->{nsid});
+    }
+    $used{$_} = 1 for keys nvmet_namespaces($cfs, $nqn)->%*;
+
+    my $highest = max(0, $inv->{last_nsid} // 0, keys %used);
+    return $highest + 1 if $highest < $nvmet_max_nsid;
+    my $nsid = 1;
+    $nsid++ while $used{$nsid};
+    die "no free namespace ID\n" if $nsid > $nvmet_max_nsid;
+    return $nsid;
+}
+
+# Export state of (NSID, UUID, device) in the subsystem: present, disabled,
+# absent, incomplete (a disabled namespace with another identity, for example
+# one being built), stale (enabled on the other template name of the zvol) or
+# nosubsys. Dies on a conflict.
+sub _nvmet_export_state($cfs, $nqn, $nsid, $uuid, $dev) {
+    return 'nosubsys' if !$cfs->{subsystems}->{$nqn};
+    my $namespaces = nvmet_namespaces($cfs, $nqn);
+    for my $id (keys $namespaces->%*) {
+        next if $id eq $nsid;
+        die "namespace UUID '$uuid' is already in use\n"
+            if $namespaces->{$id}->{device_uuid} eq $uuid;
+    }
+    my $ns = $namespaces->{$nsid} // return 'absent';
+    if ($ns->{device_uuid} eq $uuid && $ns->{device_path} eq $dev) {
+        return $ns->{enable} eq '1' ? 'present' : 'disabled';
+    }
+    # A template conversion or its undo that ssh gave up on can rename the
+    # zvol on the target after an activation exported it under its old name.
+    my $twin = nvmet_template_twin($dev);
+    return 'stale'
+        if $ns->{enable} eq '1'
+        && $ns->{device_uuid} eq $uuid
+        && $twin ne ''
+        && $ns->{device_path} eq $twin;
+    die "NVMe namespace ID '$nsid' has a different identity; refusing to replace it\n"
+        if $ns->{enable} ne '0';
+    return 'incomplete';
+}
+
+# Units that export (NSID, UUID, device) and the state they start from:
+# absent, present, disabled, reclaim or stale. A disabled namespace with
+# another identity, or a stale one, is rebuilt, which is only correct under
+# the target lock. A stale namespace belongs to a template, or to a volume
+# whose template conversion failed, so no guest uses it.
+sub _nvmet_plan_export($cfs, $nqn, $nsid, $uuid, $dev) {
+    die "NVMe subsystem does not exist\n" if !$cfs->{subsystems}->{$nqn};
+    my $state = _nvmet_export_state($cfs, $nqn, $nsid, $uuid, $dev);
+    return ([nvmet_build_unit($nqn, $nsid, $uuid, $dev)], 'absent') if $state eq 'absent';
+    return ([], 'present') if $state eq 'present';
+    return ([nvmet_enable_unit($nqn, $nsid, $dev)], 'disabled') if $state eq 'disabled';
+    my $unexport = nvmet_unexport_unit($nqn, $nsid, $state eq 'stale');
+    my $build = nvmet_build_unit($nqn, $nsid, $uuid, $dev);
+    return ([[$unexport->@*, $build->@*]], $state eq 'stale' ? 'stale' : 'reclaim');
+}
+
+# Units that remove the namespace exporting $uuid, and its previous state.
+sub _nvmet_plan_unexport($cfs, $nqn, $uuid) {
+    my $namespaces = nvmet_namespaces($cfs, $nqn);
+    my @matches = grep { $namespaces->{$_}->{device_uuid} eq $uuid }
+        sort { $a <=> $b } keys $namespaces->%*;
+    die "duplicate namespace UUID '$uuid'\n" if @matches > 1;
+    return ([], undef) if !@matches;
+    my $enabled = $namespaces->{ $matches[0] }->{enable} eq '1';
+    return (
+        [nvmet_unexport_unit($nqn, $matches[0], $enabled)],
+        { nsid => $matches[0], enabled => $enabled },
+    );
+}
+
+# The port serving a portal ({ family, address, port } of parse_nvme_portals).
+# Among complete, matching ports, the one already linked to $nqn wins, then
+# the lowest id. Incomplete ports never match.
+sub _nvmet_find_port($cfs, $nqn, $portal) {
+    my $wanted = nvmet_canonical_address($portal->{address});
+    my @matches = grep {
+        my $port = $cfs->{ports}->{$_};
+        $port->{addr_trtype} eq 'tcp'
+            && $port->{addr_adrfam} eq $portal->{family}
+            && $port->{addr_traddr} ne ''
+            && $port->{addr_trsvcid} eq "$portal->{port}"
+            && nvmet_canonical_address($port->{addr_traddr}) eq $wanted
+        } sort {
+            $a <=> $b
+        } keys $cfs->{ports}->%*;
+    my @linked = grep { nvmet_port_linked($cfs, $_, $nqn) } @matches;
+    return $linked[0] // $matches[0];
+}
+
+# Units that bring the target to the configured state, except publishing:
+# objects (subsystem, ports, hosts), keys, ACLs, namespaces in NSID order,
+# then removal of undesired namespaces.
+# $conf: { nqn, portals (of parse_nvme_portals), hostnqns, keysha }
+sub _nvmet_plan_activation($inv, $cfs, $conf) {
+    my ($nqn, $keysha) = $conf->@{qw(nqn keysha)};
+    my $desired = _nvmet_desired_namespaces($inv, $nqn); # refuse before any unit
+    my $subsys_path = nvmet_subsys_path($nqn);
+    my $subsys = $cfs->{subsystems}->{$nqn};
+    my (@objects, @key_hosts, @acls, @namespaces, %port_ids);
+
+    if (!$subsys) {
+        push @objects,
+            [
+                ['mkdir', $subsys_path],
+                nvmet_write("$subsys_path/attr_model", $nvmet_model),
+                nvmet_write("$subsys_path/attr_serial", _nvmet_serial($nqn)),
+                nvmet_write("$subsys_path/attr_allow_any_host", 0),
+            ];
+    } else {
+        my $serial = _nvmet_serial($nqn);
+        my @attrs;
+        if ($subsys->{attr_model} ne $nvmet_model || $subsys->{attr_serial} ne $serial) {
+            die "refusing to take over existing NVMe subsystem '$nqn': it was not created by"
+                . " Proxmox VE; remove it on the target or use another NQN\n"
+                if !nvmet_unfinished_subsystem($cfs, $nqn);
+            push @attrs, nvmet_write("$subsys_path/attr_serial", $serial);
+        }
+        push @attrs, nvmet_write("$subsys_path/attr_allow_any_host", 0)
+            if $subsys->{attr_allow_any_host} ne '0';
+        push @objects, [@attrs] if @attrs;
+    }
+
+    my %taken = map { $_ => 1 } keys $cfs->{ports}->%*;
+    for my $portal ($conf->{portals}->@*) {
+        my $id = _nvmet_find_port($cfs, $nqn, $portal);
+        if (!defined($id)) {
+            $id = 1;
+            $id++ while $taken{$id};
+            $taken{$id} = 1;
+            my $port_path = nvmet_port_path($id);
+            push @objects,
+                [
+                    ['mkdir', $port_path],
+                    nvmet_write("$port_path/addr_trtype", 'tcp'),
+                    nvmet_write("$port_path/addr_adrfam", $portal->{family}),
+                    nvmet_write("$port_path/addr_traddr", $portal->{address}),
+                    nvmet_write("$port_path/addr_trsvcid", $portal->{port}),
+                ];
+        }
+        $port_ids{$id} = 1;
+    }
+
+    # nvmet keeps the key on the global host object. A key that differs from
+    # ours is only replaced while no subsystem links the host.
+    my $linked = nvmet_linked_hosts($cfs);
+    for my $hostnqn ($conf->{hostnqns}->@*) {
+        my $host_path = nvmet_host_path($hostnqn);
+        if (my $host = $cfs->{hosts}->{$hostnqn}) {
+            my $sha = $host->{key_sha256}
+                // die "NVMe target does not expose a DH-HMAC-CHAP key for host '$hostnqn'\n";
+            if ($sha ne $keysha) {
+                die "refusing to replace an in-use DH-HMAC-CHAP key for '$hostnqn'\n"
+                    if $linked->{$hostnqn};
+                push @key_hosts, $host_path;
+            }
+        } else {
+            push @objects, [['mkdir', $host_path]];
+            push @key_hosts, $host_path;
+        }
+        push @acls, [['ln', '-s', $host_path, "$subsys_path/allowed_hosts/$hostnqn"]]
+            if !$subsys || !($subsys->{acl} // {})->{$hostnqn};
+    }
+    # nvmet creates the key attributes world-readable. The chmod before the
+    # write keeps every later open from reading the key; a descriptor opened
+    # while an attribute was still world-readable is not affected by it.
+    my @keys;
+    if (@key_hosts) {
+        my @attrs = map { ("$_/dhchap_key", "$_/dhchap_ctrl_key") } @key_hosts;
+        @keys = ([['chmod', '0600', @attrs], { key => [map { "$_/dhchap_key" } @key_hosts] }]);
+    }
+
+    my $current = { subsystems => { $nqn => $subsys // { acl => {}, namespaces => {} } } };
+    for my $nsid (sort { $a <=> $b } keys $desired->%*) {
+        my ($units) = _nvmet_plan_export($current, $nqn, $nsid, $desired->{$nsid}->@*);
+        push @namespaces, $units->@*;
+    }
+    my $existing = nvmet_namespaces($cfs, $nqn);
+    for my $nsid (sort { $a <=> $b } keys $existing->%*) {
+        next if $desired->{$nsid};
+        push @namespaces, nvmet_unexport_unit($nqn, $nsid, $existing->{$nsid}->{enable} eq '1');
+    }
+
+    return {
+        prepublish => [@objects, @keys, @acls, @namespaces],
+        port_ids => [sort { $a <=> $b } keys %port_ids],
+    };
+}
+
+# One call per port: each link repeats the guards, so it only succeeds while
+# the subsystem denies unknown hosts and every configured host has its ACL.
+# Under the target lock, only a manager of the target outside this cluster
+# can change that after the verify read; the guards stop the publish then.
+# Links to ports that are no longer configured are removed.
+# Returns [{ port => $id, action => 'link' | 'unlink', steps => [...] }, ...].
+sub _nvmet_plan_publish($cfs, $nqn, $port_ids, $hostnqns) {
+    my $subsys_path = nvmet_subsys_path($nqn);
+    my %wanted = map { $_ => 1 } $port_ids->@*;
+    my @guards = (
+        ['grep', '-qx', '0', "$subsys_path/attr_allow_any_host"],
+        map { ['test', '-L', "$subsys_path/allowed_hosts/$_"] } $hostnqns->@*,
+    );
+
+    my @units;
+    for my $id ($port_ids->@*) {
+        next if nvmet_port_linked($cfs, $id, $nqn);
+        my $link = ['ln', '-s', $subsys_path, nvmet_port_path($id) . "/subsystems/$nqn"];
+        push @units, { port => $id, action => 'link', steps => [@guards, $link] };
+    }
+    for my $id (sort { $a <=> $b } keys $cfs->{ports}->%*) {
+        next if $wanted{$id} || !nvmet_port_linked($cfs, $id, $nqn);
+        my $unlink = ['rm', nvmet_port_path($id) . "/subsystems/$nqn"];
+        push @units, { port => $id, action => 'unlink', steps => [$unlink] };
+    }
+    return \@units;
+}
+
+# Teardown units for the subsystem and the hosts that may become orphans.
+sub _nvmet_plan_delete_target($inv, $cfs, $nqn, $hostnqns) {
+    for my $dataset (sort keys $inv->{volumes}->%*) {
+        die "refusing to delete NVMe subsystem '$nqn': owned ZFS volume '$dataset' exists\n"
+            if $inv->{volumes}->{$dataset}->{nqn} eq $nqn;
+    }
+    my $subsys = $cfs->{subsystems}->{$nqn} // return ([], [$hostnqns->@*]);
+    die "refusing to delete foreign NVMe subsystem '$nqn'\n"
+        if $subsys->{attr_model} ne $nvmet_model || $subsys->{attr_serial} ne _nvmet_serial($nqn);
+    die "refusing to delete NVMe subsystem '$nqn': namespaces remain\n"
+        if nvmet_namespaces($cfs, $nqn)->%*;
+
+    my $subsys_path = nvmet_subsys_path($nqn);
+    my @acl = sort keys(($subsys->{acl} // {})->%*);
+    die "refusing to delete NVMe subsystem '$nqn': it allows a malformed host name\n"
+        if grep { $_ !~ $RE_NQN } @acl;
+    my @links = (
+        (
+            map { nvmet_port_path($_) . "/subsystems/$nqn" }
+            grep { nvmet_port_linked($cfs, $_, $nqn) }
+            sort { $a <=> $b } keys $cfs->{ports}->%*
+        ),
+        (map { "$subsys_path/allowed_hosts/$_" } @acl),
+    );
+    # A retry after a teardown that failed at its rmdir finds no ACL left, so
+    # the configured hosts are candidates as well.
+    my %candidates = map { $_ => 1 } @acl, $hostnqns->@*;
+    return (
+        [[(@links ? (['rm', @links]) : ()), ['rmdir', $subsys_path]]], [sort keys %candidates],
+    );
+}
+
+# Candidate hosts that exist and that no subsystem links any more.
+sub _nvmet_plan_orphan_hosts($cfs, $candidates) {
+    my $linked = nvmet_linked_hosts($cfs);
+    my %seen;
+    my @orphans = map { nvmet_host_path($_) }
+        grep { !$seen{$_}++ && $_ =~ $RE_NQN && $cfs->{hosts}->{$_} && !$linked->{$_} }
+        sort $candidates->@*;
+    return @orphans ? [[['rmdir', @orphans]]] : [];
+}
+
+# ---------------------------------------------------------------------------
+# Target lock and remote execution
+# ---------------------------------------------------------------------------
+
+my %nvmet_lock_owner; # lock id => pid of the process holding the domain lock
+
+# Runs $code under the pmxcfs domain lock of the target. Like the pmxcfs
+# lock itself, it is not re-entrant, and a child forked inside is not the
+# owner.
+my sub nvmet_locked($scfg, $code) {
+    my $id = nvmet_lock_id($scfg);
+    die "cluster not quorate - refusing NVMe target changes\n"
+        if !PVE::Cluster::check_cfs_quorum(1);
+    my $res = PVE::Cluster::cfs_lock_domain(
+        $id,
+        $nvmet_lock_wait,
+        sub {
+            local $nvmet_lock_owner{$id} = $$;
+            return $code->();
+        },
+    );
+    die $@ if $@;
+    return $res;
+}
+
+# Runs steps on the target. %opts: op, timeout, input, change (the steps
+# change the target, which needs the lock).
+#
+# ssh does not stop a command on the target when it gives up on it: a change
+# whose connection broke (exit code 255), any call under the lock that ssh
+# did not run to its end (-1: its timeout, or ssh was killed), and a call
+# during which the task was stopped (it dies with the task marker) may still
+# be running there. Its outcome is unknown, so it is never compensated from a
+# read that could come before the rest of it. The next activation repairs
+# from the target state.
+my sub nvmet_exec($scfg, $steps, %opts) {
+    my $locked = ($nvmet_lock_owner{ nvmet_lock_id($scfg) } // 0) == $$;
+    die "internal error: NVMe target change without target lock\n"
+        if $opts{change} && !$locked;
+    my $res = _nvmet_run(
+        $scfg, $steps,
+        op => $opts{op},
+        timeout => $opts{timeout} // 15,
+        (defined($opts{input}) ? (input => $opts{input}) : ()),
+    );
+    die "NVMe target operation '$opts{op}' did not complete ($res->{err});"
+        . " the target state is unknown\n"
+        if ($locked && $res->{rc} == -1)
+        || ($opts{change} && $res->{rc} == 255);
+    return $res;
+}
+
+# An abandoned operation is passed on instead of being compensated.
+my sub nvmet_rethrow_abandoned($error) {
+    die $error if $error =~ $RE_NVMET_ABANDONED;
+}
+
+my sub nvmet_unreachable($res) {
+    return $res->{rc} == 255 || $res->{rc} == -1;
+}
+
+my sub nvmet_check($res, $op) {
+    die "NVMe target operation '$op' failed: $res->{err}\n" if $res->{rc};
+}
+
+my sub nvmet_read_failed($scfg, $res) {
+    die "NVMe target '" . nvmet_server($scfg) . "' is unreachable: $res->{err}\n"
+        if nvmet_unreachable($res);
+    die "cannot read NVMe target state: $res->{err}\n";
+}
+
+my sub nvmet_output($res) {
+    return join('', map { "$_\n" } $res->{out}->@*);
+}
+
+# A read, retried twice after 200 ms unless the target is unreachable: a read
+# without the lock can meet a configfs object that another node removes
+# while find lists it, which makes find fail once.
+my sub nvmet_read_call($scfg, $steps, %opts) {
+    my $res;
+    for my $attempt (0 .. 2) {
+        _sleep(0.2) if $attempt;
+        $res = nvmet_exec($scfg, $steps, %opts, timeout => 15);
+        last if !$res->{rc} || nvmet_unreachable($res);
+    }
+    return $res;
+}
+
+# Reads the target state and returns (ZFS inventory, configfs state).
+# keys => [hostnqns] adds the digests of their keys. modprobe => 1, for locked
+# reads only, loads nvmet first and mounts configfs if that is what failed.
+# With tolerant => 1, a failed or unparsable read returns an empty list,
+# unless the target is unreachable.
+my sub nvmet_read($scfg, %opts) {
+    my $pool = nvmet_pool($scfg);
+    my $hostnqns = $opts{keys} // [];
+    # loading a module that may still be loading is harmless
+    my @modprobe = $opts{modprobe} ? (change => 1) : ();
+    my $read = sub () {
+        return nvmet_read_call(
+            $scfg,
+            nvmet_state_steps($pool, $hostnqns, $opts{modprobe}),
+            op => 'read target state',
+            @modprobe,
+        );
+    };
+    my $res = $read->();
+    if ($res->{rc} && $opts{modprobe} && !nvmet_unreachable($res)) {
+        my $mounts = nvmet_exec($scfg, [['cat', '/proc/mounts']], op => 'read mounts');
+        nvmet_check($mounts, 'read mounts');
+        if (!_nvmet_parse_mounts(nvmet_output($mounts))) {
+            my $mount = ['mount', '-t', 'configfs', 'none', '/sys/kernel/config'];
+            my $mounted = nvmet_exec($scfg, [$mount], op => 'mount configfs', change => 1);
+            nvmet_check($mounted, 'mount configfs');
+            $res = $read->();
+        }
+    }
+    if ($res->{rc}) {
+        return () if $opts{tolerant} && !nvmet_unreachable($res);
+        nvmet_read_failed($scfg, $res);
+    }
+    my @state = eval {
+        my ($zfs, $configfs) = _nvmet_split_state(nvmet_output($res));
+        (_nvmet_parse_zfs_inventory($zfs, $pool), _nvmet_parse_configfs($configfs));
+    };
+    if (my $error = $@) {
+        return () if $opts{tolerant};
+        die $error;
+    }
+    return @state;
+}
+
+my sub nvmet_read_zfs($scfg) {
+    my $pool = nvmet_pool($scfg);
+    my $res = nvmet_read_call($scfg, [nvmet_zfs_inventory_step($pool)], op => 'read ZFS inventory');
+    nvmet_read_failed($scfg, $res) if $res->{rc};
+    return _nvmet_parse_zfs_inventory(nvmet_output($res), $pool);
+}
+
+# Applies units in as few calls as possible and stops at the first failing
+# call, whose result it returns. A chain only saves SSH round trips: every
+# decision was taken locally before, and `&&` only stops at the first
+# failing command. Only the call with the key step gets stdin. The default
+# timeout suits configfs chains; ZFS changes pass their own.
+my sub nvmet_apply($scfg, $units, %opts) {
+    my $input = delete $opts{input};
+    $opts{timeout} //= 30;
+    for my $chunk (_nvmet_chunk($units)->@*) {
+        my $with_key = grep { ref($_) eq 'HASH' && exists($_->{key}) } $chunk->@*;
+        my @input = $with_key ? (input => $input) : ();
+        my $res = nvmet_exec($scfg, $chunk, %opts, @input, change => 1);
+        return $res if $res->{rc};
+    }
+    return { rc => 0, out => [], err => '' };
+}
+
+my sub nvmet_change($scfg, $units, %opts) {
+    nvmet_check(nvmet_apply($scfg, $units, %opts), $opts{op});
+}
+
+my sub nvmet_wait_device($scfg, $dev) {
+    my $deadline = _now() + 10;
+    while (1) {
+        my $res = nvmet_exec($scfg, [['test', '-b', $dev]], op => 'wait for zvol');
+        return if !$res->{rc};
+        nvmet_read_failed($scfg, $res) if nvmet_unreachable($res);
+        last if _now() >= $deadline;
+        _sleep(0.25);
+    }
+    die "zvol '$dev' is not a block device after 10 seconds\n";
+}
+
+my sub nvmet_new_uuid($inv, $cfs) {
+    my %used = map { lc($_->{uuid}) => 1 } values $inv->{volumes}->%*;
+    for my $subsys (values $cfs->{subsystems}->%*) {
+        $used{ lc($_->{device_uuid}) } = 1 for values(($subsys->{namespaces} // {})->%*);
+    }
+    for (1 .. 10) {
+        my $uuid = lc(file_read_firstline('/proc/sys/kernel/random/uuid') // '');
+        return $uuid if $uuid =~ $RE_NVMET_UUID && !$used{$uuid};
+    }
+    die "cannot generate an unused namespace UUID\n";
+}
+
+my sub nvmet_has_identity($row, $nqn, $nsid, $uuid) {
+    return $row && $row->{nqn} eq $nqn && $row->{nsid} eq "$nsid" && $row->{uuid} eq $uuid;
+}
+
+# The configfs state after a successful unexport of namespace $pre->{nsid}.
+my sub nvmet_without_namespace($cfs, $nqn, $pre) {
+    return $cfs if !$pre;
+    my $subsys = $cfs->{subsystems}->{$nqn};
+    my %namespaces = $subsys->{namespaces}->%*;
+    delete $namespaces{ $pre->{nsid} };
+    my $changed = { $subsys->%*, namespaces => \%namespaces };
+    return { $cfs->%*, subsystems => { $cfs->{subsystems}->%*, $nqn => $changed } };
+}
+
+# Links the subsystem to each desired port separately. One port that cannot
+# be bound (for example, its address is missing on the target) only warns
+# while another desired port serves the subsystem.
+my sub nvmet_publish($scfg, $cfs, $nqn, $port_ids, $hostnqns) {
+    my @failed;
+    for my $unit (_nvmet_plan_publish($cfs, $nqn, $port_ids, $hostnqns)->@*) {
+        my $op = "$unit->{action} NVMe subsystem on port $unit->{port}";
+        my $res = nvmet_apply($scfg, [$unit->{steps}], op => $op);
+        push @failed, { $unit->%*, err => $res->{err} } if $res->{rc};
+    }
+    return if !@failed;
+
+    my (undef, $fresh) = eval { nvmet_read($scfg) };
+    nvmet_rethrow_abandoned($@) if $@;
+    my $published = $fresh && grep { nvmet_port_linked($fresh, $_, $nqn) } $port_ids->@*;
+    my ($first) = grep { $_->{action} eq 'link' } @failed;
+    $first //= $failed[0];
+    die "cannot publish NVMe subsystem '$nqn': $first->{err}\n" if !$published;
+    for my $failure (@failed) {
+        my $what = $failure->{action} eq 'link' ? 'publish' : 'unpublish';
+        log_warn("cannot $what NVMe subsystem '$nqn' on port $failure->{port}: $failure->{err}");
+    }
+}
+
+# ---------------------------------------------------------------------------
+# Target flows
+# ---------------------------------------------------------------------------
+
+# Brings the target to the configured state and publishes the subsystem.
+# A converged target costs one read and no lock. Otherwise, under the lock:
+# read, apply, verify with a fresh read, and publish only once the verify
+# read plans nothing more. $portals and $hostnqns are the parsed and validated
+# storage configuration.
+sub _nvmet_activate_target($storeid, $scfg, $portals, $hostnqns, $key) {
+    my $nqn = nvmet_nqn($scfg);
+    my $conf = {
+        nqn => $nqn,
+        portals => $portals,
+        hostnqns => $hostnqns,
+        keysha => sha256_hex("$key\n"),
+    };
+    my $converged = sub($inv, $cfs) {
+        my $plan = _nvmet_plan_activation($inv, $cfs, $conf);
+        return !$plan->{prepublish}->@*
+            && !_nvmet_plan_publish($cfs, $nqn, $plan->{port_ids}, $hostnqns)->@*;
+    };
+
+    # Without the lock, a refusal or failed read only means "take the lock".
+    my @state = nvmet_read($scfg, keys => $hostnqns, tolerant => 1);
+    return if @state && eval { $converged->(@state) };
+
+    nvmet_locked(
+        $scfg,
+        sub {
+            # Removing the storage deletes its key before it removes the target
+            # configuration, so an activation that waited for the lock meanwhile
+            # does not restore it.
+            die "storage '$storeid' is being removed or its key changed\n"
+                if (file_read_firstline(secret_path($storeid)) // '') ne $key;
+            my ($inv, $cfs) = nvmet_read($scfg, keys => $hostnqns, modprobe => 1);
+            my $plan = _nvmet_plan_activation($inv, $cfs, $conf);
+            my $error;
+            for my $round (1 .. 2) {
+                last if !$plan->{prepublish}->@*;
+                my $apply = nvmet_apply(
+                    $scfg, $plan->{prepublish},
+                    op => 'configure NVMe target',
+                    input => "$key\n",
+                );
+                $error //= $apply->{err} if $apply->{rc};
+                ($inv, $cfs) = nvmet_read($scfg, keys => $hostnqns);
+                $plan = _nvmet_plan_activation($inv, $cfs, $conf);
+            }
+            die "NVMe target did not converge: " . ($error // 'state mismatch') . "\n"
+                if $plan->{prepublish}->@*;
+            nvmet_publish($scfg, $cfs, $nqn, $plan->{port_ids}, $hostnqns);
+        },
+    );
+    return;
+}
+
+# The namespace UUID of an owned volume, from one read-only ZFS query.
+sub _nvmet_volume_uuid($scfg, $name) {
+    my $dataset = nvmet_dataset($scfg, $name);
+    my $inv = nvmet_read_zfs($scfg);
+    my (undef, $uuid) = _nvmet_owned_identity($inv, nvmet_nqn($scfg), $dataset);
+    return $uuid;
+}
+
+# Creates (size => KiB) or clones (origin => pool/base@snap) a zvol with its
+# identity set at creation, then exports it. A failure removes only what this
+# call created, decided from a fresh read.
+my sub nvmet_create_volume($scfg, $name, %opts) {
+    my $nqn = nvmet_nqn($scfg);
+    my $dataset = nvmet_dataset($scfg, $name);
+    my $dev = "/dev/zvol/$dataset";
+    die "internal error: invalid zvol creation request\n"
+        if defined($opts{size}) == defined($opts{origin})
+        || (defined($opts{size}) && $opts{size} !~ $RE_UNSIGNED_INTEGER);
+
+    return nvmet_locked(
+        $scfg,
+        sub {
+            my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1);
+            die "zvol '$dataset' already exists\n" if exists($inv->{types}->{$dataset});
+            die "NVMe subsystem does not exist; activate the storage first\n"
+                if !$cfs->{subsystems}->{$nqn};
+            my $nsid = _nvmet_allocate_nsid($inv, $cfs, $nqn);
+            my $uuid = nvmet_new_uuid($inv, $cfs);
+            my (undef, $state) = _nvmet_plan_export($cfs, $nqn, $nsid, $uuid, $dev);
+            die "internal error: namespace ID '$nsid' is in use\n" if $state ne 'absent';
+
+            # Record the NSID on the pool before the zvol exists: a creation that
+            # completes on the target only after this call gave up on it must
+            # never share its NSID with a later volume.
+            my $timeout = nvmet_long_timeout();
+            if ($nsid > $inv->{last_nsid}) {
+                my $reserve = ['zfs', 'set', "$nvmet_last_nsid=$nsid", nvmet_pool($scfg)];
+                nvmet_change(
+                    $scfg, [[$reserve]],
+                    op => 'reserve namespace ID',
+                    timeout => $timeout,
+                );
+            }
+
+            my @identity = map { ('-o', $_) } nvmet_identity($nqn, $nsid, $uuid);
+            my @create_opts = (
+                ($scfg->{sparse} ? ('-s') : ()),
+                ($scfg->{blocksize} ? ('-b', $scfg->{blocksize}) : ()),
+            );
+            my $create =
+                defined($opts{origin})
+                ? ['zfs', 'clone', @identity, $opts{origin}, $dataset]
+                : ['zfs', 'create', @create_opts, @identity, '-V', "$opts{size}k", $dataset];
+            my $res = nvmet_apply(
+                $scfg, [[$create]],
+                op => 'create zvol',
+                timeout => $timeout,
+            );
+            if ($res->{rc}) {
+                # A definite command failure may still have created the zvol.
+                # Only a zvol with this call's identity counts. Ambiguous SSH
+                # outcomes have already aborted the transaction in nvmet_exec.
+                my $check = eval { nvmet_read_zfs($scfg) };
+                nvmet_rethrow_abandoned($@) if $@;
+                my $row = $check ? $check->{volumes}->{$dataset} : undef;
+                die "NVMe target operation 'create zvol' failed: $res->{err}\n"
+                    if !nvmet_has_identity($row, $nqn, $nsid, $uuid);
+            }
+
+            my $verified;
+            eval {
+                nvmet_wait_device($scfg, $dev);
+                my $build_unit = nvmet_build_unit($nqn, $nsid, $uuid, $dev);
+                my $build = nvmet_apply($scfg, [$build_unit], op => 'export namespace');
+                $verified = [nvmet_read($scfg)];
+                my ($vinv, $vcfs) = $verified->@*;
+                my ($vnsid, $vuuid) = _nvmet_owned_identity($vinv, $nqn, $dataset);
+                die "NVMe identity of zvol '$dataset' changed\n"
+                    if $vnsid ne $nsid || $vuuid ne $uuid;
+                if (_nvmet_export_state($vcfs, $nqn, $nsid, $uuid, $dev) ne 'present') {
+                    nvmet_check($build, 'export namespace');
+                    die "NVMe namespace '$nsid' was not exported\n";
+                }
+            };
+            my $error = $@ or return $uuid;
+            nvmet_rethrow_abandoned($error);
+
+            eval {
+                my ($cinv, $ccfs) = $verified ? $verified->@* : nvmet_read($scfg);
+                my $namespaces = nvmet_namespaces($ccfs, $nqn);
+                # The build writes the UUID before it enables the namespace, so an
+                # enabled namespace without our UUID was not built by this call.
+                my @unexport =
+                    map { nvmet_unexport_unit($nqn, $_, $namespaces->{$_}->{enable} eq '1') }
+                    grep {
+                        my $ns = $namespaces->{$_};
+                        $ns->{device_uuid} eq $uuid
+                        || ($ns->{enable} ne '1' && $ns->{device_path} eq $dev)
+                    } sort {
+                    $a <=> $b
+                    } keys $namespaces->%*;
+                nvmet_change($scfg, \@unexport, op => 'remove namespace') if @unexport;
+                # -o at creation proves that this call created the zvol
+                if (nvmet_has_identity($cinv->{volumes}->{$dataset}, $nqn, $nsid, $uuid)) {
+                    my $destroy = ['zfs', 'destroy', '-r', $dataset];
+                    nvmet_change(
+                        $scfg, [[$destroy]],
+                        op => 'destroy zvol',
+                        timeout => $timeout,
+                    );
+                }
+            };
+            log_warn("failed to clean up zvol '$name': $@") if $@;
+            die $error;
+        },
+    );
+}
+
+# Unexports and destroys an owned zvol. A missing zvol is already destroyed.
+my sub nvmet_destroy_volume($scfg, $name) {
+    my $nqn = nvmet_nqn($scfg);
+    my $dataset = nvmet_dataset($scfg, $name);
+    my $dev = "/dev/zvol/$dataset";
+    my $destroy = ['zfs', 'destroy', '-r', $dataset];
+    my $timeout = nvmet_long_timeout();
+
+    nvmet_locked(
+        $scfg,
+        sub {
+            my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1);
+            return if !exists($inv->{types}->{$dataset});
+            my ($nsid, $uuid) = _nvmet_owned_identity($inv, $nqn, $dataset);
+            my ($unexport, $pre) = ([], undef);
+            if ($cfs->{subsystems}->{$nqn}) {
+                _nvmet_plan_export($cfs, $nqn, $nsid, $uuid, $dev); # refusals only
+                ($unexport, $pre) = _nvmet_plan_unexport($cfs, $nqn, $uuid);
+            }
+            my $units = [$unexport->@*, [$destroy]];
+            my $res = nvmet_apply($scfg, $units, op => 'destroy zvol', timeout => $timeout);
+            return if !$res->{rc};
+            my $error = $res->{err};
+
+            # Decide from a fresh read what the failed call changed.
+            my ($finv, $fcfs) = eval { nvmet_read($scfg) };
+            if (my $read_error = $@) {
+                nvmet_rethrow_abandoned($read_error);
+                die "failed to destroy '$dataset': $error\n";
+            }
+            return if !exists($finv->{types}->{$dataset});
+            my $namespaces = nvmet_namespaces($fcfs, $nqn);
+            my ($current) =
+                grep { $namespaces->{$_}->{device_uuid} eq $uuid } keys $namespaces->%*;
+            if (defined($current)) {
+                die "cannot disable namespace '$current': $error\n"
+                    if $namespaces->{$current}->{enable} eq '1';
+                die "cannot remove namespace '$current': $error\n" if !$pre || !$pre->{enabled};
+                my $enable = [nvmet_enable_unit($nqn, $current, $dev)];
+                my $restore = nvmet_apply($scfg, $enable, op => 'enable namespace');
+                die "cannot remove namespace '$current': $error"
+                    . ($restore->{rc} ? "; cannot restore namespace: $restore->{err}" : '') . "\n";
+            }
+
+            # The namespace is gone and the zvol is still there, typically busy.
+            for my $attempt (1 .. 5) {
+                _sleep(1);
+                my $retry =
+                    nvmet_apply($scfg, [[$destroy]], op => 'destroy zvol', timeout => $timeout);
+                return if !$retry->{rc};
+                $error = $retry->{err};
+            }
+            die "failed to destroy '$dataset': $error\n" if !$pre;
+            eval {
+                my (undef, $rcfs) = nvmet_read($scfg);
+                my ($units) = _nvmet_plan_export($rcfs, $nqn, $nsid, $uuid, $dev);
+                nvmet_change($scfg, $units, op => 'restore namespace');
+            };
+            die "failed to destroy '$dataset': $error; namespace restoration also failed: $@"
+                if $@;
+            die "failed to destroy '$dataset': $error; namespace restored\n";
+        },
+    );
+    return;
+}
+
+# Rolls an owned zvol back and always exports it again with the same identity.
+my sub nvmet_rollback_volume($scfg, $name, $snap) {
+    my $snapshot = nvmet_snapshot($scfg, $name, $snap);
+    my $nqn = nvmet_nqn($scfg);
+    my $dataset = nvmet_dataset($scfg, $name);
+    my $dev = "/dev/zvol/$dataset";
+
+    my $timeout = nvmet_long_timeout();
+
+    nvmet_locked(
+        $scfg,
+        sub {
+            my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1);
+            my ($nsid, $uuid) = _nvmet_owned_identity($inv, $nqn, $dataset);
+            die "NVMe subsystem does not exist\n" if !$cfs->{subsystems}->{$nqn};
+            _nvmet_plan_export($cfs, $nqn, $nsid, $uuid, $dev); # refusals only
+            my ($unexport, $pre) = _nvmet_plan_unexport($cfs, $nqn, $uuid);
+            my $set_identity = ['zfs', 'set', nvmet_identity($nqn, $nsid, $uuid), $dataset];
+            my $rollback = ['zfs', 'rollback', $snapshot];
+            my $res = nvmet_apply(
+                $scfg, [$unexport->@*, [$rollback], [$set_identity]],
+                op => 'rollback zvol',
+                timeout => $timeout,
+            );
+            my $error;
+            $error = "NVMe target operation 'rollback zvol' failed: $res->{err}" if $res->{rc};
+
+            # Always export the volume again, from a fresh read if anything failed.
+            eval {
+                my ($rinv, $rcfs) =
+                    $error ? nvmet_read($scfg) : ($inv, nvmet_without_namespace($cfs, $nqn, $pre));
+                nvmet_wait_device($scfg, $dev);
+                my @units;
+                push @units, [$set_identity]
+                    if !nvmet_has_identity($rinv->{volumes}->{$dataset}, $nqn, $nsid, $uuid);
+                my ($export) = _nvmet_plan_export($rcfs, $nqn, $nsid, $uuid, $dev);
+                push @units, $export->@*;
+                nvmet_change($scfg, \@units, op => 'restore namespace', timeout => $timeout);
+            };
+            chomp(my $restore_error = $@);
+
+            die "rollback failed: $error; namespace restoration failed: $restore_error\n"
+                if defined($error) && $restore_error;
+            die "namespace restoration failed after rollback: $restore_error\n"
+                if $restore_error;
+            die "$error\n" if defined($error);
+        },
+    );
+    return;
+}
+
+# Renames vm-* to base-*, exports it under the same identity and takes the
+# __base__ snapshot. A failed conversion is undone from a fresh read.
+my sub nvmet_template_volume($scfg, $name) {
+    my $nqn = nvmet_nqn($scfg);
+    my $old = nvmet_dataset($scfg, $name);
+    my $new = _nvmet_template_name($old);
+    my ($old_dev, $new_dev) = ("/dev/zvol/$old", "/dev/zvol/$new");
+    my $timeout = nvmet_long_timeout();
+
+    nvmet_locked(
+        $scfg,
+        sub {
+            my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1);
+            my ($nsid, $uuid) = _nvmet_owned_identity($inv, $nqn, $old);
+            die "template zvol '$new' already exists\n" if exists($inv->{types}->{$new});
+            die "NVMe subsystem does not exist\n" if !$cfs->{subsystems}->{$nqn};
+            _nvmet_plan_export($cfs, $nqn, $nsid, $uuid, $old_dev); # refusals only
+            my ($unexport) = _nvmet_plan_unexport($cfs, $nqn, $uuid);
+
+            eval {
+                my $rename = [$unexport->@*, [['zfs', 'rename', $old, $new]]];
+                nvmet_change($scfg, $rename, op => 'rename zvol', timeout => $timeout);
+                nvmet_wait_device($scfg, $new_dev);
+                my $build = nvmet_build_unit($nqn, $nsid, $uuid, $new_dev);
+                my $template = [$build, [['zfs', 'snapshot', "$new\@__base__"]]];
+                nvmet_change($scfg, $template, op => 'create template', timeout => $timeout);
+            };
+            my $error = $@ or return;
+            nvmet_rethrow_abandoned($error);
+            chomp($error);
+
+            eval {
+                my ($rinv, $rcfs) = nvmet_read($scfg);
+                if (exists($rinv->{types}->{$new})) {
+                    my ($units) = _nvmet_plan_unexport($rcfs, $nqn, $uuid);
+                    my $back = [$units->@*, [['zfs', 'rename', $new, $old]]];
+                    nvmet_change($scfg, $back, op => 'rename zvol back', timeout => $timeout);
+                    (undef, $rcfs) = nvmet_read($scfg);
+                } elsif (!exists($rinv->{types}->{$old})) {
+                    die "zvol '$old' does not exist\n";
+                }
+                nvmet_wait_device($scfg, $old_dev);
+                my ($units) = _nvmet_plan_export($rcfs, $nqn, $nsid, $uuid, $old_dev);
+                nvmet_change($scfg, $units, op => 'restore namespace');
+            };
+            die "template conversion failed: $error; restoration failed: $@" if $@;
+            die "template conversion failed: $error\n";
+        },
+    );
+    return;
+}
+
+# Grows an owned zvol and lets an exported namespace pick up the new size.
+my sub nvmet_resize_volume($scfg, $name, $size) {
+    my $nqn = nvmet_nqn($scfg);
+    my $dataset = nvmet_dataset($scfg, $name);
+    die "internal error: invalid zvol size\n" if $size !~ $RE_UNSIGNED_INTEGER;
+
+    # The volume is checked before it changes; under the lock, the namespace
+    # read here is still the one to revalidate after the resize. It is
+    # revalidated in the same call, so whenever the zvol grows, also after a
+    # lost reply, the hosts see the new size. A namespace that is not enabled
+    # reads the new size when it is enabled.
+    nvmet_locked(
+        $scfg,
+        sub {
+            my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1);
+            my (undef, $uuid) = _nvmet_owned_identity($inv, $nqn, $dataset);
+            my $namespaces = nvmet_namespaces($cfs, $nqn);
+            my @matches =
+                grep { $namespaces->{$_}->{device_uuid} eq $uuid } keys $namespaces->%*;
+            die "duplicate namespace UUID '$uuid'\n" if @matches > 1;
+            my @resize = (['zfs', 'set', "volsize=${size}k", $dataset]);
+            push @resize, nvmet_write(nvmet_ns_path($nqn, $matches[0]) . '/revalidate_size', 1)
+                if @matches && $namespaces->{ $matches[0] }->{enable} eq '1';
+            nvmet_change(
+                $scfg, [\@resize],
+                op => 'resize zvol',
+                timeout => nvmet_long_timeout(),
+            );
+        },
+    );
+    return;
+}
+
+# Removes the subsystem once it has no volumes and namespaces, then the hosts
+# that no subsystem links any more. Ports are never removed.
+sub _nvmet_delete_target($scfg) {
+    my $nqn = nvmet_nqn($scfg);
+    my $hostnqns = parse_nvme_host_nqns($scfg->{'nvme-host-nqns'}, 1) // [];
+
+    nvmet_locked(
+        $scfg,
+        sub {
+            my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1);
+            my ($units, $candidates) = _nvmet_plan_delete_target($inv, $cfs, $nqn, $hostnqns);
+            my $fresh = $cfs;
+            if ($units->@*) {
+                my $res = nvmet_apply($scfg, $units, op => 'remove NVMe subsystem');
+                (undef, $fresh) = eval { nvmet_read($scfg) };
+                my $read_error = $@;
+                nvmet_rethrow_abandoned($read_error) if $read_error;
+                die "cannot remove NVMe subsystem '$nqn': $res->{err}\n"
+                    if $res->{rc} && (!$fresh || $fresh->{subsystems}->{$nqn});
+                die $read_error if !$fresh;
+            }
+            my $orphans = _nvmet_plan_orphan_hosts($fresh, $candidates);
+            return if !$orphans->@*;
+            my $res = nvmet_apply($scfg, $orphans, op => 'remove orphan NVMe hosts');
+            log_warn("could not remove orphan NVMe host(s): $res->{err}") if $res->{rc};
+        },
+    );
+    return;
+}
+
+# ---------------------------------------------------------------------------
+# Configuration
+# ---------------------------------------------------------------------------
+
+sub verify_nvme_nqn($value, $noerr = undef) {
+
+    if (
+        length($value) > 223
+        || $value !~ $RE_NQN
+    ) {
+        return undef if $noerr;
+        die "value is not a valid NVMe qualified name\n";
+    }
+
+    return $value;
+}
+
+sub parse_nvme_portals($value, $noerr = undef) {
+    my $result = [];
+    my $seen = {};
+
+    for my $entry (split(/,/, $value // '')) {
+        $entry = trim($entry);
+        my ($address, $port, $family);
+        if ($entry =~ $RE_IPV6_PORTAL) {
+            ($address, $port, $family) = ($+{address}, $+{port} // 4420, 'ipv6');
+        } elsif ($entry =~ $RE_IPV4_PORTAL) {
+            ($address, $port, $family) = ($+{address}, $+{port} // 4420, 'ipv4');
+        } else {
+            return undef if $noerr;
+            die "invalid NVMe/TCP portal '$entry'\n";
+        }
+
+        if (
+            !PVE::JSONSchema::pve_verify_ip($address, 1)
+            || ($family eq 'ipv4' && index($address, ':') >= 0)
+            || ($family eq 'ipv6' && index($address, ':') < 0)
+            || $port < 1
+            || $port > 65535
+        ) {
+            return undef if $noerr;
+            die "invalid NVMe/TCP portal '$entry'\n";
+        }
+        $address = nvmet_canonical_address($address) if $family eq 'ipv6';
+        # one listener: compare like _assert_unique_target
+        my $id = join(
+            "\0",
+            nvmet_listener_address({ family => $family, address => $address }),
+            int($port),
+        );
+        if ($seen->{$id}++) {
+            return undef if $noerr;
+            die "duplicate NVMe/TCP portal '$entry'\n";
+        }
+        push $result->@*,
+            {
+                address => $address,
+                port => int($port),
+                family => $family,
+            };
+        if (scalar($result->@*) > $max_paths) {
+            return undef if $noerr;
+            die "at most $max_paths NVMe/TCP portals are supported\n";
+        }
+    }
+
+    if (!$result->@*) {
+        return undef if $noerr;
+        die "at least one NVMe/TCP portal is required\n";
+    }
+
+    return $result;
+}
+
+my sub verify_nvme_portals($value, $noerr = undef) {
+    return undef if !parse_nvme_portals($value, $noerr);
+    return $value;
+}
+
+sub parse_nvme_host_ifaces($value, $noerr = undef) {
+    my $result = [];
+
+    for my $iface (split(/,/, $value // '')) {
+        $iface = trim($iface);
+        # the kernel never names an interface '.' or '..'
+        if (
+            length($iface) < 1
+            || length($iface) > 15
+            || $iface !~ $RE_HOST_IFACE
+            || $iface eq '.'
+            || $iface eq '..'
+        ) {
+            return undef if $noerr;
+            die "invalid NVMe/TCP host interface '$iface'\n";
+        }
+        push $result->@*, $iface;
+    }
+
+    if (!$result->@*) {
+        return undef if $noerr;
+        die "at least one NVMe/TCP host interface is required\n";
+    }
+
+    return $result;
+}
+
+sub parse_nvme_host_nqns($value, $noerr = undef) {
+    my $result = [];
+    my $seen = {};
+
+    for my $hostnqn (split(/,/, $value // '')) {
+        $hostnqn = trim($hostnqn);
+        if (!verify_nvme_nqn($hostnqn, 1) || $seen->{$hostnqn}++) {
+            return undef if $noerr;
+            die "invalid or duplicate NVMe host NQN '$hostnqn'\n";
+        }
+        push $result->@*, $hostnqn;
+        if (scalar($result->@*) > $max_hosts) {
+            return undef if $noerr;
+            die "at most $max_hosts NVMe host NQNs are supported\n";
+        }
+    }
+
+    if (!$result->@*) {
+        return undef if $noerr;
+        die "at least one NVMe host NQN is required\n";
+    }
+
+    return $result;
+}
+
+my sub verify_nvme_host_ifaces($value, $noerr = undef) {
+    return undef if !parse_nvme_host_ifaces($value, $noerr);
+    return $value;
+}
+
+my sub verify_nvme_host_nqns($value, $noerr = undef) {
+    return undef if !parse_nvme_host_nqns($value, $noerr);
+    return $value;
+}
+
+sub _configured_portals($scfg) {
+    my $portals = parse_nvme_portals($scfg->{'nvme-portals'});
+    my $ifaces = parse_nvme_host_ifaces($scfg->{'nvme-host-ifaces'});
+    die "nvme-host-ifaces must contain one interface for each nvme-portals entry\n"
+        if scalar($ifaces->@*) != scalar($portals->@*);
+
+    for (my $i = 0; $i < scalar($portals->@*); $i++) {
+        $portals->[$i]->{host_iface} = $ifaces->[$i];
+    }
+    return $portals;
+}
+
+sub _local_iface_exists($iface) {
+    return -d "/sys/class/net/$iface";
+}
+
+sub _validate_local_ifaces($portals) {
+    for my $portal ($portals->@*) {
+        my $iface = $portal->{host_iface};
+        die "NVMe/TCP host interface '$iface' does not exist on this node\n"
+            if !_local_iface_exists($iface);
+    }
+}
+
+PVE::JSONSchema::register_format('pve-storage-nvme-nqn', \&verify_nvme_nqn);
+PVE::JSONSchema::register_format('pve-storage-nvme-portals', \&verify_nvme_portals);
+PVE::JSONSchema::register_format('pve-storage-nvme-host-ifaces', \&verify_nvme_host_ifaces);
+PVE::JSONSchema::register_format('pve-storage-nvme-host-nqns', \&verify_nvme_host_nqns);
+
+sub type($class) {
+    return 'zfsnvme';
+}
+
+sub plugindata($class) {
+    return {
+        content => [{ images => 1 }, { images => 1 }],
+        'sensitive-properties' => { 'dhchap-key' => 1 },
+    };
+}
+
+sub properties($class) {
+    return {
+        subsysnqn => {
+            description => "NVMe subsystem qualified name.",
+            type => 'string',
+            format => 'pve-storage-nvme-nqn',
+        },
+        'nvme-portals' => {
+            description =>
+                "Comma-separated NVMe/TCP target IP addresses, at most 16. Put IPv6 addresses in"
+                . " brackets. The default TCP port is 4420.",
+            type => 'string',
+            format => 'pve-storage-nvme-portals',
+            maxLength => 2048,
+        },
+        'nvme-host-ifaces' => {
+            description =>
+                "Comma-separated local interfaces, in portal order. Interface names must be"
+                . " identical on every cluster node.",
+            type => 'string',
+            format => 'pve-storage-nvme-host-ifaces',
+            maxLength => 512,
+        },
+        'nvme-host-nqns' => {
+            description =>
+                "Comma-separated /etc/nvme/hostnqn values for every cluster node allowed to use"
+                . " this storage, at most 64. Host NQNs can only be added.",
+            type => 'string',
+            format => 'pve-storage-nvme-host-nqns',
+            maxLength => 8192,
+        },
+        'dhchap-key' => {
+            description =>
+                "NVMe DH-HMAC-CHAP key in secret representation format (DHHC-1). Required when the"
+                . " storage is created; it cannot be changed.",
+            type => 'string',
+            maxLength => 256,
+        },
+        'nvme-iopolicy' => {
+            description => "Native NVMe multipath I/O policy.",
+            type => 'string',
+            enum => ['numa', 'round-robin', 'queue-depth'],
+            default => 'round-robin',
+        },
+        'nvme-keep-alive-tmo' => {
+            description => "NVMe keep-alive timeout in seconds.",
+            type => 'integer',
+            minimum => 1,
+            maximum => 120,
+            default => 5,
+        },
+        'nvme-reconnect-delay' => {
+            description => "Delay between NVMe reconnect attempts in seconds.",
+            type => 'integer',
+            minimum => 1,
+            maximum => 120,
+            default => 2,
+        },
+        'nvme-ctrl-loss-tmo' => {
+            description => "Time to keep retrying a lost NVMe controller in seconds.",
+            type => 'integer',
+            minimum => -1,
+            maximum => 86400,
+            default => 600,
+        },
+        'nvme-fast-io-fail-tmo' => {
+            description =>
+                "Optional time before failing I/O on a reconnecting NVMe controller. Unset queues"
+                . " I/O until controller loss timeout.",
+            type => 'integer',
+            minimum => 0,
+            maximum => 86400,
+            optional => 1,
+        },
+        'nvme-nr-io-queues' => {
+            description => "Number of NVMe/TCP I/O queues per controller.",
+            type => 'integer',
+            minimum => 1,
+            maximum => 1024,
+            optional => 1,
+        },
+    };
+}
+
+sub options($class) {
+    return {
+        server => { fixed => 1 },
+        subsysnqn => { fixed => 1 },
+        'nvme-portals' => { fixed => 1 },
+        'nvme-host-ifaces' => {},
+        'nvme-host-nqns' => {},
+        pool => { fixed => 1 },
+        blocksize => { fixed => 1 },
+        sparse => { optional => 1 },
+        'dhchap-key' => { optional => 1 },
+        'nvme-iopolicy' => { optional => 1 },
+        'nvme-keep-alive-tmo' => { optional => 1 },
+        'nvme-reconnect-delay' => { optional => 1 },
+        'nvme-ctrl-loss-tmo' => { optional => 1 },
+        'nvme-fast-io-fail-tmo' => { optional => 1 },
+        'nvme-nr-io-queues' => { optional => 1 },
+        nodes => { optional => 1 },
+        disable => { optional => 1 },
+        content => { optional => 1 },
+        bwlimit => { optional => 1 },
+    };
+}
+
+sub _validate_fail_fast_timeout($config, $default_ctrl_loss_tmo = undef) {
+    my $fast = $config->{'nvme-fast-io-fail-tmo'};
+    return if !defined($fast);
+
+    my $ctrl = $config->{'nvme-ctrl-loss-tmo'};
+    $ctrl = $default_ctrl_loss_tmo if !defined($ctrl);
+    return if !defined($ctrl) || $ctrl < 0;
+
+    die "nvme-fast-io-fail-tmo must not exceed nvme-ctrl-loss-tmo\n"
+        if $fast > $ctrl;
+}
+
+# Defaults are not injected here: this also runs when storage.cfg is parsed.
+# Every use site applies the schema default itself.
+sub check_config($class, $section_id, $config, $create, $skip_schema_check = undef) {
+    nvmet_pool($config) if defined($config->{pool});
+    verify_nvme_nqn($config->{subsysnqn}) if defined($config->{subsysnqn});
+    parse_nvme_portals($config->{'nvme-portals'}) if defined($config->{'nvme-portals'});
+    if (defined($config->{'nvme-host-ifaces'}) && defined($config->{'nvme-portals'})) {
+        _configured_portals($config);
+    } elsif (defined($config->{'nvme-host-ifaces'})) {
+        parse_nvme_host_ifaces($config->{'nvme-host-ifaces'});
+    }
+    parse_nvme_host_nqns($config->{'nvme-host-nqns'})
+        if defined($config->{'nvme-host-nqns'});
+    _validate_fail_fast_timeout($config, $create ? 600 : undef);
+    return $class->SUPER::check_config($section_id, $config, $create, $skip_schema_check);
+}
+
+# ---------------------------------------------------------------------------
+# ZFS helpers (same formats as the other ZFS plugins)
+# ---------------------------------------------------------------------------
+
+my sub zfs_request($scfg, $timeout, $method, @args) {
+    $timeout = PVE::RPCEnvironment->is_worker() ? 60 * 60 : 10 if !$timeout;
+    my $change = $method ne 'get' && $method ne 'list';
+    my $res = nvmet_exec(
+        $scfg, [['zfs', $method, @args]],
+        op => "zfs $method",
+        timeout => $timeout,
+        change => $change,
+    );
+    die "zfs error on '" . nvmet_server($scfg) . "': $res->{err}\n" if $res->{rc};
+    return nvmet_output($res);
+}
+
+# The direct children of the pool with PVE volume names. The plugin creates
+# and owns zvols only; a filesystem child holds its name.
+my sub zfs_parse_zvol_list($text, $pool) {
+    my $list = [];
+    for my $line (split /\n/, $text // '') {
+        my ($dataset, $size, $origin, $type) = split(/\s+/, $line);
+        next if !defined($type) || ($type ne 'volume' && $type ne 'filesystem');
+        my @parts = split /\//, $dataset;
+        next if @parts < 2;
+        my $name = pop @parts;
+        next if join('/', @parts) ne $pool;
+        next if $name !~ $RE_ZVOL_OWNER;
+        push $list->@*,
+            {
+                name => $name,
+                owner => $+{owner},
+                type => $type,
+                ($type eq 'volume' ? (size => $size + 0) : ()),
+                ($origin ne '-' ? (origin => $origin) : ()),
+            };
+    }
+    return $list;
+}
+
+# The pool's direct children with PVE volume names, so that a name is never
+# handed out twice; with $owned_only, the zvols owned by this storage's NVMe
+# subsystem.
+my sub zfs_list_zvol($scfg, $owned_only) {
+    my $pool = nvmet_pool($scfg);
+    my $text = zfs_request(
+        $scfg,
+        10,
+        'list',
+        '-o',
+        'name,volsize,origin,type',
+        '-t',
+        'volume,filesystem',
+        '-d1',
+        '-Hp',
+        $pool,
+    );
+    my $list = {};
+    my $prefix = "$pool/";
+    for my $zvol (zfs_parse_zvol_list($text, $pool)->@*) {
+        next if $owned_only && $zvol->{type} ne 'volume';
+        my $parent = $zvol->{origin};
+        # an origin in the pool is named relative to it, like the ZFS plugins do
+        $parent = substr($parent, length($prefix)) if $parent && index($parent, $prefix) == 0;
+        $list->{ $zvol->{name} } = {
+            name => $zvol->{name},
+            size => $zvol->{size},
+            parent => $parent,
+            format => 'raw',
+            vmid => $zvol->{owner},
+        };
+    }
+    return $list if !$owned_only;
+
+    my $properties = zfs_request(
+        $scfg,
+        10,
+        'get',
+        '-H',
+        '-d',
+        '1',
+        '-o',
+        'name,value,source',
+        'proxmox:nvme-subsys',
+        $pool,
+    );
+    my $owned = {};
+    for my $line (split(/\n/, $properties)) {
+        my ($dataset, $nqn, $source) = split(/\t/, $line, 3);
+        next if !defined($source) || ($source ne 'local' && $source ne 'received');
+        next if !defined($nqn) || $nqn ne $scfg->{subsysnqn};
+        next if index($dataset, $prefix) != 0;
+        $owned->{ substr($dataset, length($prefix)) } = 1;
+    }
+    for my $name (keys $list->%*) {
+        delete $list->{$name} if !$owned->{$name};
+    }
+    return $list;
+}
+
+my sub zfs_get_properties($scfg, $properties, $dataset, $timeout = undef) {
+    my $text = zfs_request($scfg, $timeout, 'get', '-o', 'value', '-Hp', $properties, $dataset);
+    my @values = split /\n/, $text;
+    return wantarray ? @values : $values[0];
+}
+
+my sub zfs_get_sorted_snapshot_list($scfg, $name, $sort_params) {
+    return [
+        map { s/^.*\@//r } split /\n/,
+        zfs_request(
+            $scfg,
+            undef,
+            'list',
+            '-H',
+            '-r',
+            '-t',
+            'snapshot',
+            '-o',
+            'name',
+            $sort_params->@*,
+            nvmet_dataset($scfg, $name),
+        ),
+    ];
+}
+
+# ---------------------------------------------------------------------------
+# Storage API: volumes
+# ---------------------------------------------------------------------------
+
+sub parse_volname($class, $volname) {
+    if ($volname =~ $RE_VOLNAME) {
+        my $format = ($+{type} eq 'subvol' || $+{type} eq 'basevol') ? 'subvol' : 'raw';
+        my $is_base = $+{type} eq 'base' || $+{type} eq 'basevol';
+        return ('images', $+{name}, $+{vmid}, $+{base}, $+{base_vmid}, $is_base, $format);
+    }
+    die "unable to parse zfs volume name '$volname'\n";
+}
+
+sub list_images($class, $storeid, $scfg, $vmid = undef, $vollist = undef, $cache = undef) {
+    my $res = [];
+    for my $info (values zfs_list_zvol($scfg, 1)->%*) {
+        my $volname =
+            $info->{parent} && $info->{parent} =~ $RE_BASE_SNAPSHOT
+            ? "$storeid:$+{base}/$info->{name}"
+            : "$storeid:$info->{name}";
+        next
+            if $vollist
+            ? !grep { $_ eq $volname } $vollist->@*
+            : defined($vmid) && $info->{vmid} ne $vmid;
+        $info->{volid} = $volname;
+        push $res->@*, $info;
+    }
+    return $res;
+}
+
+# Names are allocated against every direct child of the pool, owned or not, so
+# a leftover zvol that list_images hides is never handed out again.
+sub find_free_diskname($class, $storeid, $scfg, $vmid, $fmt = undef, $add_fmt_suffix = undef) {
+    my $disk_list = [map { "$storeid:$_" } sort keys zfs_list_zvol($scfg, 0)->%*];
+    return PVE::Storage::Plugin::get_next_vm_diskname(
+        $disk_list, $storeid, $vmid, $fmt, $scfg, $add_fmt_suffix,
+    );
+}
+
+sub status($class, $storeid, $scfg, $cache = undef) {
+    my ($available, $used) =
+        eval { zfs_get_properties($scfg, 'available,used', nvmet_pool($scfg)) };
+    if (my $err = $@) {
+        warn "storage '$storeid': $err";
+        return (0, 0, 0, 0);
+    }
+    if (
+        !defined($available)
+        || !defined($used)
+        || $available !~ $RE_UNSIGNED_INTEGER
+        || $used !~ $RE_UNSIGNED_INTEGER
+    ) {
+        warn "unexpected ZFS pool usage for storage '$storeid'\n";
+        return (0, 0, 0, 0);
+    }
+    return ($available + $used, $available, $used, 1);
+}
+
+sub volume_size_info($class, $scfg, $storeid, $volname, $timeout = undef) {
+    my (undef, $name, undef, $parent, undef, undef, $format) = $class->parse_volname($volname);
+    die "volume_size_info requires a ZFS volume\n" if $format ne 'raw';
+    my ($size, $used) = zfs_get_properties(
+        $scfg, 'volsize,usedbydataset', nvmet_dataset($scfg, $name), $timeout,
+    );
+    die "Could not get zfs volume size\n" if !defined($size) || $size !~ $RE_UNSIGNED_INTEGER;
+    $used = defined($used) && $used =~ $RE_UNSIGNED_INTEGER ? $used + 0 : 0;
+    return wantarray ? ($size + 0, 'raw', $used, $parent) : $size + 0;
+}
+
+sub volume_snapshot($class, $scfg, $storeid, $volname, $snap) {
+    my (undef, $name, undef, undef, undef, undef, $format) = $class->parse_volname($volname);
+    die "volume_snapshot requires a ZFS volume\n" if $format ne 'raw';
+    my $snapshot = nvmet_snapshot($scfg, $name, $snap);
+    nvmet_locked($scfg, sub { zfs_request($scfg, undef, 'snapshot', $snapshot) });
+    return;
+}
+
+sub volume_snapshot_delete($class, $scfg, $storeid, $volname, $snap, $running = undef) {
+    my $name = ($class->parse_volname($volname))[1];
+    my $snapshot = nvmet_snapshot($scfg, $name, $snap);
+    nvmet_locked($scfg, sub { zfs_request($scfg, undef, 'destroy', $snapshot) });
+    return;
+}
+
+sub volume_rollback_is_possible($class, $scfg, $storeid, $volname, $snap, $blockers = undef) {
+    my $name = ($class->parse_volname($volname))[1];
+    nvmet_snapshot($scfg, $name, $snap); # validation only
+    my $found;
+    $blockers //= [];
+    for my $snapshot (zfs_get_sorted_snapshot_list($scfg, $name, ['-s', 'creation'])->@*) {
+        $found = 1 if $snapshot eq $snap;
+        push $blockers->@*, $snapshot if $found && $snapshot ne $snap;
+    }
+    die "can't rollback, snapshot '$snap' does not exist on '${storeid}:${volname}'\n" if !$found;
+    die "can't rollback, '$snap' is not most recent snapshot on '${storeid}:${volname}'\n"
+        if $blockers->@*;
+    return 1;
+}
+
+sub volume_snapshot_rollback($class, $scfg, $storeid, $volname, $snap) {
+    my (undef, $name, undef, undef, undef, undef, $format) = $class->parse_volname($volname);
+    die "snapshot rollback requires a ZFS volume\n" if $format ne 'raw';
+    nvmet_rollback_volume($scfg, $name, $snap);
+}
+
+sub volume_snapshot_info($class, $scfg, $storeid, $volname) {
+    my $name = ($class->parse_volname($volname))[1];
+    my $info = {};
+    my $text = zfs_request(
+        $scfg,
+        undef,
+        'list',
+        '-Hp',
+        '-r',
+        '-t',
+        'snapshot',
+        '-o',
+        'name,guid,creation',
+        nvmet_dataset($scfg, $name),
+    );
+    for my $line (split /\n/, $text) {
+        my ($snapshot, $guid, $creation) = split /\s+/, $line;
+        $snapshot =~ s/^.*\@//;
+        $info->{$snapshot} = { id => $guid, timestamp => $creation };
+    }
+    return $info;
+}
+
+sub free_image($class, $storeid, $scfg, $volname, $is_base = undef, $format = undef) {
+    my (undef, $name, undef, undef, undef, undef, $volformat) = $class->parse_volname($volname);
+    die "free_image requires a ZFS volume\n" if $volformat ne 'raw';
+    nvmet_destroy_volume($scfg, $name);
+    return undef;
+}
+
+sub create_base($class, $storeid, $scfg, $volname) {
+    my (undef, $name, undef, $basename, undef, $is_base, $format) = $class->parse_volname($volname);
+    die "create_base not possible with base image\n" if $is_base;
+    die "create_base requires a ZFS volume\n" if $format ne 'raw';
+    my $newname = $name =~ s/^vm-/base-/r;
+    nvmet_template_volume($scfg, $name);
+    return $basename ? "$basename/$newname" : $newname;
+}
+
+sub clone_image($class, $scfg, $storeid, $volname, $vmid, $snap = undef) {
+    $snap ||= '__base__';
+    my (undef, $basename, undef, undef, undef, $is_base, $format) = $class->parse_volname($volname);
+    die "clone_image only works on base images\n" if !$is_base;
+    die "clone_image requires a ZFS volume\n" if $format ne 'raw';
+    my $origin = nvmet_snapshot($scfg, $basename, $snap);
+    my $name = $class->find_free_diskname($storeid, $scfg, $vmid, $format);
+    nvmet_create_volume($scfg, $name, origin => $origin);
+    return "$basename/$name";
+}
+
+sub alloc_image($class, $storeid, $scfg, $vmid, $fmt, $name, $size) {
+    die "unsupported format '$fmt'" if $fmt ne 'raw';
+    die "illegal name '$name' - should be 'vm-$vmid-*'\n"
+        if $name && index($name, "vm-$vmid-") != 0;
+    my $volname = $name || $class->find_free_diskname($storeid, $scfg, $vmid, $fmt);
+    $size += (1024 - $size % 1024) % 1024;
+    nvmet_create_volume($scfg, $volname, size => $size);
+    return $volname;
+}
+
+sub volume_resize($class, $scfg, $storeid, $volname, $size, $running = undef, $snapname = undef) {
+    # QEMU can resize a host_device once the device itself has grown, but the
+    # plugin cannot yet wait for the local NVMe namespace to report the new
+    # size. Until it can, refuse online resize before changing the zvol, so a
+    # running VM never sees a size different from its configuration.
+    die "online resize is not supported for NVMe/TCP block devices; stop the VM first\n"
+        if $running;
+    die "resizing a snapshot is not supported for $class\n" if $snapname;
+    my $new_size = int($size / 1024);
+    $new_size += (1024 - $new_size % 1024) % 1024;
+    my $name = ($class->parse_volname($volname))[1];
+    nvmet_resize_volume($scfg, $name, $new_size);
+    return $new_size;
+}
+
+sub volume_has_feature(
+    $class, $scfg, $feature, $storeid, $volname,
+    $snapname = undef,
+    $running = undef,
+    $opts = undef,
+) {
+    my $features = {
+        snapshot => { current => 1, snap => 1 },
+        clone => { base => 1 },
+        template => { current => 1 },
+        copy => { base => 1, current => 1 },
+    };
+    my $is_base = ($class->parse_volname($volname))[5];
+    my $key = $snapname ? 'snap' : $is_base ? 'base' : 'current';
+    return $features->{$feature}->{$key};
+}
+
+# No stream transfers or renames; the storage has no path, so the base class
+# offers no export or import format either.
+sub volume_export($class, @args) {
+    die "ZFS stream export is not supported for NVMe/TCP storage\n";
+}
+
+sub volume_import($class, @args) {
+    die "ZFS stream import is not supported for NVMe/TCP storage\n";
+}
+
+sub rename_volume($class, @args) {
+    die "renaming volumes is not supported for NVMe/TCP storage\n";
+}
+
+sub rename_snapshot($class, @args) {
+    die "rename_snapshot is not supported for $class";
+}
+
+# ---------------------------------------------------------------------------
+# Secrets and configuration hooks
+# ---------------------------------------------------------------------------
+
+# Each storage needs its own NVMe/TCP listener (address family, address and
+# port), and a disabled storage keeps its listener. nvmet answers a connect to
+# a listener that does not publish the subsystem yet with DNR, and the host
+# then deletes the controller whatever its loss timeout: after a target
+# restart, the storage published second would lose its paths. Storages that
+# reach the same data address must spell the server alike, as it names the
+# target lock.
+sub _assert_unique_target($storeid, $scfg, $cfg = undef) {
+    $cfg //= PVE::Storage::config();
+    my (%listeners, %addresses);
+    for my $portal (parse_nvme_portals($scfg->{'nvme-portals'})->@*) {
+        my $address = join("\0", nvmet_listener_address($portal));
+        $listeners{"$address\0$portal->{port}"} = 1;
+        $addresses{$address} = 1;
+    }
+    my $ids = $cfg->{ids} // {};
+    for my $other_id (sort keys $ids->%*) {
+        next if $other_id eq $storeid;
+        my $other = $ids->{$other_id};
+        next if ($other->{type} // '') ne 'zfsnvme';
+
+        die "NVMe subsystem NQN is already used by storage '$other_id'\n"
+            if ($other->{subsysnqn} // '') eq $scfg->{subsysnqn};
+        my $other_server = $other->{server} // '';
+        # a storage whose portals do not parse cannot be activated
+        for my $portal ((parse_nvme_portals($other->{'nvme-portals'}, 1) // [])->@*) {
+            my $address = join("\0", nvmet_listener_address($portal));
+            die "NVMe/TCP portal '$portal->{address}' port $portal->{port} is already used by"
+                . " storage '$other_id'; give each storage its own address or port\n"
+                if $listeners{"$address\0$portal->{port}"};
+            die "storage '$other_id' reaches target address '$portal->{address}' through server"
+                . " '$other_server'; use the same server value for both storages\n"
+                if $addresses{$address} && $other_server ne $scfg->{server};
+        }
+        next if $other_server ne $scfg->{server};
+        my ($pool, $other_pool) = ($scfg->{pool}, $other->{pool} // '');
+        die "ZFS pool '$pool' on '$scfg->{server}' is already used by storage '$other_id'\n"
+            if $other_pool eq $pool;
+        die "ZFS pool '$pool' on '$scfg->{server}' overlaps the pool of storage '$other_id'\n"
+            if index("$other_pool/", "$pool/") == 0 || index("$pool/", "$other_pool/") == 0;
+    }
+}
+
+# nvmet keeps one DH-HMAC-CHAP key per host NQN, so storages on the same target
+# whose host NQNs overlap must use the same key.
+sub _assert_shared_host_key($storeid, $scfg, $key, $cfg = undef) {
+    $cfg //= PVE::Storage::config();
+    my %hosts = map { $_ => 1 } (parse_nvme_host_nqns($scfg->{'nvme-host-nqns'}, 1) // [])->@*;
+    my $ids = $cfg->{ids} // {};
+    for my $other_id (sort keys $ids->%*) {
+        next if $other_id eq $storeid;
+        my $other = $ids->{$other_id};
+        next if ($other->{type} // '') ne 'zfsnvme';
+        next if ($other->{server} // '') ne ($scfg->{server} // '');
+        my $other_hosts = parse_nvme_host_nqns($other->{'nvme-host-nqns'}, 1) // [];
+        next if !grep { $hosts{$_} } $other_hosts->@*;
+        my $other_key = file_read_firstline(secret_path($other_id));
+        die "storage '$other_id' uses a different DH-HMAC-CHAP key for the same NVMe host NQNs"
+            . " on '$scfg->{server}'\n"
+            if defined($other_key) && $other_key ne $key;
+    }
+}
+
+# DHHC-1:<hash>:<base64 of key and its little-endian CRC-32>: as generated by
+# `nvme gen-dhchap-key`. Messages never contain the key.
+sub _validate_secret($key) {
+    die "missing NVMe DH-HMAC-CHAP key\n" if !defined($key) || $key eq '';
+    my $invalid = "invalid NVMe DH-HMAC-CHAP key representation\n";
+    die $invalid if $key !~ $RE_DHCHAP_KEY;
+    my ($hash, $encoded) = @+{qw(hash secret)};
+    die $invalid if length($encoded) % 4;
+    my $decoded = decode_base64($encoded);
+    my $length = length($decoded) - 4;
+    my %lengths = ('00' => [32, 48, 64], '01' => [32], '02' => [48], '03' => [64]);
+    die $invalid if !grep { $_ == $length } $lengths{$hash}->@*;
+    die $invalid if pack('V', crc32(substr($decoded, 0, $length))) ne substr($decoded, $length);
+    return $key;
+}
+
+my sub set_secret($storeid, $key) {
+    _validate_secret($key);
+    make_path($secret_dir, { mode => 0700 });
+    file_set_contents(secret_path($storeid), "$key\n", 0600);
+}
+
+my sub get_secret($storeid) {
+    my $key = file_read_firstline(secret_path($storeid));
+    return _validate_secret($key);
+}
+
+# Removes the key file of a storage (a test seam).
+sub _unlink_file($path) {
+    return unlink($path);
+}
+
+my sub delete_secret($storeid) {
+    _unlink_file(secret_path($storeid));
+}
+
+sub on_add_hook($class, $storeid, $scfg, %sensitive) {
+    _configured_portals($scfg);
+    parse_nvme_host_nqns($scfg->{'nvme-host-nqns'});
+    _assert_unique_target($storeid, $scfg);
+    my $key = _validate_secret($sensitive{'dhchap-key'});
+    _assert_shared_host_key($storeid, $scfg, $key);
+    set_secret($storeid, $key);
+    return;
+}
+
+sub on_update_hook_full($class, $storeid, $scfg, $update, $delete = undef, $sensitive = undef) {
+    $sensitive //= {};
+    my %prospective = ($scfg->%*, $update->%*);
+    delete @prospective{ $delete->@* } if $delete;
+    verify_nvme_nqn($prospective{subsysnqn});
+    _configured_portals(\%prospective);
+    my $new_hostnqns = parse_nvme_host_nqns($prospective{'nvme-host-nqns'});
+    my %new_hosts = map { $_ => 1 } $new_hostnqns->@*;
+    for my $old_hostnqn (parse_nvme_host_nqns($scfg->{'nvme-host-nqns'})->@*) {
+        die "removing NVMe host NQN '$old_hostnqn' is not supported; move every volume off"
+            . " the storage and recreate it with a new subsystem NQN and a new DH-HMAC-CHAP"
+            . " key\n"
+            if !$new_hosts{$old_hostnqn};
+    }
+    _validate_fail_fast_timeout(\%prospective, 600);
+    _assert_unique_target($storeid, \%prospective);
+
+    my $old_key = file_read_firstline(secret_path($storeid));
+    my $key = exists($sensitive->{'dhchap-key'}) ? $sensitive->{'dhchap-key'} : $old_key;
+    _validate_secret($key);
+
+    # nvmet stores authentication on the global Host NQN object, which can be
+    # shared by multiple subsystems. Replacing an active key in place can make
+    # unrelated storages unrecoverable on their next reconnect. Until PVE can
+    # coordinate a rolling rotation on every node, fail before mutating the
+    # cluster-wide secret.
+    die "NVMe DH-HMAC-CHAP key rotation is not supported; create a new storage with a new"
+        . " subsystem NQN\n"
+        if defined($old_key) && $key ne $old_key;
+    _assert_shared_host_key($storeid, \%prospective, $key);
+
+    set_secret($storeid, $key) if exists($sensitive->{'dhchap-key'});
+    return;
+}
+
+# ---------------------------------------------------------------------------
+# Local NVMe host side
+# ---------------------------------------------------------------------------
+
+my %slow_path_last; # storeid => time of the last slow-path activation
+
+# Probes of the local host (test seams).
+sub _block_device($path) {
+    return -b $path;
+}
+
+sub _link_target($path) {
+    return readlink($path);
+}
+
+sub _fabrics_device() {
+    return '/dev/nvme-fabrics';
+}
+
+# Runs $code in a child and gives up on it after $timeout seconds. Returns the
+# child's result; the child's errors pass through as exceptions, a timeout
+# becomes one. The bound is not hard: as with PVE::Tools::run_command, a
+# child stuck in the kernel (a write or delete in D state) is only reaped
+# once the kernel returns, and a timeout does not mean that the operation did
+# not happen.
+#
+# A stopped task (TERM, also HUP, INT and QUIT) stays pending in the parent
+# while the child runs, and the child inherits the mask: a connect that ends
+# before PVE escalates the stop to KILL, 5 seconds after the TERM, creates
+# the controller and restricts its secret attributes without interruption,
+# and run_fork_with_timeout, whose own handlers would consume the signal when
+# the child delivers a result, never sees it. Once the child is reaped and the
+# mask is restored, the handler of the worker dies with the task marker, also
+# over a result or an error of the child. ALRM, the timeout, is not blocked.
+# The KILL ends the parent and the child together; a controller that the
+# child had created but not restricted yet keeps its secret attributes
+# world-readable until the next activation restricts them.
+#
+# run_fork_with_timeout is called in list context on purpose: only there does
+# run_with_timeout report a timeout as its second return value; in scalar
+# context the timeout would be a warning and an undef result.
+my sub run_bounded($timeout, $what, $code) {
+    my $previous = POSIX::SigSet->new();
+    sigprocmask(SIG_BLOCK, POSIX::SigSet->new(SIGHUP, SIGINT, SIGQUIT, SIGTERM), $previous)
+        or die "cannot block signals: $!\n";
+    my $outer_warn = $SIG{__WARN__};
+    my ($res, $timed_out) = eval {
+        # Only a child that leaves no result makes run_fork_with_timeout warn
+        # in the parent; that case is an exception below. The child keeps the
+        # handler of the caller, so its own warnings reach the task log.
+        local $SIG{__WARN__} = sub($message) { };
+        PVE::Tools::run_fork_with_timeout(
+            $timeout,
+            sub {
+                local $SIG{__WARN__} = $outer_warn;
+                return $code->();
+            },
+        );
+    };
+    my $error = $@;
+    sigprocmask(SIG_SETMASK, $previous) or die "cannot restore signals: $!\n";
+    if ($error) {
+        rethrow_task_interrupt($error);
+        die $error;
+    }
+    die "$what did not complete within $timeout seconds\n" if $timed_out;
+    die "$what left no result\n" if !defined($res);
+    return $res;
+}
+
+# The connect options of one controller, as the fabrics device takes them,
+# and their names. The values are validated configuration and host identities
+# checked by the caller; the kernel reads each one up to the next ','.
+sub _fabrics_options($scfg, $portal, $hostnqn, $hostid, $key) {
+    my @options = (
+        [transport => 'tcp'],
+        [traddr => $portal->{address}],
+        [trsvcid => $portal->{port}],
+        [host_iface => $portal->{host_iface}],
+        [nqn => $scfg->{subsysnqn}],
+        [hostnqn => $hostnqn],
+        [hostid => $hostid],
+        [dhchap_secret => $key],
+        [keep_alive_tmo => $scfg->{'nvme-keep-alive-tmo'} // 5],
+        [reconnect_delay => $scfg->{'nvme-reconnect-delay'} // 2],
+        [ctrl_loss_tmo => $scfg->{'nvme-ctrl-loss-tmo'} // 600],
+    );
+    # unlike unset, which leaves fast I/O failure off, 0 is a real timeout
+    push @options, [fast_io_fail_tmo => $scfg->{'nvme-fast-io-fail-tmo'}]
+        if defined($scfg->{'nvme-fast-io-fail-tmo'});
+    push @options, [nr_io_queues => $scfg->{'nvme-nr-io-queues'}]
+        if defined($scfg->{'nvme-nr-io-queues'});
+    for my $option (@options) {
+        die "internal error: invalid NVMe connect option '$option->[0]'\n"
+            if !defined($option->[1]) || $option->[1] !~ $RE_FABRICS_VALUE;
+    }
+    return (join(',', map { "$_->[0]=$_->[1]" } @options), [map { $_->[0] } @options]);
+}
+
+# The child of a connect (a test seam): creates exactly one controller with
+# one write of $options to the fabrics device and returns its instance. The
+# write is never repeated, because the kernel may have created the controller
+# even when it reports an error. Errors carry the errno text of the failed
+# operation and never the options, which hold the key.
+sub _fabrics_connect($options, $names) {
+    my $device = _fabrics_device();
+    my $probe;
+    if (!sysopen($probe, $device, O_RDONLY | O_NOFOLLOW)) {
+        die "'$device' does not exist; load the nvme-tcp kernel module\n" if $! == ENOENT;
+        die "cannot open '$device': $!\n";
+    }
+    my @stat = stat($probe);
+    die "'$device' is not a character device\n" if !@stat || !S_ISCHR($stat[2]);
+    # without a controller, a read lists the supported options as patterns
+    my $supported = '';
+    sysread($probe, $supported, 4096) // die "cannot read '$device': $!\n";
+    close($probe);
+    my %supported = map { s/=.*//r => 1 } split(/,/, $supported =~ s/\n\z//r);
+    for my $name ($names->@*) {
+        die "the running kernel does not support the NVMe connect option '$name'\n"
+            if !$supported{$name};
+    }
+
+    # Not PVE::SysFSTools::file_write: the kernel returns the instance of the
+    # new controller on a read of the file descriptor that wrote the options.
+    sysopen(my $fh, $device, O_RDWR | O_NOFOLLOW) or die "cannot open '$device': $!\n";
+    my $written = _fabrics_write($fh, $options);
+    die "connect failed: $!\n" if !defined($written);
+    die "connect failed: short write\n" if $written != length($options);
+    my $result = '';
+    sysread($fh, $result, 128) // die "cannot read the connect result: $!\n";
+    die "unexpected connect result\n" if $result !~ $RE_FABRICS_RESULT;
+    my $instance = int($+{instance});
+    # The kernel creates the secret attributes of a controller world-readable;
+    # the caller keeps a stopped task from interrupting before this, short of
+    # a KILL.
+    _restrict_attr("/sys/class/nvme/nvme$instance/$_") for qw(dhchap_secret dhchap_ctrl_secret);
+    close($fh);
+    return $instance;
+}
+
+# The write that creates a controller (a test seam).
+sub _fabrics_write($fh, $options) {
+    return syswrite($fh, $options);
+}
+
+# Writes a sysfs attribute and returns the errno text of a failure, or undef.
+# PVE::SysFSTools::file_write returns undef without a warning when the open
+# fails, leaving its errno in $!, and 0 after warning "error writing ...:
+# <errno text>" when the write fails; that warning is not passed on. With
+# $missing_ok, an attribute that does not exist, because its controller went
+# away, is not a failure.
+my sub write_attribute($path, $value, $missing_ok = 0) {
+    my $warning;
+    my $written = do {
+        local $SIG{__WARN__} = sub($message) { $warning = $message };
+        PVE::SysFSTools::file_write($path, $value);
+    };
+    return if $written;
+    return if !defined($warning) && $missing_ok && $! == ENOENT;
+    return "$!" if !defined($warning);
+    return $warning =~ s/\A.*: //sr =~ s/\n\z//r;
+}
+
+# Deletes controllers (the body of a child, a test seam): the deletion can
+# block in the kernel. The error names the controller and the errno text; a
+# controller that is already gone counts as deleted.
+sub _delete_controllers($controllers) {
+    for my $controller ($controllers->@*) {
+        my $path = "/sys/class/nvme/$controller->{name}/delete_controller";
+        my $error = write_attribute($path, '1', 1) // next;
+        die "cannot delete NVMe controller '$controller->{name}': $error\n";
+    }
+    return scalar($controllers->@*);
+}
+
+# Every local controller of the subsystem.
+my sub nvme_controllers($nqn) {
+    my $controllers = [];
+    PVE::File::dir_glob_foreach(
+        '/sys/class/nvme',
+        $RE_NVME_CONTROLLER,
+        sub($entry) {
+            my $base = "/sys/class/nvme/$entry";
+            my $subsys = file_read_firstline("$base/subsysnqn");
+            return if !defined($subsys) || $subsys ne $nqn;
+            my $address = file_read_firstline("$base/address") // '';
+            my $traddr = $address =~ $RE_TRADDR ? $+{value} : undef;
+            my $trsvcid = $address =~ $RE_TRSVCID ? $+{value} : undef;
+            my $host_iface = $address =~ $RE_HOST_IFACE_ADDRESS ? $+{value} : undef;
+            push $controllers->@*,
+                {
+                    name => $entry,
+                    state => file_read_firstline("$base/state") // 'unknown',
+                    traddr => nvmet_canonical_address($traddr),
+                    trsvcid => $trsvcid,
+                    host_iface => $host_iface,
+                };
+        },
+    );
+    return $controllers;
+}
+
+# Controllers by "address:port", with canonical IPv6 addresses. A portal has
+# more than one controller while its path moves to another interface.
+my sub controller_states($nqn) {
+    my $states = {};
+    for my $controller (nvme_controllers($nqn)->@*) {
+        next if !defined($controller->{traddr}) || !defined($controller->{trsvcid});
+        push $states->{"$controller->{traddr}:$controller->{trsvcid}"}->@*, $controller;
+    }
+    return $states;
+}
+
+# The multipath namespace devices of the subsystem and their partitions. One
+# dir_glob_foreach per directory level, because PVE::File::dir_glob_regex
+# returns only the first match.
+my sub namespace_devices($nqn) {
+    my $devices = {};
+
+    PVE::File::dir_glob_foreach(
+        '/sys/class/nvme-subsystem',
+        $RE_NVME_SUBSYSTEM,
+        sub($entry) {
+            my $base = "/sys/class/nvme-subsystem/$entry";
+            my $subsys = file_read_firstline("$base/subsysnqn");
+            return if !defined($subsys) || $subsys ne $nqn;
+
+            PVE::File::dir_glob_foreach(
+                $base,
+                $RE_NVME_NAMESPACE,
+                sub($device) {
+                    $devices->{"/dev/$device"} = $device if _block_device("/dev/$device");
+                    PVE::File::dir_glob_foreach(
+                        "/sys/class/block/$device",
+                        $RE_NVME_PARTITION,
+                        sub($partition) {
+                            $devices->{"/dev/$partition"} = $partition
+                                if _block_device("/dev/$partition");
+                        },
+                    );
+                },
+            );
+        },
+    );
+    return $devices;
+}
+
+sub _namespace_openers($nqn) {
+    my $devices = namespace_devices($nqn);
+    return [] if !$devices->%*;
+
+    my %openers;
+    PVE::File::dir_glob_foreach(
+        '/proc',
+        '[0-9]+',
+        sub($pid) {
+            PVE::File::dir_glob_foreach(
+                "/proc/$pid/fd",
+                '[0-9]+',
+                sub($fd) {
+                    my $target = _link_target("/proc/$pid/fd/$fd");
+                    return if !defined($target) || !exists($devices->{$target});
+                    my $comm = eval { file_read_firstline("/proc/$pid/comm") } // 'unknown';
+                    $openers{"$pid:$target"} = "$comm (PID $pid, $target)";
+                },
+            );
+        },
+    );
+
+    # Kernel consumers such as device-mapper do not necessarily keep a userspace
+    # file descriptor open, but expose their dependency in the holders directory.
+    for my $path (keys $devices->%*) {
+        my $device = $devices->{$path};
+        my $holders = "/sys/class/block/$device/holders";
+        PVE::File::dir_glob_foreach(
+            $holders,
+            '[^\\.].*',
+            sub($holder) {
+                $openers{"holder:$device:$holder"} = "$path held by $holder";
+            },
+        );
+    }
+
+    return [sort values %openers];
+}
+
+my sub portal_reachable($portal) {
+    return PVE::Network::tcp_ping($portal->{address}, $portal->{port}, 2) // 0;
+}
+
+# The number of portals with a live controller on their configured interface,
+# or, with $any_iface, on any interface: such a path carries I/O, even while
+# it waits to be moved to its configured interface.
+sub _live_portal_count($states, $portals, $any_iface = 0) {
+    my $live = 0;
+    for my $portal ($portals->@*) {
+        my $id = "$portal->{address}:$portal->{port}";
+        $live++ if grep {
+            $_->{state} eq 'live'
+                && ($any_iface || ($_->{host_iface} // '') eq $portal->{host_iface})
+        } ($states->{$id} // [])->@*;
+    }
+    return $live;
+}
+
+# Creates the controller of a portal in a child, giving up after 10 seconds.
+sub _connect_portal($scfg, $portal, $hostnqn, $hostid, $key) {
+    my ($options, $names) = _fabrics_options($scfg, $portal, $hostnqn, $hostid, $key);
+    return run_bounded(10, 'NVMe/TCP connect', sub { return _fabrics_connect($options, $names) });
+}
+
+my sub set_iopolicy($nqn, $policy) {
+    PVE::File::dir_glob_foreach(
+        '/sys/class/nvme-subsystem',
+        $RE_NVME_SUBSYSTEM,
+        sub($entry) {
+            my $base = "/sys/class/nvme-subsystem/$entry";
+            my $subsys = file_read_firstline("$base/subsysnqn");
+            return if !defined($subsys) || $subsys ne $nqn;
+            my $error = write_attribute("$base/iopolicy", "$policy\n");
+            die "unable to set NVMe multipath policy: $error\n" if defined($error);
+        },
+    );
+}
+
+# Applies the reconnect timeouts to connected controllers, which otherwise
+# keep the values of their connect. sysfs shows "off" for -1, and shows the
+# loss timeout rounded up to a multiple of the reconnect delay.
+my sub set_tunables($scfg) {
+    my $delay = $scfg->{'nvme-reconnect-delay'} // 2;
+    my $loss = $scfg->{'nvme-ctrl-loss-tmo'} // 600;
+    my $fast = $scfg->{'nvme-fast-io-fail-tmo'};
+    my $shown_loss = $loss < 0 ? 'off' : int(($loss + $delay - 1) / $delay) * $delay;
+    my $shown_fast = defined($fast) && $fast >= 0 ? $fast : 'off';
+
+    for my $controller (nvme_controllers($scfg->{subsysnqn})->@*) {
+        my $base = "/sys/class/nvme/$controller->{name}";
+        my $current_delay = file_read_firstline("$base/reconnect_delay") // next;
+        my $write = sub($attr, $value) {
+            my $error = write_attribute("$base/$attr", "$value\n") // return;
+            log_warn("cannot set $attr of NVMe controller '$controller->{name}': $error");
+        };
+        my $delay_changed = $current_delay ne "$delay";
+        $write->('reconnect_delay', $delay) if $delay_changed;
+        # the loss timeout is stored as a number of reconnects of the delay
+        $write->('ctrl_loss_tmo', $loss < 0 ? -1 : $loss)
+            if $delay_changed
+            || (file_read_firstline("$base/ctrl_loss_tmo") // '') ne "$shown_loss";
+        $write->('fast_io_fail_tmo', $shown_fast eq 'off' ? -1 : $fast)
+            if (file_read_firstline("$base/fast_io_fail_tmo") // '') ne "$shown_fast";
+    }
+}
+
+# Attempts to make a controller attribute readable by root only (a test seam).
+sub _restrict_attr($path) {
+    my @stat = stat($path);
+    if (!@stat) {
+        log_warn("cannot stat NVMe controller attribute '$path': $!") if $! != ENOENT;
+        return;
+    }
+    return if !($stat[2] & 077);
+    chmod(0600, $path) or log_warn("cannot restrict permissions of '$path': $!");
+}
+
+# The kernel creates the DH-HMAC-CHAP secret attributes world-readable.
+my sub restrict_secret_attrs($nqn) {
+    for my $controller (nvme_controllers($nqn)->@*) {
+        _restrict_attr("/sys/class/nvme/$controller->{name}/$_")
+            for qw(dhchap_secret dhchap_ctrl_secret);
+    }
+}
+
+# Best effort: a controller that went away since it was listed needs no scan,
+# and one that cannot be rescanned is a warning.
+my sub rescan_namespaces($nqn) {
+    for my $controller (nvme_controllers($nqn)->@*) {
+        next if $controller->{state} ne 'live';
+        my $path = "/sys/class/nvme/$controller->{name}/rescan_controller";
+        my $error = write_attribute($path, "1\n", 1) // next;
+        log_warn("cannot rescan NVMe controller '$controller->{name}': $error");
+    }
+}
+
+# Deletes a controller that is dead or that moved to another interface. The
+# activation reconnects or keeps the path either way, so this only warns.
+my sub disconnect_controller($controller) {
+    my $name = $controller->{name};
+    eval {
+        run_bounded(
+            5,
+            "deleting NVMe controller '$name'",
+            sub { return _delete_controllers([$controller]) },
+        );
+    };
+    if (my $error = $@) {
+        rethrow_task_interrupt($error);
+        log_warn($error);
+    }
+}
+
+# Waits up to 10 seconds for a live controller of the portal on its
+# configured interface.
+my sub wait_for_live_path($nqn, $portal) {
+    for (my $attempt = 0; $attempt < 40; $attempt++) {
+        return 1 if _live_portal_count(controller_states($nqn), [$portal]);
+        _sleep(0.25);
+    }
+    return 0;
+}
+
+# Connects every missing or dead path on its configured interface. A path
+# that is connected on another interface moves make-before-break: its old
+# controller is only disconnected once the new one is live, so changing an
+# interface mapping never takes away a path that carries I/O.
+my sub connect_portals($scfg, $portals, $hostnqn, $hostid, $key) {
+    my $nqn = $scfg->{subsysnqn};
+    for my $portal ($portals->@*) {
+        my $id = "$portal->{address}:$portal->{port}";
+        my $iface = $portal->{host_iface};
+        my @controllers = (controller_states($nqn)->{$id} // [])->@*;
+        my ($configured) = grep { ($_->{host_iface} // '') eq $iface } @controllers;
+        my @moved = grep { ($_->{host_iface} // '') ne $iface } @controllers;
+
+        if (!$configured || $configured->{state} eq 'dead') {
+            disconnect_controller($configured) if $configured;
+            if (!portal_reachable($portal)) {
+                log_warn("NVMe/TCP portal '$id' is unreachable");
+                next;
+            }
+            eval { _connect_portal($scfg, $portal, $hostnqn, $hostid, $key) };
+            my $connect_error = $@;
+            # A failed or timed-out connect can still have created a controller.
+            eval { restrict_secret_attrs($nqn); 1 };
+            if (my $error = $@) {
+                rethrow_task_interrupt($error);
+                log_warn("cannot restrict NVMe controller attributes for '$nqn'");
+            }
+            rethrow_task_interrupt($connect_error) if $connect_error;
+            if ($connect_error) {
+                log_warn("$id: $connect_error");
+                next;
+            }
+        }
+        next if !@moved;
+        if (!wait_for_live_path($nqn, $portal)) {
+            log_warn("NVMe/TCP portal '$id' is not live on '$iface' yet; keeping its path on"
+                . " another interface");
+            next;
+        }
+        disconnect_controller($_) for @moved;
+    }
+}
+
+sub activate_storage($class, $storeid, $scfg, $cache = undef) {
+    $cache //= {};
+    my $nqn = $scfg->{subsysnqn};
+    # First, whatever the rest of the activation does, also for controllers
+    # connected by hand. Also try after each connect attempt below, including
+    # failures that may have created a controller.
+    restrict_secret_attrs($nqn);
+
+    my $multipath = file_read_firstline('/sys/module/nvme_core/parameters/multipath')
+        // die "the NVMe kernel modules are not loaded; load the nvme-tcp kernel module\n";
+    die "native NVMe multipath is disabled in the running kernel\n" if $multipath ne 'Y';
+
+    my $portals = _configured_portals($scfg);
+    verify_nvme_nqn($nqn);
+    _validate_local_ifaces($portals);
+    _assert_unique_target($storeid, $scfg);
+    my $hostnqn = file_read_firstline('/etc/nvme/hostnqn')
+        // die "missing /etc/nvme/hostnqn; the nvme-cli package generates it\n";
+    verify_nvme_nqn($hostnqn);
+    my $hostnqns = parse_nvme_host_nqns($scfg->{'nvme-host-nqns'});
+    die "local NVMe host NQN '$hostnqn' is missing from nvme-host-nqns\n"
+        if !grep { $_ eq $hostnqn } $hostnqns->@*;
+    my $policy = $scfg->{'nvme-iopolicy'} // 'round-robin';
+    my $force_reconcile = delete($cache->{'zfsnvme-force-reconcile'}->{$storeid}) // 0;
+    my $states = controller_states($nqn);
+    my $healthy = _live_portal_count($states, $portals);
+    my $usable = _live_portal_count($states, $portals, 1);
+    my $last = $slow_path_last{$storeid};
+    my $recent = defined($last) && _now() - $last < $slow_path_backoff;
+
+    # This method is called by the periodic storage status loop. Once every
+    # configured path is live, lifecycle operations keep the target converged,
+    # so no remote call is needed. A degraded storage retries the target at
+    # most once a minute: in between, one with a live path (also one waiting
+    # to move to its configured interface) is usable, and one without fails
+    # at once instead of waiting for a path again. Missing namespaces
+    # explicitly force the slow path from activate_volume().
+    if (!$force_reconcile && ($healthy == scalar($portals->@*) || $recent)) {
+        die "no live NVMe/TCP path for storage '$storeid'\n" if !$usable;
+        die "missing NVMe DH-HMAC-CHAP key\n"
+            if (file_read_firstline(secret_path($storeid)) // '') eq '';
+        set_iopolicy($nqn, $policy);
+        set_tunables($scfg);
+        return 1;
+    }
+
+    $slow_path_last{$storeid} = _now();
+    my $hostid = file_read_firstline('/etc/nvme/hostid')
+        // die "missing /etc/nvme/hostid; the nvme-cli package generates it\n";
+    die "invalid NVMe host ID\n"
+        if $hostid !~ $RE_NVMET_UUID || $hostid =~ $RE_ZERO_UUID;
+    my $key = get_secret($storeid);
+    for my $name (qw(keep-alive-tmo reconnect-delay ctrl-loss-tmo fast-io-fail-tmo nr-io-queues)) {
+        my $value = $scfg->{"nvme-$name"};
+        die "invalid NVMe connection parameter '$name'\n"
+            if defined($value) && $value !~ $RE_CONFIG_INT;
+    }
+    _nvmet_activate_target($storeid, $scfg, $portals, $hostnqns, $key);
+    connect_portals($scfg, $portals, $hostnqn, $hostid, $key);
+
+    # Existing controllers can be in the middle of their kernel reconnect delay
+    # after a target restart. Do not create duplicates, but give that recovery
+    # cycle enough time to complete before declaring the storage unavailable.
+    for (my $attempt = 0; $attempt < 60; $attempt++) {
+        $states = controller_states($nqn);
+        last if _live_portal_count($states, $portals, 1);
+        _sleep(0.25);
+    }
+    die "no live NVMe/TCP path for storage '$storeid'\n"
+        if !_live_portal_count($states, $portals, 1);
+    $healthy = _live_portal_count($states, $portals);
+    log_warn("storage '$storeid' is degraded: $healthy/" . scalar($portals->@*) . " paths live")
+        if $healthy < scalar($portals->@*);
+
+    set_iopolicy($nqn, $policy);
+    set_tunables($scfg);
+    return 1;
+}
+
+sub deactivate_storage($class, $storeid, $scfg, $cache = undef) {
+    my $nqn = $scfg->{subsysnqn};
+    my $openers = _namespace_openers($nqn);
+    die "refusing to disconnect NVMe storage '$storeid': namespace in use by "
+        . join(', ', $openers->@*) . "\n"
+        if $openers->@*;
+
+    # One child deletes every controller, bounded by 15 seconds in all. When it
+    # gives up, the controllers that are left stay connected and the
+    # deactivation fails.
+    my $controllers = nvme_controllers($nqn);
+    run_bounded(
+        15,
+        "disconnecting NVMe subsystem '$nqn'",
+        sub { return _delete_controllers($controllers) },
+    ) if $controllers->@*;
+    return 1;
+}
+
+# Removing a storage definition never destroys data and never depends on the
+# target, like for every other storage type. Target cleanup is best effort
+# and only happens once the storage owns no volume. Other nodes keep their
+# connections until they are disconnected there.
+sub on_delete_hook($class, $storeid, $scfg) {
+    # First: an activation elsewhere that waits for the target lock then
+    # refuses instead of restoring what is removed below.
+    delete_secret($storeid);
+    eval { $class->deactivate_storage($storeid, $scfg) };
+    if (my $err = $@) {
+        rethrow_task_interrupt($err);
+        chomp($err);
+        log_warn("not disconnecting NVMe storage '$storeid': $err");
+    }
+    eval { _nvmet_delete_target($scfg) };
+    if (my $err = $@) {
+        rethrow_task_interrupt($err);
+        chomp($err);
+        log_warn("keeping the NVMe target configuration of storage '$storeid': $err");
+    }
+    return;
+}
+
+sub path($class, $scfg, $volname, $storeid, $snapname = undef) {
+    die "direct access to snapshots not implemented\n" if defined($snapname);
+    my ($vtype, $name, $vmid) = $class->parse_volname($volname);
+    my $uuid = _nvmet_volume_uuid($scfg, $name);
+    my $path = "/dev/disk/by-id/nvme-uuid.$uuid";
+    return ($path, $vmid, $vtype);
+}
+
+sub qemu_blockdev_options(
+    $class, $scfg, $storeid, $volname,
+    $machine_version = undef,
+    $options = undef,
+) {
+    die "direct access to snapshots not implemented\n" if $options->{'snapshot-name'};
+    my ($path) = $class->path($scfg, $volname, $storeid);
+    return { driver => 'host_device', filename => $path };
+}
+
+# Whether the local device behind a namespace link is the namespace (NSID,
+# UUID) of this subsystem, so a UUID claimed by another subsystem is never
+# handed to a guest.
+sub _nvmet_local_namespace_ok($path, $nqn, $nsid, $uuid) {
+    my $link = _link_target($path) // return 0;
+    my $device = (split m{/}, $link)[-1];
+    return 0 if !defined($device) || $device !~ $RE_NVME_NAMESPACE;
+    my $sys = "/sys/block/$device";
+    return 0 if lc(file_read_firstline("$sys/uuid") // '') ne lc($uuid);
+    return 0 if (file_read_firstline("$sys/nsid") // '') ne "$nsid";
+    return 0 if (file_read_firstline("$sys/device/subsysnqn") // '') ne $nqn;
+    return 1;
+}
+
+sub activate_volume(
+    $class, $storeid, $scfg, $volname,
+    $snapname = undef,
+    $cache = undef,
+    $hints = undef,
+) {
+    die "unable to activate snapshot from remote zfs storage\n" if $snapname;
+    my (undef, $name) = $class->parse_volname($volname);
+    my $nqn = nvmet_nqn($scfg);
+    my $dataset = nvmet_dataset($scfg, $name);
+
+    my ($inv, $cfs) = nvmet_read($scfg);
+    my ($nsid, $uuid) = _nvmet_owned_identity($inv, $nqn, $dataset);
+    my $state = _nvmet_export_state($cfs, $nqn, $nsid, $uuid, "/dev/zvol/$dataset");
+    my $path = "/dev/disk/by-id/nvme-uuid.$uuid";
+    my $ok = sub {
+        return _block_device($path) && _nvmet_local_namespace_ok($path, $nqn, $nsid, $uuid);
+    };
+    return 1 if $state eq 'present' && $ok->();
+
+    $cache //= {};
+    $cache->{'zfsnvme-force-reconcile'}->{$storeid} = 1;
+    $class->activate_storage($storeid, $scfg, $cache);
+    for (my $attempt = 0; $attempt < 40 && !_block_device($path); $attempt++) {
+        # a host can miss the change notice; rescanning is harmless
+        rescan_namespaces($nqn) if $attempt % 8 == 0;
+        _sleep(0.25);
+    }
+    die "NVMe namespace for '$volname' did not appear\n" if !_block_device($path);
+    die "NVMe namespace for '$volname' has an unexpected identity\n" if !$ok->();
+    return 1;
+}
+
+sub deactivate_volume(
+    $class, $storeid, $scfg, $volname, $snapname = undef, $cache = undef,
+) {
+    die "unable to deactivate snapshot from remote zfs storage\n" if $snapname;
+    return 1;
+}
+
+1;



  reply	other threads:[~2026-10-06  8:54 UTC|newest]

Thread overview: 7+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-05  0:26 [PATCH storage v3 0/4] add ZFS over NVMe/TCP storage plugin Joaquin Varela
2026-10-05  0:26 ` Joaquin Varela [this message]
2026-10-05  0:26 ` [PATCH storage v3 2/4] test: add zfsnvme plugin tests Joaquin Varela
2026-10-05  0:26 ` [PATCH storage v3 3/4] zfsnvme: fence target commands of abandoned transactions Joaquin Varela
2026-10-05  0:26 ` [PATCH storage v3 4/4] zfsnvme: wait up to 30 seconds for the shared storage lock Joaquin Varela
2026-10-05  0:26 ` [PATCH docs v3] storage: document ZFS over NVMe/TCP Joaquin Varela
2026-10-05  0:26 ` [PATCH manager v3] ui: storage: add ZFS over NVMe/TCP editor Joaquin Varela

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261005002609.571-2-joaquinvarela@neatech.ar \
    --to=joaquinvarela@neatech.ar \
    --cc=pve-devel@lists.proxmox.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Service provided by Proxmox Server Solutions GmbH | Privacy | Legal