all lists on lists.proxmox.com
 help / color / mirror / Atom feed
From: Joaquin Varela <joaquinvarela@neatech.ar>
To: pve-devel@lists.proxmox.com
Subject: [PATCH storage v3 1/4] zfsnvme: add ZFS over NVMe/TCP storage plugin
Date: Sun,  4 Oct 2026 21:26:04 -0300	[thread overview]
Message-ID: <20261005002609.571-2-joaquinvarela@neatech.ar> (raw)
In-Reply-To: <20261005002609.571-1-joaquinvarela@neatech.ar>

Add the shared storage type 'zfsnvme': ZFS zvols on a remote Linux
target, exported with the kernel NVMe target (nvmet, configured through
configfs) over NVMe/TCP, and consumed by the cluster nodes with native
NVMe multipath and DH-HMAC-CHAP authentication.

The ZFS dataset is the source of truth for a namespace: the subsystem
NQN, NSID and UUID of every volume are stored in the ZFS user properties
'proxmox:nvme-subsys', 'proxmox:nvme-nsid' and 'proxmox:nvme-uuid', and
the pool keeps the highest NSID handed out in 'proxmox:nvme-last-nsid'.
The nodes find a namespace by its UUID. Everything else on the target
(nvmet_tcp module, configfs mount, subsystem, namespaces, ports, host
ACLs) is derived state that an activation rebuilds after a target
restart.

All parsing, validation and planning happen locally. The target state
is read with one 'zfs get' and one 'find' over configfs; other reads
are 'zfs get' and 'zfs list', 'test -b' while a new zvol appears, and
'cat /proc/mounts' to decide whether configfs needs to be mounted. The
target runs quoted simple commands over SSH, joined with '&&' where a
step depends on the previous one: an export starts with 'test -b' on
its zvol, and each publish link repeats two read-only checks right
before it links the subsystem to a port. There are no scripts, loops
or variables on the target. As with ZFS over iSCSI, the SSH key is
/etc/pve/priv/zfs/<server>_id_rsa. Target changes are serialized
cluster-wide by a pmxcfs domain lock per target server, and an
operation whose SSH connection broke is reported as unknown and
repaired by the next activation rather than compensated from a read
that could overtake it.

Each storage needs its own NVMe/TCP listener (address family, address
and port) on the target. nvmet answers a connect to a listener that
does not publish the subsystem yet with DNR, and the host then deletes
the controller whatever its loss timeout, so after a target restart the
storage published second on a shared listener would lose its paths.
Adding, updating and activating a storage therefore refuse a portal
whose listener another zfsnvme storage uses, disabled ones included,
and a target address that another storage reaches through a different
'server' value, since that value names the target lock. A data address
thus identifies one target for all zfsnvme storages of the cluster.

The DH-HMAC-CHAP key is a sensitive property, stored like the other
storage secrets in /etc/pve/priv/storage/<storeid>.nvme-dhchap. It
reaches the target through stdin, and the plugin restricts the
dhchap_key and dhchap_ctrl_key attributes there to 0600 whenever it
writes the key. On the nodes, controllers are created with exactly one
write to /dev/nvme-fabrics in a bounded child (run_fork_with_timeout),
so the key never appears on a command line, and the secret attributes of
a new controller are restricted before the child returns. Rescans and
disconnects are sysfs writes; nvme-cli is not run.

Keep-alive, reconnect delay, controller loss and optional fast I/O
failure timeouts are configurable, with defaults that prefer waiting for
the target over failing guest I/O (ctrl_loss_tmo 600, fast_io_fail_tmo
unset).

Outside the plugin: register the type in PVE::Storage and the Makefile,
add it to @SHARED_STORAGE in PVE::Storage::Plugin, let 'pvesm add' and
'pvesm set' read --dhchap-key from a file, and depend on nvme-cli, which
creates /etc/nvme/hostnqn and /etc/nvme/hostid and remains the
administration tool, like the client packages of the other shared types.

Signed-off-by: Joaquin Varela <joaquinvarela@neatech.ar>
---
 debian/control                   |    1 +
 src/PVE/CLI/pvesm.pm             |   21 +-
 src/PVE/Storage.pm               |    2 +
 src/PVE/Storage/Makefile         |    1 +
 src/PVE/Storage/Plugin.pm        |    2 +-
 src/PVE/Storage/ZFSNVMePlugin.pm | 3177 ++++++++++++++++++++++++++++++
 6 files changed, 3201 insertions(+), 3 deletions(-)
 create mode 100644 src/PVE/Storage/ZFSNVMePlugin.pm

diff --git a/debian/control b/debian/control
index 850cd57c..ab46e6b6 100644
--- a/debian/control
+++ b/debian/control
@@ -43,6 +43,7 @@ Depends: bzip2,
          lvm2,
          lzop,
          nfs-common,
+         nvme-cli,
          proxmox-backup-client (>= 2.1.10~),
          proxmox-backup-file-restore,
          pve-cluster (>= 5.0-32),
diff --git a/src/PVE/CLI/pvesm.pm b/src/PVE/CLI/pvesm.pm
index 1aaa9f3c..3c68ffb5 100755
--- a/src/PVE/CLI/pvesm.pm
+++ b/src/PVE/CLI/pvesm.pm
@@ -78,12 +78,29 @@ sub param_mapping {
         },
     };
 
+    my $dhchap_key_map = {
+        name => 'dhchap-key',
+        desc => 'a file containing the NVMe DH-HMAC-CHAP key',
+        func => sub {
+            my ($value) = @_;
+            # never repeat a key passed in place of the file name
+            die "dhchap-key expects the path of a file containing the key\n"
+                if $value =~ /^DHHC-1:/i || -d $value;
+            my ($key) = split(/\n/, PVE::Tools::file_get_contents($value), 2);
+            $key = PVE::Tools::trim($key // '');
+            die "DH-HMAC-CHAP key file '$value' is empty\n" if $key eq '';
+            return $key;
+        },
+    };
+
     my $mapping = {
         'cifsscan' => [$password_map],
         'cifs' => [$password_map],
         'pbs' => [$password_map],
-        'create' => [$password_map, $enc_key_map, $master_key_map, $keyring_map],
-        'update' => [$password_map, $enc_key_map, $master_key_map, $keyring_map],
+        'create' =>
+            [$password_map, $enc_key_map, $master_key_map, $keyring_map, $dhchap_key_map],
+        'update' =>
+            [$password_map, $enc_key_map, $master_key_map, $keyring_map, $dhchap_key_map],
     };
     return $mapping->{$name};
 }
diff --git a/src/PVE/Storage.pm b/src/PVE/Storage.pm
index fc3db812..2e5a8031 100755
--- a/src/PVE/Storage.pm
+++ b/src/PVE/Storage.pm
@@ -36,6 +36,7 @@ use PVE::Storage::CephFSPlugin;
 use PVE::Storage::ISCSIDirectPlugin;
 use PVE::Storage::ZFSPoolPlugin;
 use PVE::Storage::ZFSPlugin;
+use PVE::Storage::ZFSNVMePlugin;
 use PVE::Storage::PBSPlugin;
 use PVE::Storage::BTRFSPlugin;
 use PVE::Storage::ESXiPlugin;
@@ -61,6 +62,7 @@ PVE::Storage::CephFSPlugin->register();
 PVE::Storage::ISCSIDirectPlugin->register();
 PVE::Storage::ZFSPoolPlugin->register();
 PVE::Storage::ZFSPlugin->register();
+PVE::Storage::ZFSNVMePlugin->register();
 PVE::Storage::PBSPlugin->register();
 PVE::Storage::BTRFSPlugin->register();
 PVE::Storage::ESXiPlugin->register();
diff --git a/src/PVE/Storage/Makefile b/src/PVE/Storage/Makefile
index a67dc25f..d1cbfe29 100644
--- a/src/PVE/Storage/Makefile
+++ b/src/PVE/Storage/Makefile
@@ -11,6 +11,7 @@ SOURCES= \
 	ISCSIDirectPlugin.pm \
 	ZFSPoolPlugin.pm \
 	ZFSPlugin.pm \
+	ZFSNVMePlugin.pm \
 	PBSPlugin.pm \
 	BTRFSPlugin.pm \
 	LvmThinPlugin.pm \
diff --git a/src/PVE/Storage/Plugin.pm b/src/PVE/Storage/Plugin.pm
index a9e17513..a5aedc2b 100644
--- a/src/PVE/Storage/Plugin.pm
+++ b/src/PVE/Storage/Plugin.pm
@@ -35,7 +35,7 @@ our @COMMON_TAR_FLAGS = qw(
 );
 
 our @SHARED_STORAGE = (
-    'iscsi', 'nfs', 'cifs', 'rbd', 'cephfs', 'iscsidirect', 'zfs', 'drbd', 'pbs',
+    'iscsi', 'nfs', 'cifs', 'rbd', 'cephfs', 'iscsidirect', 'zfs', 'zfsnvme', 'drbd', 'pbs',
 );
 
 our $QCOW2_PREALLOCATION = {
diff --git a/src/PVE/Storage/ZFSNVMePlugin.pm b/src/PVE/Storage/ZFSNVMePlugin.pm
new file mode 100644
index 00000000..127680fb
--- /dev/null
+++ b/src/PVE/Storage/ZFSNVMePlugin.pm
@@ -0,0 +1,3177 @@
+package PVE::Storage::ZFSNVMePlugin;
+
+use v5.36;
+
+use Compress::Zlib qw(crc32);
+use Digest::SHA qw(sha256_hex);
+use Errno qw(ENOENT);
+use Fcntl qw(O_NOFOLLOW O_RDONLY O_RDWR S_ISCHR);
+use File::Path qw(make_path);
+use List::Util qw(max);
+use MIME::Base64 qw(decode_base64);
+use POSIX qw(SIG_BLOCK SIG_SETMASK SIGHUP SIGINT SIGQUIT SIGTERM sigprocmask);
+use Socket qw(AF_INET6 inet_ntop inet_pton);
+use Time::HiRes;
+
+use PVE::Cluster;
+use PVE::File;
+use PVE::JSONSchema;
+use PVE::Network;
+use PVE::RESTEnvironment qw(log_warn);
+use PVE::RPCEnvironment;
+use PVE::SysFSTools;
+use PVE::Tools qw(file_read_firstline file_set_contents run_command trim);
+
+use base qw(PVE::Storage::Plugin);
+
+# ZFS zvols on a remote Linux target, exported through the kernel NVMe target
+# (nvmet, configured through configfs) over NVMe/TCP and consumed with native
+# NVMe multipath.
+#
+# The ZFS dataset is the source of truth for a namespace: its owner NQN, NSID
+# and UUID are ZFS user properties. Everything else on the target is derived
+# state that the plugin rebuilds after a target restart.
+#
+# Storage parsing, validation and planning happen locally. SSH carries quoted
+# commands, sometimes joined with `&&`, to read or change ZFS/configfs state.
+# A pmxcfs domain lock per target serializes the changes within the cluster,
+# held like the storage lock of the other shared storage types.
+
+# ---------------------------------------------------------------------------
+# Constants and regular expressions
+# ---------------------------------------------------------------------------
+
+# LogLevel=ERROR keeps a login banner out of the first stderr line, which is
+# the error text of a failed call.
+my @ssh_cmd = (
+    '/usr/bin/ssh', '-o', 'BatchMode=yes', '-o', 'ConnectTimeout=10', '-o', 'LogLevel=ERROR',
+);
+my $id_rsa_path = '/etc/pve/priv/zfs';
+my $secret_dir = '/etc/pve/priv/storage';
+my $max_paths = 16;
+my $max_hosts = 64;
+my $slow_path_backoff = 60;
+
+my $nvmet_root = '/sys/kernel/config/nvmet';
+my $nvmet_model = 'Proxmox ZFS NVMe';
+my $nvmet_marker = 'ZFSNVME-CONFIGFS';
+my $nvmet_max_nsid = 0xfffffffe;
+my $nvmet_max_command = 65536;
+# Seconds to wait for the target lock while another node changes the target;
+# a destroy or rollback can hold it for several seconds.
+my $nvmet_lock_wait = 30;
+# On the pool dataset: the highest NSID handed out for a volume of the pool.
+my $nvmet_last_nsid = 'proxmox:nvme-last-nsid';
+
+my @nvmet_zfs_props =
+    ('type', 'proxmox:nvme-subsys', 'proxmox:nvme-nsid', 'proxmox:nvme-uuid', $nvmet_last_nsid);
+my %nvmet_zfs_keys = (
+    type => 'type',
+    'proxmox:nvme-subsys' => 'nqn',
+    'proxmox:nvme-nsid' => 'nsid',
+    'proxmox:nvme-uuid' => 'uuid',
+    $nvmet_last_nsid => 'last_nsid',
+);
+my @nvmet_cfs_attrs = qw(
+    enable device_path device_uuid buffered_io
+    attr_model attr_serial attr_allow_any_host
+    addr_trtype addr_adrfam addr_traddr addr_trsvcid
+);
+# Every command the plugin runs on the target, besides the `dd | tee` pipe of
+# a key write. The configfs read runs find through env to set its locale.
+my %nvmet_commands = map { $_ => 1 } qw(
+    cat chmod env grep ln mkdir modprobe mount printf rm rmdir test zfs
+);
+
+my $RE_NQN = qr{
+    \A
+    nqn \.
+    [A-Za-z0-9] [A-Za-z0-9.-]*
+    :
+    [A-Za-z0-9] [A-Za-z0-9._:-]*
+    \z
+}nxx;
+my $RE_IPV4_PORTAL = qr{
+    \A
+    (?<address> [^:]+)
+    (?: : (?<port> [0-9]+))?
+    \z
+}nxx;
+my $RE_IPV6_PORTAL = qr{
+    \A
+    \[ (?<address> [^\]]+) \]
+    (?: : (?<port> [0-9]+))?
+    \z
+}nxx;
+# A canonical IPv4-mapped IPv6 address.
+my $RE_IPV4_MAPPED = qr{\A ::ffff: (?<address> [0-9.]+) \z}nxx;
+my $RE_HOST_IFACE = qr{\A [A-Za-z0-9_.-]+ \z}nxx;
+my $RE_DHCHAP_KEY = qr{
+    \A DHHC-1 : (?<hash> 0[0-3]) : (?<secret> [A-Za-z0-9+/]+ ={0,2}) : \z
+}nxx;
+# A stopped PVE task: the worker's signal handler dies with the first text,
+# run_command adds the command context, and run_fork_with_timeout reports
+# the signal in its own words.
+my $RE_TASK_INTERRUPT = qr{
+    \A (?: (?: command \x20 ' .* ' \x20 failed: \x20 )? received \x20 interrupt
+    | interrupted \x20 by \x20 unexpected \x20 signal ) \n \z
+}nsxx;
+my $RE_COMMAND_TIMEOUT = qr{
+    \A (?: command \x20 ' .* ' \x20 failed: \x20 )? got \x20 timeout \n \z
+}nsxx;
+my $RE_COMMAND_EXIT =
+    qr{\A command \x20 ' .* ' \x20 failed: \x20 exit \x20 code \x20 (?<code> [0-9]+) \n \z}nsxx;
+my $RE_FABRICS_VALUE = qr{\A [^\s,\x00]+ \z}nxx;
+my $RE_FABRICS_RESULT = qr{\A instance=(?<instance> [0-9]+) ,cntlid= [0-9]+ \n \z}nxx;
+my $RE_ZERO_UUID = qr{\A 0{8} (?: -0{4}){3} -0{12} \z}nxx;
+my $RE_CONFIG_INT = qr{\A -?[0-9]+ \z}nxx;
+my $RE_NVME_CONTROLLER = qr{\A nvme [0-9]+ \z}nxx;
+my $RE_NVME_SUBSYSTEM = qr{\A nvme-subsys [0-9]+ \z}nxx;
+my $RE_NVME_NAMESPACE = qr{\A nvme [0-9]+ n [0-9]+ \z}nxx;
+my $RE_NVME_PARTITION = qr{\A nvme [0-9]+ n [0-9]+ p [0-9]+ \z}nxx;
+my $RE_TRADDR = qr{(?: \A | ,) traddr=(?<value>[^,]+)}nxx;
+my $RE_TRSVCID = qr{(?: \A | ,) trsvcid=(?<value>[^,]+)}nxx;
+my $RE_HOST_IFACE_ADDRESS = qr{(?: \A | ,) host_iface=(?<value>[^,]+)}nxx;
+my $RE_ZVOL_OWNER = qr{^ (?:vm|base|subvol|basevol)- (?<owner>\d+) - \S+ $}nxx;
+my $RE_VOLNAME = qr{
+    ^
+    (?: (?<base> (?:base|basevol)- (?<base_vmid>\d+) - \S+) /)?
+    (?<name> (?<type>base|basevol|vm|subvol)- (?<vmid>\d+) - \S+)
+    $
+}nxx;
+my $RE_BASE_SNAPSHOT = qr{^ (?<base>\S+) \@__base__ $}nxx;
+my $RE_UNSIGNED_INTEGER = qr{^ (?<value>\d+) $}nxx;
+
+my $RE_NVMET_UUID = qr{\A [0-9a-fA-F]{8} (?: - [0-9a-fA-F]{4}){3} - [0-9a-fA-F]{12} \z}nxx;
+my $RE_NVMET_NSID = qr{\A [1-9] [0-9]* \z}nxx;
+my $RE_NVMET_POOL = qr{\A [A-Za-z0-9] [A-Za-z0-9_.:/-]* \z}nxx;
+my $RE_NVMET_DATASET_NAME = qr{\A [A-Za-z0-9] [A-Za-z0-9_.:-]* \z}nxx;
+my $RE_NVMET_SNAPSHOT = qr{\A [A-Za-z0-9_.:-]+ \z}nxx;
+my $RE_NVMET_TEMPLATE_ZVOL = qr{/ (?<type> vm | base) - (?<suffix> [^/]+) \z}nxx;
+my $RE_NVMET_INHERITED = qr{\A inherited \x20 from \x20 (?<source> .+) \z}nxx;
+# Errors after which the target state is not known: an interrupted task, or a
+# change whose command may still be running on the target.
+my $RE_NVMET_ABANDONED = qr{
+    (?: \A received \x20 interrupt | ; \x20 the \x20 target \x20 state \x20 is \x20 unknown) \n \z
+}nxx;
+
+# Lines of the configfs read (nvmet_configfs_step below): `D <dir>`,
+# `L <link>`, `<file>:<value>` from grep, and `<sha256>  <file>` from
+# sha256sum. Attribute names anchor the split, so an NQN containing ':' is
+# never ambiguous.
+my $cfs_root = quotemeta($nvmet_root);
+my $RE_CFS_TOP = qr{\A D \x20 $cfs_root (?: / (?<top> hosts | ports | subsystems))? \z}nxx;
+my $RE_CFS_SUBSYS_DIR = qr{\A D \x20 $cfs_root /subsystems/ (?<nqn> [^/]+) \z}nxx;
+my $RE_CFS_SUBSYS_GROUP = qr{
+    \A D \x20 $cfs_root /subsystems/ [^/]+ / (?: namespaces | allowed_hosts) \z
+}nxx;
+my $RE_CFS_NS_DIR = qr{
+    \A D \x20 $cfs_root /subsystems/ (?<nqn> [^/]+) /namespaces/ (?<nsid> [0-9]+) \z
+}nxx;
+my $RE_CFS_PORT_DIR = qr{\A D \x20 $cfs_root /ports/ (?<port> [0-9]+) \z}nxx;
+my $RE_CFS_PORT_GROUP = qr{\A D \x20 $cfs_root /ports/ [0-9]+ /subsystems \z}nxx;
+my $RE_CFS_HOST_DIR = qr{\A D \x20 $cfs_root /hosts/ (?<host> [^/]+) \z}nxx;
+my $RE_CFS_PORT_LINK = qr{
+    \A L \x20 $cfs_root /ports/ (?<port> [0-9]+) /subsystems/ (?<nqn> [^/]+) \z
+}nxx;
+my $RE_CFS_ACL_LINK = qr{
+    \A L \x20 $cfs_root /subsystems/ (?<nqn> [^/]+) /allowed_hosts/ (?<host> [^/]+) \z
+}nxx;
+my $RE_CFS_OTHER_ENTRY = qr{\A [DL] \x20 $cfs_root /}nxx;
+my $RE_CFS_NS_ATTR = qr{
+    \A $cfs_root /subsystems/ (?<nqn> [^/]+) /namespaces/ (?<nsid> [0-9]+)
+    / (?<attr> enable | device_path | device_uuid | buffered_io) : (?<value> .*) \z
+}nxx;
+my $RE_CFS_SUBSYS_ATTR = qr{
+    \A $cfs_root /subsystems/ (?<nqn> [^/]+)
+    / (?<attr> attr_model | attr_serial | attr_allow_any_host) : (?<value> .*) \z
+}nxx;
+my $RE_CFS_PORT_ATTR = qr{
+    \A $cfs_root /ports/ (?<port> [0-9]+)
+    / (?<attr> addr_trtype | addr_adrfam | addr_traddr | addr_trsvcid) : (?<value> .*) \z
+}nxx;
+my $RE_CFS_KEY_DIGEST = qr{
+    \A (?<sha> [0-9a-f]{64}) \x20 [\x20*] $cfs_root /hosts/ (?<host> [^/]+) /dhchap_key \z
+}nxx;
+my $RE_CFS_OTHER_ATTR = qr{\A $cfs_root /}nxx;
+
+# ---------------------------------------------------------------------------
+# Target addressing and validated names
+# ---------------------------------------------------------------------------
+
+my sub nvmet_server($scfg) {
+    return $scfg->{server};
+}
+
+my sub nvmet_ssh_key($scfg) {
+    return "$id_rsa_path/" . nvmet_server($scfg) . '_id_rsa';
+}
+
+my sub nvmet_lock_id($scfg) {
+    my $server = nvmet_server($scfg) // '';
+    return 'zfsnvme-' . ($server =~ s/[^A-Za-z0-9.-]/_/gr);
+}
+
+my sub secret_path($storeid) {
+    return "$secret_dir/$storeid.nvme-dhchap";
+}
+
+my sub nvmet_nqn($scfg) {
+    my $nqn = $scfg->{subsysnqn};
+    die "invalid NVMe subsystem NQN\n" if !defined($nqn) || !verify_nvme_nqn($nqn, 1);
+    return $nqn;
+}
+
+my sub nvmet_pool($scfg) {
+    my $pool = $scfg->{pool};
+    die "invalid ZFS pool name\n" if !defined($pool) || $pool !~ $RE_NVMET_POOL;
+    return $pool;
+}
+
+my sub nvmet_dataset($scfg, $name) {
+    die "invalid ZFS volume name\n" if !defined($name) || $name !~ $RE_NVMET_DATASET_NAME;
+    return nvmet_pool($scfg) . "/$name";
+}
+
+# zfs destroy reads '%' and ',' in a snapshot name as a range and a list.
+my sub nvmet_snapshot($scfg, $name, $snap) {
+    die "invalid snapshot name\n" if !defined($snap) || $snap !~ $RE_NVMET_SNAPSHOT;
+    return nvmet_dataset($scfg, $name) . "\@$snap";
+}
+
+my sub nvmet_valid_nsid($value) {
+    return defined($value) && $value =~ $RE_NVMET_NSID && $value <= $nvmet_max_nsid;
+}
+
+# IPv6 addresses have many spellings; compare and store only the canonical one.
+my sub nvmet_canonical_address($address) {
+    return $address if !defined($address) || index($address, ':') < 0;
+    my $packed = inet_pton(AF_INET6, $address) // return $address;
+    return inet_ntop(AF_INET6, $packed);
+}
+
+# The address family and address that a portal of parse_nvme_portals reaches:
+# an IPv4-mapped IPv6 address reaches the IPv4 address.
+my sub nvmet_listener_address($portal) {
+    return ('ipv4', $+{address})
+        if $portal->{family} eq 'ipv6' && $portal->{address} =~ $RE_IPV4_MAPPED;
+    return $portal->@{qw(family address)};
+}
+
+# Target-provided names only appear in messages when they are well-formed.
+my sub nvmet_label($name) {
+    return $name =~ $RE_NQN ? " '$name'" : '';
+}
+
+# The other name of a zvol that a template conversion renames: vm-* and base-*.
+my sub nvmet_template_twin($name) {
+    return '' if $name !~ $RE_NVMET_TEMPLATE_ZVOL;
+    my $twin = ($+{type} eq 'vm' ? 'base' : 'vm') . "-$+{suffix}";
+    return substr($name, 0, $-[0]) . "/$twin";
+}
+
+# The timeout of a call that changes ZFS: ZFS waits for transaction groups,
+# which a busy pool can take minutes to sync.
+my sub nvmet_long_timeout() {
+    return PVE::RPCEnvironment->is_worker() ? 60 * 60 : 60;
+}
+
+# ---------------------------------------------------------------------------
+# Remote steps
+#
+# A step is an argv array, a write `{ write => $path, value => $value }` or a
+# key write `{ key => [$path, ...] }` fed from stdin. A unit is a list of steps
+# that must stay in one call.
+# ---------------------------------------------------------------------------
+
+my sub nvmet_subsys_path($nqn) {
+    return "$nvmet_root/subsystems/$nqn";
+}
+
+my sub nvmet_ns_path($nqn, $nsid) {
+    return "$nvmet_root/subsystems/$nqn/namespaces/$nsid";
+}
+
+my sub nvmet_port_path($id) {
+    return "$nvmet_root/ports/$id";
+}
+
+my sub nvmet_host_path($hostnqn) {
+    return "$nvmet_root/hosts/$hostnqn";
+}
+
+my sub nvmet_write($path, $value) {
+    return { write => $path, value => $value };
+}
+
+# Parts of a parsed configfs state; missing parts read as empty.
+my sub nvmet_namespaces($cfs, $nqn) {
+    my $subsys = $cfs->{subsystems}->{$nqn};
+    return ($subsys ? $subsys->{namespaces} : undef) // {};
+}
+
+my sub nvmet_port_linked($cfs, $id, $nqn) {
+    my $port = $cfs->{ports}->{$id};
+    return $port && ($port->{links} // {})->{$nqn};
+}
+
+my sub nvmet_linked_hosts($cfs) {
+    my %linked;
+    for my $subsys (values $cfs->{subsystems}->%*) {
+        $linked{$_} = 1 for keys(($subsys->{acl} // {})->%*);
+    }
+    return \%linked;
+}
+
+# A subsystem whose creation by this plugin stopped after its model was
+# written, before its serial: nobody else writes our model, and it exports,
+# allows and publishes nothing, so completing it takes nothing from anybody.
+# A subsystem with the kernel's default model may belong to another tool.
+my sub nvmet_unfinished_subsystem($cfs, $nqn) {
+    my $subsys = $cfs->{subsystems}->{$nqn};
+    return 0 if !$subsys;
+    return
+        $subsys->{attr_model} eq $nvmet_model
+        && !($subsys->{namespaces} // {})->%*
+        && !($subsys->{acl} // {})->%*
+        && !grep { nvmet_port_linked($cfs, $_, $nqn) } keys $cfs->{ports}->%*;
+}
+
+# Same write order as the kernel requires: the device before enabling it.
+my sub nvmet_build_unit($nqn, $nsid, $uuid, $dev) {
+    my $ns = nvmet_ns_path($nqn, $nsid);
+    return [
+        ['test', '-b', $dev],
+        ['mkdir', $ns],
+        nvmet_write("$ns/device_path", $dev),
+        nvmet_write("$ns/device_uuid", $uuid),
+        nvmet_write("$ns/buffered_io", 0),
+        nvmet_write("$ns/enable", 1),
+    ];
+}
+
+my sub nvmet_enable_unit($nqn, $nsid, $dev) {
+    return [['test', '-b', $dev], nvmet_write(nvmet_ns_path($nqn, $nsid) . '/enable', 1)];
+}
+
+my sub nvmet_unexport_unit($nqn, $nsid, $enabled) {
+    my $ns = nvmet_ns_path($nqn, $nsid);
+    return [($enabled ? (nvmet_write("$ns/enable", 0)) : ()), ['rmdir', $ns]];
+}
+
+my sub nvmet_identity($nqn, $nsid, $uuid) {
+    return ("proxmox:nvme-subsys=$nqn", "proxmox:nvme-nsid=$nsid", "proxmox:nvme-uuid=$uuid");
+}
+
+# The pool and its direct children with their identity properties.
+my sub nvmet_zfs_inventory_step($pool) {
+    return [
+        'zfs',
+        'get',
+        '-H',
+        '-p',
+        '-d',
+        '1',
+        '-t',
+        'filesystem,volume',
+        '-o',
+        'name,property,value,source',
+        join(',', @nvmet_zfs_props),
+        $pool,
+    ];
+}
+
+# Every configfs object the plugin uses, in one traversal. Keys are never
+# read back; with host NQNs, only their sha256 digests are returned. Only the
+# kernel's own groups of objects the plugin does not use are skipped, so an
+# object of another tool with one of their names is still listed.
+my sub nvmet_configfs_step($hostnqns) {
+    my @names = map { ('-o', '-name', $_) } @nvmet_cfs_attrs;
+    shift @names;
+    my @keys = map { ('-o', '-path', nvmet_host_path($_) . '/dhchap_key') } $hostnqns->@*;
+    shift @keys;
+    return [
+        'env',
+        'LC_ALL=C',
+        'find',
+        $nvmet_root,
+        '(',
+        '-path',
+        "$nvmet_root/subsystems/*/passthru",
+        '-o',
+        '-path',
+        "$nvmet_root/ports/*/ana_groups",
+        '-o',
+        '-path',
+        "$nvmet_root/ports/*/referrals",
+        ')',
+        '-type',
+        'd',
+        '-prune',
+        '-o',
+        '-type',
+        'd',
+        '-exec',
+        'printf',
+        'D %s\n',
+        '{}',
+        '+',
+        '-o',
+        '-type',
+        'l',
+        '-exec',
+        'printf',
+        'L %s\n',
+        '{}',
+        '+',
+        '-o',
+        '-type',
+        'f',
+        '(',
+        @names,
+        ')',
+        '-exec',
+        'grep',
+        '',
+        '/dev/null',
+        '{}',
+        '+',
+        (@keys ? ('-o', '-type', 'f', '(', @keys, ')', '-exec', 'sha256sum', '{}', '+') : ()),
+    ];
+}
+
+# The whole target state: the ZFS inventory, a marker and configfs. Only a
+# locked read loads nvmet first (modprobe).
+my sub nvmet_state_steps($pool, $hostnqns, $modprobe) {
+    return [
+        ($modprobe ? (['modprobe', 'nvmet_tcp']) : ()),
+        nvmet_zfs_inventory_step($pool),
+        ['printf', '%s\n', $nvmet_marker],
+        nvmet_configfs_step($hostnqns),
+    ];
+}
+
+# ---------------------------------------------------------------------------
+# Renderer and runner
+# ---------------------------------------------------------------------------
+
+my sub nvmet_valid_word($word) {
+    return defined($word) && !ref($word) && $word !~ /[\0\r\n]/;
+}
+
+my sub nvmet_valid_path($path) {
+    return nvmet_valid_word($path) && $path =~ m{\A/};
+}
+
+my sub nvmet_quote($word) {
+    return PVE::Tools::shellquote($word);
+}
+
+# Renders steps into one POSIX shell command line: simple commands joined with
+# `&&`, plus `>` for configfs writes and one `dd | tee` pipe for a key read from
+# stdin. There are no loops, variables, conditionals or substitutions. Every
+# word is quoted; all operands are built from validated configuration.
+sub _nvmet_render($steps) {
+    my $invalid = "internal error: invalid NVMe target step\n";
+    die $invalid if ref($steps) ne 'ARRAY' || !$steps->@*;
+
+    my $keys = 0;
+    my @commands;
+    for my $step ($steps->@*) {
+        if (ref($step) eq 'ARRAY') {
+            die $invalid if !$step->@* || grep { !nvmet_valid_word($_) } $step->@*;
+            die $invalid if !$nvmet_commands{ $step->[0] };
+            push @commands, join(' ', map { nvmet_quote($_) } $step->@*);
+        } elsif (ref($step) eq 'HASH' && exists($step->{write})) {
+            die $invalid
+                if keys($step->%*) != 2
+                || !nvmet_valid_path($step->{write})
+                || !nvmet_valid_word($step->{value});
+            push @commands,
+                q{printf '%s\n' }
+                . nvmet_quote($step->{value}) . ' > '
+                . nvmet_quote($step->{write});
+        } elsif (ref($step) eq 'HASH' && exists($step->{key})) {
+            my $paths = $step->{key};
+            die $invalid
+                if keys($step->%*) != 1
+                || ref($paths) ne 'ARRAY'
+                || !$paths->@*
+                || grep { !nvmet_valid_path($_) } $paths->@*;
+            $keys++;
+            die $invalid if $keys > 1;
+            # configfs needs the whole key in one write(2). dd collects its
+            # input into one output block: GNU dd without bs=, busybox dd only
+            # when ibs and obs differ. tee writes it to every host object.
+            push @commands,
+                'dd ibs=4096 obs=8192 2>/dev/null | tee '
+                . join(' ', map { nvmet_quote($_) } $paths->@*)
+                . ' >/dev/null';
+        } else {
+            die $invalid;
+        }
+    }
+    return join(' && ', @commands);
+}
+
+# Packs whole units into calls of at most $max bytes of rendered command, the
+# single argument sshd runs, below its limit. A rendered command is ASCII
+# (every operand is validated), so its length is its size in bytes.
+sub _nvmet_chunk($units, $max = $nvmet_max_command) {
+    my (@chunks, @current);
+    for my $unit ($units->@*) {
+        my @next = (@current, $unit->@*);
+        if (length(_nvmet_render(\@next)) > $max) {
+            die "internal error: NVMe target command too long\n"
+                if !@current || length(_nvmet_render($unit)) > $max;
+            push @chunks, [@current];
+            @next = $unit->@*;
+        }
+        @current = @next;
+    }
+    push @chunks, [@current] if @current;
+    return \@chunks;
+}
+
+# Never propagate the context run_command adds to the task marker: it quotes
+# the command line.
+my sub rethrow_task_interrupt($error) {
+    die "received interrupt\n" if $error =~ $RE_TASK_INTERRUPT;
+}
+
+# The only code that runs anything on the target. Returns
+# { rc => exit code, or -1 when ssh did not run to its end, out => [lines],
+# err => text } and never dies on a failing command; a stopped task dies with
+# the task marker. %opts: op (label, required), timeout and input (stdin,
+# used for keys).
+sub _nvmet_run($scfg, $steps, %opts) {
+    die "internal error: NVMe target call without label\n" if !defined($opts{op});
+    my $command = _nvmet_render($steps);
+    die "internal error: NVMe target command too long\n" if length($command) > $nvmet_max_command;
+    my $cmd = [@ssh_cmd, '-i', nvmet_ssh_key($scfg), 'root@' . nvmet_server($scfg), $command];
+    my (@out, $err);
+    my $rc = 0;
+    # Not noerr: run_command would then also swallow the interrupt of a stopped
+    # task, which must end the call. The exit code is taken from the exception.
+    eval {
+        run_command(
+            $cmd,
+            timeout => $opts{timeout} // 15,
+            outfunc => sub($line) { push @out, $line },
+            errfunc => sub($line) { $err = $line if !defined($err) && $line ne '' },
+            (defined($opts{input}) ? (input => $opts{input}) : ()),
+        );
+    };
+    if (my $error = $@) {
+        rethrow_task_interrupt($error);
+        return { rc => -1, out => [], err => 'timeout' } if $error =~ $RE_COMMAND_TIMEOUT;
+        return { rc => -1, out => [], err => 'ssh failed' } if $error !~ $RE_COMMAND_EXIT;
+        $rc = $+{code};
+    }
+    return { rc => $rc, out => \@out, err => $err // ($rc ? "exit code $rc" : '') };
+}
+
+# Monotonic clock for the activation backoff and local waits (a test seam).
+sub _now() {
+    return Time::HiRes::clock_gettime(Time::HiRes::CLOCK_MONOTONIC());
+}
+
+# Every local wait of the plugin (a test seam).
+sub _sleep($seconds) {
+    Time::HiRes::sleep($seconds);
+    return;
+}
+
+# ---------------------------------------------------------------------------
+# Parsers (pure)
+# ---------------------------------------------------------------------------
+
+sub _nvmet_split_state($text) {
+    my @parts = split /^\Q$nvmet_marker\E\n/m, $text, -1;
+    die "malformed NVMe target state\n" if @parts != 2;
+    return @parts;
+}
+
+sub _nvmet_property_value($value, $source) {
+    die "malformed ZFS property value\n"
+        if !defined($value)
+        || !defined($source)
+        || $value eq ''
+        || $source eq ''
+        || $value =~ /[\r\n\t]/
+        || $source =~ /[\r\n\t]/;
+    return $value if $source eq 'local' || $source eq 'received';
+    return '-' if $value eq '-' && $source eq '-';
+    if ($source =~ $RE_NVMET_INHERITED) {
+        die "invalid ZFS pool name\n" if $+{source} !~ $RE_NVMET_POOL;
+        return '-';
+    }
+    die "invalid ZFS property source\n";
+}
+
+# Parses the ZFS inventory into { types => { dataset => type }, volumes =>
+# { dataset => { nqn, nsid, uuid } }, last_nsid => highest NSID handed out }.
+# The inventory must be complete: the pool rows and all rows of every
+# dataset, or the read fails. Children whose names PVE never uses (for
+# example with a space) are kept as datasets but never count as owned.
+sub _nvmet_parse_zfs_inventory($text, $pool) {
+    die "invalid ZFS pool name\n" if $pool !~ $RE_NVMET_POOL;
+
+    my (%rows, %foreign);
+    for my $line (split /\n/, $text) {
+        my @fields = split /\t/, $line, -1;
+        die "malformed ZFS inventory row\n" if @fields != 4 || $line =~ /\r/;
+        my ($dataset, $property, $value, $source) = @fields;
+        if ($dataset ne $pool) {
+            die "ZFS dataset is outside configured pool\n" if index($dataset, "$pool/") != 0;
+            my $child = substr($dataset, length($pool) + 1);
+            die "ZFS dataset is outside configured pool\n"
+                if $child eq '' || index($child, '/') >= 0;
+            $foreign{$dataset} = 1 if $child !~ $RE_NVMET_DATASET_NAME;
+        }
+        my $on = $foreign{$dataset} ? '' : " on '$dataset'";
+        my $key = $nvmet_zfs_keys{$property} // die "unexpected ZFS property$on\n";
+        die "duplicate ZFS property '$property'$on\n" if exists($rows{$dataset}->{$key});
+        if ($key eq 'type') {
+            die "invalid ZFS dataset type$on\n"
+                if ($value ne 'filesystem' && $value ne 'volume') || $source ne '-';
+            $rows{$dataset}->{type} = $value;
+        } else {
+            $rows{$dataset}->{$key} = _nvmet_property_value($value, $source);
+        }
+    }
+    die "ZFS dataset inventory is missing the configured pool\n" if !$rows{$pool};
+
+    my (%types, %volumes);
+    for my $dataset (keys %rows) {
+        my $row = $rows{$dataset};
+        die "incomplete ZFS properties" . ($foreign{$dataset} ? '' : " on '$dataset'") . "\n"
+            if keys($row->%*) != scalar(@nvmet_zfs_props);
+        $types{$dataset} = $row->{type};
+        next if $row->{type} ne 'volume';
+        # nvmet shows device_uuid and udev names the by-id link in lower case,
+        # so identities are compared in lower case.
+        $volumes{$dataset} =
+            $foreign{$dataset}
+            ? { nqn => '-', nsid => '-', uuid => '-' }
+            : { nqn => $row->{nqn}, nsid => $row->{nsid}, uuid => lc($row->{uuid}) };
+    }
+    my $last_nsid = $rows{$pool}->{last_nsid};
+    return {
+        types => \%types,
+        volumes => \%volumes,
+        last_nsid => nvmet_valid_nsid($last_nsid) ? $last_nsid : 0,
+    };
+}
+
+# Parses the configfs read into
+# { subsystems => { nqn => { attr_model, attr_serial, attr_allow_any_host,
+#       acl => { hostnqn => 1 }, namespaces => { nsid => { enable, device_path,
+#       device_uuid, buffered_io } } } },
+#   ports => { id => { addr_trtype, addr_adrfam, addr_traddr, addr_trsvcid,
+#       links => { nqn => 1 } } },
+#   hosts => { hostnqn => { key_sha256 => hex or undef } } }
+# Every object must be complete; unknown objects of newer kernels are ignored.
+sub _nvmet_parse_configfs($text) {
+    my $state = { subsystems => {}, ports => {}, hosts => {} };
+    my (%top, %dir);
+
+    my $subsystem = sub($nqn) {
+        return $state->{subsystems}->{$nqn} //= { acl => {}, namespaces => {} };
+    };
+    my $namespace = sub($nqn, $nsid) {
+        return $subsystem->($nqn)->{namespaces}->{$nsid} //= {};
+    };
+    my $port = sub($id) { return $state->{ports}->{$id} //= { links => {} } };
+    my $host = sub($hostnqn) { return $state->{hosts}->{$hostnqn} //= { key_sha256 => undef } };
+    my $set = sub($object, $attr, $value) {
+        die "duplicate NVMe target configfs attribute\n" if exists($object->{$attr});
+        $object->{$attr} = $value;
+    };
+
+    for my $line (split /\n/, $text) {
+        if ($line =~ $RE_CFS_TOP) {
+            $top{ $+{top} // 'root' } = 1;
+        } elsif ($line =~ $RE_CFS_SUBSYS_DIR) {
+            my $nqn = $+{nqn};
+            $subsystem->($nqn);
+            $dir{"subsystem $nqn"} = 1;
+        } elsif ($line =~ $RE_CFS_SUBSYS_GROUP || $line =~ $RE_CFS_PORT_GROUP) {
+            # default groups, always present with their parent
+        } elsif ($line =~ $RE_CFS_NS_DIR) {
+            my ($nqn, $nsid) = @+{qw(nqn nsid)};
+            $namespace->($nqn, $nsid);
+            $dir{"namespace $nqn/$nsid"} = 1;
+        } elsif ($line =~ $RE_CFS_PORT_DIR) {
+            my $id = $+{port};
+            $port->($id);
+            $dir{"port $id"} = 1;
+        } elsif ($line =~ $RE_CFS_HOST_DIR) {
+            my $hostnqn = $+{host};
+            $host->($hostnqn);
+            $dir{"host $hostnqn"} = 1;
+        } elsif ($line =~ $RE_CFS_PORT_LINK) {
+            my ($id, $nqn) = @+{qw(port nqn)};
+            $port->($id)->{links}->{$nqn} = 1;
+        } elsif ($line =~ $RE_CFS_ACL_LINK) {
+            my ($nqn, $hostnqn) = @+{qw(nqn host)};
+            $subsystem->($nqn)->{acl}->{$hostnqn} = 1;
+        } elsif ($line =~ $RE_CFS_OTHER_ENTRY) {
+            # objects of a shape the plugin does not use
+        } elsif ($line =~ $RE_CFS_NS_ATTR) {
+            my ($nqn, $nsid, $attr, $value) = @+{qw(nqn nsid attr value)};
+            $set->($namespace->($nqn, $nsid), $attr, $value);
+        } elsif ($line =~ $RE_CFS_SUBSYS_ATTR) {
+            my ($nqn, $attr, $value) = @+{qw(nqn attr value)};
+            $set->($subsystem->($nqn), $attr, $value);
+        } elsif ($line =~ $RE_CFS_PORT_ATTR) {
+            my ($id, $attr, $value) = @+{qw(port attr value)};
+            $set->($port->($id), $attr, $value);
+        } elsif ($line =~ $RE_CFS_KEY_DIGEST) {
+            my ($sha, $hostnqn) = @+{qw(sha host)};
+            my $object = $host->($hostnqn);
+            die "duplicate DH-HMAC-CHAP key digest" . nvmet_label($hostnqn) . "\n"
+                if defined($object->{key_sha256});
+            $object->{key_sha256} = $sha;
+        } elsif ($line =~ $RE_CFS_OTHER_ATTR) {
+            # attribute of an object shape the plugin does not use
+        } else {
+            die "unexpected NVMe target configfs line\n";
+        }
+    }
+
+    for my $name (qw(root hosts ports subsystems)) {
+        die "NVMe target configfs is unavailable\n" if !$top{$name};
+    }
+    for my $nqn (keys $state->{subsystems}->%*) {
+        my $subsys = $state->{subsystems}->{$nqn};
+        die "incomplete NVMe subsystem" . nvmet_label($nqn) . "\n"
+            if !$dir{"subsystem $nqn"}
+            || grep { !defined($subsys->{$_}) } qw(attr_model attr_serial attr_allow_any_host);
+        for my $nsid (keys $subsys->{namespaces}->%*) {
+            my $ns = $subsys->{namespaces}->{$nsid};
+            die "incomplete NVMe namespace '$nsid' of subsystem" . nvmet_label($nqn) . "\n"
+                if !$dir{"namespace $nqn/$nsid"}
+                || grep { !defined($ns->{$_}) } qw(enable device_path device_uuid buffered_io);
+        }
+    }
+    for my $id (keys $state->{ports}->%*) {
+        my $object = $state->{ports}->{$id};
+        my @attrs = qw(addr_trtype addr_adrfam addr_traddr addr_trsvcid);
+        die "incomplete NVMe port '$id'\n"
+            if !$dir{"port $id"} || grep { !defined($object->{$_}) } @attrs;
+    }
+    for my $hostnqn (keys $state->{hosts}->%*) {
+        die "DH-HMAC-CHAP key digest for unknown NVMe host" . nvmet_label($hostnqn) . "\n"
+            if !$dir{"host $hostnqn"};
+    }
+    return $state;
+}
+
+sub _nvmet_parse_mounts($text) {
+    for my $line (split /\n/, $text) {
+        my (undef, $mountpoint, $type) = split / /, $line;
+        return 1 if ($mountpoint // '') eq '/sys/kernel/config' && ($type // '') eq 'configfs';
+    }
+    return 0;
+}
+
+# ---------------------------------------------------------------------------
+# Planners (pure)
+# ---------------------------------------------------------------------------
+
+sub _nvmet_serial($nqn) {
+    return 'PVEZFS' . substr(sha256_hex($nqn), 0, 14);
+}
+
+sub _nvmet_template_name($dataset) {
+    die "only VM zvols can become templates\n"
+        if $dataset !~ $RE_NVMET_TEMPLATE_ZVOL || $+{type} ne 'vm';
+    return nvmet_template_twin($dataset);
+}
+
+# The durable identity (NSID, UUID) of a volume owned by $nqn. Refuses a
+# volume that another owned volume duplicates anywhere in the pool.
+sub _nvmet_owned_identity($inv, $nqn, $dataset) {
+    my $type = $inv->{types}->{$dataset} // die "ZFS volume '$dataset' does not exist\n";
+    die "'$dataset' is not a ZFS volume\n" if $type ne 'volume';
+    my $row = $inv->{volumes}->{$dataset};
+    die "ZFS volume '$dataset' is not owned by NVMe subsystem '$nqn'\n" if $row->{nqn} ne $nqn;
+    die "invalid NSID on '$dataset'\n" if !nvmet_valid_nsid($row->{nsid});
+    die "invalid namespace UUID on '$dataset'\n" if $row->{uuid} !~ $RE_NVMET_UUID;
+    for my $other (keys $inv->{volumes}->%*) {
+        next if $other eq $dataset;
+        my $identity = $inv->{volumes}->{$other};
+        next if $identity->{nqn} ne $nqn;
+        die "duplicate NVMe identity on '$dataset'\n"
+            if $identity->{nsid} eq $row->{nsid} || $identity->{uuid} eq $row->{uuid};
+    }
+    return ($row->{nsid}, $row->{uuid});
+}
+
+# { nsid => [uuid, device] } for every volume owned by $nqn, after validating
+# all of them, so a bad identity refuses the whole plan before any change.
+sub _nvmet_desired_namespaces($inv, $nqn) {
+    my (%desired, %uuids);
+    for my $dataset (sort keys $inv->{volumes}->%*) {
+        my $row = $inv->{volumes}->{$dataset};
+        next if $row->{nqn} ne $nqn;
+        die "invalid NSID on '$dataset'\n" if !nvmet_valid_nsid($row->{nsid});
+        die "invalid namespace UUID on '$dataset'\n" if $row->{uuid} !~ $RE_NVMET_UUID;
+        die "duplicate NSID '$row->{nsid}'\n" if $desired{ $row->{nsid} };
+        die "duplicate namespace UUID '$row->{uuid}'\n" if $uuids{ lc($row->{uuid}) }++;
+        $desired{ $row->{nsid} } = [$row->{uuid}, "/dev/zvol/$dataset"];
+    }
+    return \%desired;
+}
+
+# NSIDs grow monotonically, also past the last NSID handed out in the pool,
+# so a host that missed a namespace removal never sees a new volume at an old
+# NSID. Only after the last NSID is used does the allocation fall back to the
+# lowest free one.
+sub _nvmet_allocate_nsid($inv, $cfs, $nqn) {
+    my %used;
+    for my $row (values $inv->{volumes}->%*) {
+        $used{ $row->{nsid} } = 1 if $row->{nqn} eq $nqn && nvmet_valid_nsid($row->{nsid});
+    }
+    $used{$_} = 1 for keys nvmet_namespaces($cfs, $nqn)->%*;
+
+    my $highest = max(0, $inv->{last_nsid} // 0, keys %used);
+    return $highest + 1 if $highest < $nvmet_max_nsid;
+    my $nsid = 1;
+    $nsid++ while $used{$nsid};
+    die "no free namespace ID\n" if $nsid > $nvmet_max_nsid;
+    return $nsid;
+}
+
+# Export state of (NSID, UUID, device) in the subsystem: present, disabled,
+# absent, incomplete (a disabled namespace with another identity, for example
+# one being built), stale (enabled on the other template name of the zvol) or
+# nosubsys. Dies on a conflict.
+sub _nvmet_export_state($cfs, $nqn, $nsid, $uuid, $dev) {
+    return 'nosubsys' if !$cfs->{subsystems}->{$nqn};
+    my $namespaces = nvmet_namespaces($cfs, $nqn);
+    for my $id (keys $namespaces->%*) {
+        next if $id eq $nsid;
+        die "namespace UUID '$uuid' is already in use\n"
+            if $namespaces->{$id}->{device_uuid} eq $uuid;
+    }
+    my $ns = $namespaces->{$nsid} // return 'absent';
+    if ($ns->{device_uuid} eq $uuid && $ns->{device_path} eq $dev) {
+        return $ns->{enable} eq '1' ? 'present' : 'disabled';
+    }
+    # A template conversion or its undo that ssh gave up on can rename the
+    # zvol on the target after an activation exported it under its old name.
+    my $twin = nvmet_template_twin($dev);
+    return 'stale'
+        if $ns->{enable} eq '1'
+        && $ns->{device_uuid} eq $uuid
+        && $twin ne ''
+        && $ns->{device_path} eq $twin;
+    die "NVMe namespace ID '$nsid' has a different identity; refusing to replace it\n"
+        if $ns->{enable} ne '0';
+    return 'incomplete';
+}
+
+# Units that export (NSID, UUID, device) and the state they start from:
+# absent, present, disabled, reclaim or stale. A disabled namespace with
+# another identity, or a stale one, is rebuilt, which is only correct under
+# the target lock. A stale namespace belongs to a template, or to a volume
+# whose template conversion failed, so no guest uses it.
+sub _nvmet_plan_export($cfs, $nqn, $nsid, $uuid, $dev) {
+    die "NVMe subsystem does not exist\n" if !$cfs->{subsystems}->{$nqn};
+    my $state = _nvmet_export_state($cfs, $nqn, $nsid, $uuid, $dev);
+    return ([nvmet_build_unit($nqn, $nsid, $uuid, $dev)], 'absent') if $state eq 'absent';
+    return ([], 'present') if $state eq 'present';
+    return ([nvmet_enable_unit($nqn, $nsid, $dev)], 'disabled') if $state eq 'disabled';
+    my $unexport = nvmet_unexport_unit($nqn, $nsid, $state eq 'stale');
+    my $build = nvmet_build_unit($nqn, $nsid, $uuid, $dev);
+    return ([[$unexport->@*, $build->@*]], $state eq 'stale' ? 'stale' : 'reclaim');
+}
+
+# Units that remove the namespace exporting $uuid, and its previous state.
+sub _nvmet_plan_unexport($cfs, $nqn, $uuid) {
+    my $namespaces = nvmet_namespaces($cfs, $nqn);
+    my @matches = grep { $namespaces->{$_}->{device_uuid} eq $uuid }
+        sort { $a <=> $b } keys $namespaces->%*;
+    die "duplicate namespace UUID '$uuid'\n" if @matches > 1;
+    return ([], undef) if !@matches;
+    my $enabled = $namespaces->{ $matches[0] }->{enable} eq '1';
+    return (
+        [nvmet_unexport_unit($nqn, $matches[0], $enabled)],
+        { nsid => $matches[0], enabled => $enabled },
+    );
+}
+
+# The port serving a portal ({ family, address, port } of parse_nvme_portals).
+# Among complete, matching ports, the one already linked to $nqn wins, then
+# the lowest id. Incomplete ports never match.
+sub _nvmet_find_port($cfs, $nqn, $portal) {
+    my $wanted = nvmet_canonical_address($portal->{address});
+    my @matches = grep {
+        my $port = $cfs->{ports}->{$_};
+        $port->{addr_trtype} eq 'tcp'
+            && $port->{addr_adrfam} eq $portal->{family}
+            && $port->{addr_traddr} ne ''
+            && $port->{addr_trsvcid} eq "$portal->{port}"
+            && nvmet_canonical_address($port->{addr_traddr}) eq $wanted
+        } sort {
+            $a <=> $b
+        } keys $cfs->{ports}->%*;
+    my @linked = grep { nvmet_port_linked($cfs, $_, $nqn) } @matches;
+    return $linked[0] // $matches[0];
+}
+
+# Units that bring the target to the configured state, except publishing:
+# objects (subsystem, ports, hosts), keys, ACLs, namespaces in NSID order,
+# then removal of undesired namespaces.
+# $conf: { nqn, portals (of parse_nvme_portals), hostnqns, keysha }
+sub _nvmet_plan_activation($inv, $cfs, $conf) {
+    my ($nqn, $keysha) = $conf->@{qw(nqn keysha)};
+    my $desired = _nvmet_desired_namespaces($inv, $nqn); # refuse before any unit
+    my $subsys_path = nvmet_subsys_path($nqn);
+    my $subsys = $cfs->{subsystems}->{$nqn};
+    my (@objects, @key_hosts, @acls, @namespaces, %port_ids);
+
+    if (!$subsys) {
+        push @objects,
+            [
+                ['mkdir', $subsys_path],
+                nvmet_write("$subsys_path/attr_model", $nvmet_model),
+                nvmet_write("$subsys_path/attr_serial", _nvmet_serial($nqn)),
+                nvmet_write("$subsys_path/attr_allow_any_host", 0),
+            ];
+    } else {
+        my $serial = _nvmet_serial($nqn);
+        my @attrs;
+        if ($subsys->{attr_model} ne $nvmet_model || $subsys->{attr_serial} ne $serial) {
+            die "refusing to take over existing NVMe subsystem '$nqn': it was not created by"
+                . " Proxmox VE; remove it on the target or use another NQN\n"
+                if !nvmet_unfinished_subsystem($cfs, $nqn);
+            push @attrs, nvmet_write("$subsys_path/attr_serial", $serial);
+        }
+        push @attrs, nvmet_write("$subsys_path/attr_allow_any_host", 0)
+            if $subsys->{attr_allow_any_host} ne '0';
+        push @objects, [@attrs] if @attrs;
+    }
+
+    my %taken = map { $_ => 1 } keys $cfs->{ports}->%*;
+    for my $portal ($conf->{portals}->@*) {
+        my $id = _nvmet_find_port($cfs, $nqn, $portal);
+        if (!defined($id)) {
+            $id = 1;
+            $id++ while $taken{$id};
+            $taken{$id} = 1;
+            my $port_path = nvmet_port_path($id);
+            push @objects,
+                [
+                    ['mkdir', $port_path],
+                    nvmet_write("$port_path/addr_trtype", 'tcp'),
+                    nvmet_write("$port_path/addr_adrfam", $portal->{family}),
+                    nvmet_write("$port_path/addr_traddr", $portal->{address}),
+                    nvmet_write("$port_path/addr_trsvcid", $portal->{port}),
+                ];
+        }
+        $port_ids{$id} = 1;
+    }
+
+    # nvmet keeps the key on the global host object. A key that differs from
+    # ours is only replaced while no subsystem links the host.
+    my $linked = nvmet_linked_hosts($cfs);
+    for my $hostnqn ($conf->{hostnqns}->@*) {
+        my $host_path = nvmet_host_path($hostnqn);
+        if (my $host = $cfs->{hosts}->{$hostnqn}) {
+            my $sha = $host->{key_sha256}
+                // die "NVMe target does not expose a DH-HMAC-CHAP key for host '$hostnqn'\n";
+            if ($sha ne $keysha) {
+                die "refusing to replace an in-use DH-HMAC-CHAP key for '$hostnqn'\n"
+                    if $linked->{$hostnqn};
+                push @key_hosts, $host_path;
+            }
+        } else {
+            push @objects, [['mkdir', $host_path]];
+            push @key_hosts, $host_path;
+        }
+        push @acls, [['ln', '-s', $host_path, "$subsys_path/allowed_hosts/$hostnqn"]]
+            if !$subsys || !($subsys->{acl} // {})->{$hostnqn};
+    }
+    # nvmet creates the key attributes world-readable. The chmod before the
+    # write keeps every later open from reading the key; a descriptor opened
+    # while an attribute was still world-readable is not affected by it.
+    my @keys;
+    if (@key_hosts) {
+        my @attrs = map { ("$_/dhchap_key", "$_/dhchap_ctrl_key") } @key_hosts;
+        @keys = ([['chmod', '0600', @attrs], { key => [map { "$_/dhchap_key" } @key_hosts] }]);
+    }
+
+    my $current = { subsystems => { $nqn => $subsys // { acl => {}, namespaces => {} } } };
+    for my $nsid (sort { $a <=> $b } keys $desired->%*) {
+        my ($units) = _nvmet_plan_export($current, $nqn, $nsid, $desired->{$nsid}->@*);
+        push @namespaces, $units->@*;
+    }
+    my $existing = nvmet_namespaces($cfs, $nqn);
+    for my $nsid (sort { $a <=> $b } keys $existing->%*) {
+        next if $desired->{$nsid};
+        push @namespaces, nvmet_unexport_unit($nqn, $nsid, $existing->{$nsid}->{enable} eq '1');
+    }
+
+    return {
+        prepublish => [@objects, @keys, @acls, @namespaces],
+        port_ids => [sort { $a <=> $b } keys %port_ids],
+    };
+}
+
+# One call per port: each link repeats the guards, so it only succeeds while
+# the subsystem denies unknown hosts and every configured host has its ACL.
+# Under the target lock, only a manager of the target outside this cluster
+# can change that after the verify read; the guards stop the publish then.
+# Links to ports that are no longer configured are removed.
+# Returns [{ port => $id, action => 'link' | 'unlink', steps => [...] }, ...].
+sub _nvmet_plan_publish($cfs, $nqn, $port_ids, $hostnqns) {
+    my $subsys_path = nvmet_subsys_path($nqn);
+    my %wanted = map { $_ => 1 } $port_ids->@*;
+    my @guards = (
+        ['grep', '-qx', '0', "$subsys_path/attr_allow_any_host"],
+        map { ['test', '-L', "$subsys_path/allowed_hosts/$_"] } $hostnqns->@*,
+    );
+
+    my @units;
+    for my $id ($port_ids->@*) {
+        next if nvmet_port_linked($cfs, $id, $nqn);
+        my $link = ['ln', '-s', $subsys_path, nvmet_port_path($id) . "/subsystems/$nqn"];
+        push @units, { port => $id, action => 'link', steps => [@guards, $link] };
+    }
+    for my $id (sort { $a <=> $b } keys $cfs->{ports}->%*) {
+        next if $wanted{$id} || !nvmet_port_linked($cfs, $id, $nqn);
+        my $unlink = ['rm', nvmet_port_path($id) . "/subsystems/$nqn"];
+        push @units, { port => $id, action => 'unlink', steps => [$unlink] };
+    }
+    return \@units;
+}
+
+# Teardown units for the subsystem and the hosts that may become orphans.
+sub _nvmet_plan_delete_target($inv, $cfs, $nqn, $hostnqns) {
+    for my $dataset (sort keys $inv->{volumes}->%*) {
+        die "refusing to delete NVMe subsystem '$nqn': owned ZFS volume '$dataset' exists\n"
+            if $inv->{volumes}->{$dataset}->{nqn} eq $nqn;
+    }
+    my $subsys = $cfs->{subsystems}->{$nqn} // return ([], [$hostnqns->@*]);
+    die "refusing to delete foreign NVMe subsystem '$nqn'\n"
+        if $subsys->{attr_model} ne $nvmet_model || $subsys->{attr_serial} ne _nvmet_serial($nqn);
+    die "refusing to delete NVMe subsystem '$nqn': namespaces remain\n"
+        if nvmet_namespaces($cfs, $nqn)->%*;
+
+    my $subsys_path = nvmet_subsys_path($nqn);
+    my @acl = sort keys(($subsys->{acl} // {})->%*);
+    die "refusing to delete NVMe subsystem '$nqn': it allows a malformed host name\n"
+        if grep { $_ !~ $RE_NQN } @acl;
+    my @links = (
+        (
+            map { nvmet_port_path($_) . "/subsystems/$nqn" }
+            grep { nvmet_port_linked($cfs, $_, $nqn) }
+            sort { $a <=> $b } keys $cfs->{ports}->%*
+        ),
+        (map { "$subsys_path/allowed_hosts/$_" } @acl),
+    );
+    # A retry after a teardown that failed at its rmdir finds no ACL left, so
+    # the configured hosts are candidates as well.
+    my %candidates = map { $_ => 1 } @acl, $hostnqns->@*;
+    return (
+        [[(@links ? (['rm', @links]) : ()), ['rmdir', $subsys_path]]], [sort keys %candidates],
+    );
+}
+
+# Candidate hosts that exist and that no subsystem links any more.
+sub _nvmet_plan_orphan_hosts($cfs, $candidates) {
+    my $linked = nvmet_linked_hosts($cfs);
+    my %seen;
+    my @orphans = map { nvmet_host_path($_) }
+        grep { !$seen{$_}++ && $_ =~ $RE_NQN && $cfs->{hosts}->{$_} && !$linked->{$_} }
+        sort $candidates->@*;
+    return @orphans ? [[['rmdir', @orphans]]] : [];
+}
+
+# ---------------------------------------------------------------------------
+# Target lock and remote execution
+# ---------------------------------------------------------------------------
+
+my %nvmet_lock_owner; # lock id => pid of the process holding the domain lock
+
+# Runs $code under the pmxcfs domain lock of the target. Like the pmxcfs
+# lock itself, it is not re-entrant, and a child forked inside is not the
+# owner.
+my sub nvmet_locked($scfg, $code) {
+    my $id = nvmet_lock_id($scfg);
+    die "cluster not quorate - refusing NVMe target changes\n"
+        if !PVE::Cluster::check_cfs_quorum(1);
+    my $res = PVE::Cluster::cfs_lock_domain(
+        $id,
+        $nvmet_lock_wait,
+        sub {
+            local $nvmet_lock_owner{$id} = $$;
+            return $code->();
+        },
+    );
+    die $@ if $@;
+    return $res;
+}
+
+# Runs steps on the target. %opts: op, timeout, input, change (the steps
+# change the target, which needs the lock).
+#
+# ssh does not stop a command on the target when it gives up on it: a change
+# whose connection broke (exit code 255), any call under the lock that ssh
+# did not run to its end (-1: its timeout, or ssh was killed), and a call
+# during which the task was stopped (it dies with the task marker) may still
+# be running there. Its outcome is unknown, so it is never compensated from a
+# read that could come before the rest of it. The next activation repairs
+# from the target state.
+my sub nvmet_exec($scfg, $steps, %opts) {
+    my $locked = ($nvmet_lock_owner{ nvmet_lock_id($scfg) } // 0) == $$;
+    die "internal error: NVMe target change without target lock\n"
+        if $opts{change} && !$locked;
+    my $res = _nvmet_run(
+        $scfg, $steps,
+        op => $opts{op},
+        timeout => $opts{timeout} // 15,
+        (defined($opts{input}) ? (input => $opts{input}) : ()),
+    );
+    die "NVMe target operation '$opts{op}' did not complete ($res->{err});"
+        . " the target state is unknown\n"
+        if ($locked && $res->{rc} == -1)
+        || ($opts{change} && $res->{rc} == 255);
+    return $res;
+}
+
+# An abandoned operation is passed on instead of being compensated.
+my sub nvmet_rethrow_abandoned($error) {
+    die $error if $error =~ $RE_NVMET_ABANDONED;
+}
+
+my sub nvmet_unreachable($res) {
+    return $res->{rc} == 255 || $res->{rc} == -1;
+}
+
+my sub nvmet_check($res, $op) {
+    die "NVMe target operation '$op' failed: $res->{err}\n" if $res->{rc};
+}
+
+my sub nvmet_read_failed($scfg, $res) {
+    die "NVMe target '" . nvmet_server($scfg) . "' is unreachable: $res->{err}\n"
+        if nvmet_unreachable($res);
+    die "cannot read NVMe target state: $res->{err}\n";
+}
+
+my sub nvmet_output($res) {
+    return join('', map { "$_\n" } $res->{out}->@*);
+}
+
+# A read, retried twice after 200 ms unless the target is unreachable: a read
+# without the lock can meet a configfs object that another node removes
+# while find lists it, which makes find fail once.
+my sub nvmet_read_call($scfg, $steps, %opts) {
+    my $res;
+    for my $attempt (0 .. 2) {
+        _sleep(0.2) if $attempt;
+        $res = nvmet_exec($scfg, $steps, %opts, timeout => 15);
+        last if !$res->{rc} || nvmet_unreachable($res);
+    }
+    return $res;
+}
+
+# Reads the target state and returns (ZFS inventory, configfs state).
+# keys => [hostnqns] adds the digests of their keys. modprobe => 1, for locked
+# reads only, loads nvmet first and mounts configfs if that is what failed.
+# With tolerant => 1, a failed or unparsable read returns an empty list,
+# unless the target is unreachable.
+my sub nvmet_read($scfg, %opts) {
+    my $pool = nvmet_pool($scfg);
+    my $hostnqns = $opts{keys} // [];
+    # loading a module that may still be loading is harmless
+    my @modprobe = $opts{modprobe} ? (change => 1) : ();
+    my $read = sub () {
+        return nvmet_read_call(
+            $scfg,
+            nvmet_state_steps($pool, $hostnqns, $opts{modprobe}),
+            op => 'read target state',
+            @modprobe,
+        );
+    };
+    my $res = $read->();
+    if ($res->{rc} && $opts{modprobe} && !nvmet_unreachable($res)) {
+        my $mounts = nvmet_exec($scfg, [['cat', '/proc/mounts']], op => 'read mounts');
+        nvmet_check($mounts, 'read mounts');
+        if (!_nvmet_parse_mounts(nvmet_output($mounts))) {
+            my $mount = ['mount', '-t', 'configfs', 'none', '/sys/kernel/config'];
+            my $mounted = nvmet_exec($scfg, [$mount], op => 'mount configfs', change => 1);
+            nvmet_check($mounted, 'mount configfs');
+            $res = $read->();
+        }
+    }
+    if ($res->{rc}) {
+        return () if $opts{tolerant} && !nvmet_unreachable($res);
+        nvmet_read_failed($scfg, $res);
+    }
+    my @state = eval {
+        my ($zfs, $configfs) = _nvmet_split_state(nvmet_output($res));
+        (_nvmet_parse_zfs_inventory($zfs, $pool), _nvmet_parse_configfs($configfs));
+    };
+    if (my $error = $@) {
+        return () if $opts{tolerant};
+        die $error;
+    }
+    return @state;
+}
+
+my sub nvmet_read_zfs($scfg) {
+    my $pool = nvmet_pool($scfg);
+    my $res = nvmet_read_call($scfg, [nvmet_zfs_inventory_step($pool)], op => 'read ZFS inventory');
+    nvmet_read_failed($scfg, $res) if $res->{rc};
+    return _nvmet_parse_zfs_inventory(nvmet_output($res), $pool);
+}
+
+# Applies units in as few calls as possible and stops at the first failing
+# call, whose result it returns. A chain only saves SSH round trips: every
+# decision was taken locally before, and `&&` only stops at the first
+# failing command. Only the call with the key step gets stdin. The default
+# timeout suits configfs chains; ZFS changes pass their own.
+my sub nvmet_apply($scfg, $units, %opts) {
+    my $input = delete $opts{input};
+    $opts{timeout} //= 30;
+    for my $chunk (_nvmet_chunk($units)->@*) {
+        my $with_key = grep { ref($_) eq 'HASH' && exists($_->{key}) } $chunk->@*;
+        my @input = $with_key ? (input => $input) : ();
+        my $res = nvmet_exec($scfg, $chunk, %opts, @input, change => 1);
+        return $res if $res->{rc};
+    }
+    return { rc => 0, out => [], err => '' };
+}
+
+my sub nvmet_change($scfg, $units, %opts) {
+    nvmet_check(nvmet_apply($scfg, $units, %opts), $opts{op});
+}
+
+my sub nvmet_wait_device($scfg, $dev) {
+    my $deadline = _now() + 10;
+    while (1) {
+        my $res = nvmet_exec($scfg, [['test', '-b', $dev]], op => 'wait for zvol');
+        return if !$res->{rc};
+        nvmet_read_failed($scfg, $res) if nvmet_unreachable($res);
+        last if _now() >= $deadline;
+        _sleep(0.25);
+    }
+    die "zvol '$dev' is not a block device after 10 seconds\n";
+}
+
+my sub nvmet_new_uuid($inv, $cfs) {
+    my %used = map { lc($_->{uuid}) => 1 } values $inv->{volumes}->%*;
+    for my $subsys (values $cfs->{subsystems}->%*) {
+        $used{ lc($_->{device_uuid}) } = 1 for values(($subsys->{namespaces} // {})->%*);
+    }
+    for (1 .. 10) {
+        my $uuid = lc(file_read_firstline('/proc/sys/kernel/random/uuid') // '');
+        return $uuid if $uuid =~ $RE_NVMET_UUID && !$used{$uuid};
+    }
+    die "cannot generate an unused namespace UUID\n";
+}
+
+my sub nvmet_has_identity($row, $nqn, $nsid, $uuid) {
+    return $row && $row->{nqn} eq $nqn && $row->{nsid} eq "$nsid" && $row->{uuid} eq $uuid;
+}
+
+# The configfs state after a successful unexport of namespace $pre->{nsid}.
+my sub nvmet_without_namespace($cfs, $nqn, $pre) {
+    return $cfs if !$pre;
+    my $subsys = $cfs->{subsystems}->{$nqn};
+    my %namespaces = $subsys->{namespaces}->%*;
+    delete $namespaces{ $pre->{nsid} };
+    my $changed = { $subsys->%*, namespaces => \%namespaces };
+    return { $cfs->%*, subsystems => { $cfs->{subsystems}->%*, $nqn => $changed } };
+}
+
+# Links the subsystem to each desired port separately. One port that cannot
+# be bound (for example, its address is missing on the target) only warns
+# while another desired port serves the subsystem.
+my sub nvmet_publish($scfg, $cfs, $nqn, $port_ids, $hostnqns) {
+    my @failed;
+    for my $unit (_nvmet_plan_publish($cfs, $nqn, $port_ids, $hostnqns)->@*) {
+        my $op = "$unit->{action} NVMe subsystem on port $unit->{port}";
+        my $res = nvmet_apply($scfg, [$unit->{steps}], op => $op);
+        push @failed, { $unit->%*, err => $res->{err} } if $res->{rc};
+    }
+    return if !@failed;
+
+    my (undef, $fresh) = eval { nvmet_read($scfg) };
+    nvmet_rethrow_abandoned($@) if $@;
+    my $published = $fresh && grep { nvmet_port_linked($fresh, $_, $nqn) } $port_ids->@*;
+    my ($first) = grep { $_->{action} eq 'link' } @failed;
+    $first //= $failed[0];
+    die "cannot publish NVMe subsystem '$nqn': $first->{err}\n" if !$published;
+    for my $failure (@failed) {
+        my $what = $failure->{action} eq 'link' ? 'publish' : 'unpublish';
+        log_warn("cannot $what NVMe subsystem '$nqn' on port $failure->{port}: $failure->{err}");
+    }
+}
+
+# ---------------------------------------------------------------------------
+# Target flows
+# ---------------------------------------------------------------------------
+
+# Brings the target to the configured state and publishes the subsystem.
+# A converged target costs one read and no lock. Otherwise, under the lock:
+# read, apply, verify with a fresh read, and publish only once the verify
+# read plans nothing more. $portals and $hostnqns are the parsed and validated
+# storage configuration.
+sub _nvmet_activate_target($storeid, $scfg, $portals, $hostnqns, $key) {
+    my $nqn = nvmet_nqn($scfg);
+    my $conf = {
+        nqn => $nqn,
+        portals => $portals,
+        hostnqns => $hostnqns,
+        keysha => sha256_hex("$key\n"),
+    };
+    my $converged = sub($inv, $cfs) {
+        my $plan = _nvmet_plan_activation($inv, $cfs, $conf);
+        return !$plan->{prepublish}->@*
+            && !_nvmet_plan_publish($cfs, $nqn, $plan->{port_ids}, $hostnqns)->@*;
+    };
+
+    # Without the lock, a refusal or failed read only means "take the lock".
+    my @state = nvmet_read($scfg, keys => $hostnqns, tolerant => 1);
+    return if @state && eval { $converged->(@state) };
+
+    nvmet_locked(
+        $scfg,
+        sub {
+            # Removing the storage deletes its key before it removes the target
+            # configuration, so an activation that waited for the lock meanwhile
+            # does not restore it.
+            die "storage '$storeid' is being removed or its key changed\n"
+                if (file_read_firstline(secret_path($storeid)) // '') ne $key;
+            my ($inv, $cfs) = nvmet_read($scfg, keys => $hostnqns, modprobe => 1);
+            my $plan = _nvmet_plan_activation($inv, $cfs, $conf);
+            my $error;
+            for my $round (1 .. 2) {
+                last if !$plan->{prepublish}->@*;
+                my $apply = nvmet_apply(
+                    $scfg, $plan->{prepublish},
+                    op => 'configure NVMe target',
+                    input => "$key\n",
+                );
+                $error //= $apply->{err} if $apply->{rc};
+                ($inv, $cfs) = nvmet_read($scfg, keys => $hostnqns);
+                $plan = _nvmet_plan_activation($inv, $cfs, $conf);
+            }
+            die "NVMe target did not converge: " . ($error // 'state mismatch') . "\n"
+                if $plan->{prepublish}->@*;
+            nvmet_publish($scfg, $cfs, $nqn, $plan->{port_ids}, $hostnqns);
+        },
+    );
+    return;
+}
+
+# The namespace UUID of an owned volume, from one read-only ZFS query.
+sub _nvmet_volume_uuid($scfg, $name) {
+    my $dataset = nvmet_dataset($scfg, $name);
+    my $inv = nvmet_read_zfs($scfg);
+    my (undef, $uuid) = _nvmet_owned_identity($inv, nvmet_nqn($scfg), $dataset);
+    return $uuid;
+}
+
+# Creates (size => KiB) or clones (origin => pool/base@snap) a zvol with its
+# identity set at creation, then exports it. A failure removes only what this
+# call created, decided from a fresh read.
+my sub nvmet_create_volume($scfg, $name, %opts) {
+    my $nqn = nvmet_nqn($scfg);
+    my $dataset = nvmet_dataset($scfg, $name);
+    my $dev = "/dev/zvol/$dataset";
+    die "internal error: invalid zvol creation request\n"
+        if defined($opts{size}) == defined($opts{origin})
+        || (defined($opts{size}) && $opts{size} !~ $RE_UNSIGNED_INTEGER);
+
+    return nvmet_locked(
+        $scfg,
+        sub {
+            my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1);
+            die "zvol '$dataset' already exists\n" if exists($inv->{types}->{$dataset});
+            die "NVMe subsystem does not exist; activate the storage first\n"
+                if !$cfs->{subsystems}->{$nqn};
+            my $nsid = _nvmet_allocate_nsid($inv, $cfs, $nqn);
+            my $uuid = nvmet_new_uuid($inv, $cfs);
+            my (undef, $state) = _nvmet_plan_export($cfs, $nqn, $nsid, $uuid, $dev);
+            die "internal error: namespace ID '$nsid' is in use\n" if $state ne 'absent';
+
+            # Record the NSID on the pool before the zvol exists: a creation that
+            # completes on the target only after this call gave up on it must
+            # never share its NSID with a later volume.
+            my $timeout = nvmet_long_timeout();
+            if ($nsid > $inv->{last_nsid}) {
+                my $reserve = ['zfs', 'set', "$nvmet_last_nsid=$nsid", nvmet_pool($scfg)];
+                nvmet_change(
+                    $scfg, [[$reserve]],
+                    op => 'reserve namespace ID',
+                    timeout => $timeout,
+                );
+            }
+
+            my @identity = map { ('-o', $_) } nvmet_identity($nqn, $nsid, $uuid);
+            my @create_opts = (
+                ($scfg->{sparse} ? ('-s') : ()),
+                ($scfg->{blocksize} ? ('-b', $scfg->{blocksize}) : ()),
+            );
+            my $create =
+                defined($opts{origin})
+                ? ['zfs', 'clone', @identity, $opts{origin}, $dataset]
+                : ['zfs', 'create', @create_opts, @identity, '-V', "$opts{size}k", $dataset];
+            my $res = nvmet_apply(
+                $scfg, [[$create]],
+                op => 'create zvol',
+                timeout => $timeout,
+            );
+            if ($res->{rc}) {
+                # A definite command failure may still have created the zvol.
+                # Only a zvol with this call's identity counts. Ambiguous SSH
+                # outcomes have already aborted the transaction in nvmet_exec.
+                my $check = eval { nvmet_read_zfs($scfg) };
+                nvmet_rethrow_abandoned($@) if $@;
+                my $row = $check ? $check->{volumes}->{$dataset} : undef;
+                die "NVMe target operation 'create zvol' failed: $res->{err}\n"
+                    if !nvmet_has_identity($row, $nqn, $nsid, $uuid);
+            }
+
+            my $verified;
+            eval {
+                nvmet_wait_device($scfg, $dev);
+                my $build_unit = nvmet_build_unit($nqn, $nsid, $uuid, $dev);
+                my $build = nvmet_apply($scfg, [$build_unit], op => 'export namespace');
+                $verified = [nvmet_read($scfg)];
+                my ($vinv, $vcfs) = $verified->@*;
+                my ($vnsid, $vuuid) = _nvmet_owned_identity($vinv, $nqn, $dataset);
+                die "NVMe identity of zvol '$dataset' changed\n"
+                    if $vnsid ne $nsid || $vuuid ne $uuid;
+                if (_nvmet_export_state($vcfs, $nqn, $nsid, $uuid, $dev) ne 'present') {
+                    nvmet_check($build, 'export namespace');
+                    die "NVMe namespace '$nsid' was not exported\n";
+                }
+            };
+            my $error = $@ or return $uuid;
+            nvmet_rethrow_abandoned($error);
+
+            eval {
+                my ($cinv, $ccfs) = $verified ? $verified->@* : nvmet_read($scfg);
+                my $namespaces = nvmet_namespaces($ccfs, $nqn);
+                # The build writes the UUID before it enables the namespace, so an
+                # enabled namespace without our UUID was not built by this call.
+                my @unexport =
+                    map { nvmet_unexport_unit($nqn, $_, $namespaces->{$_}->{enable} eq '1') }
+                    grep {
+                        my $ns = $namespaces->{$_};
+                        $ns->{device_uuid} eq $uuid
+                        || ($ns->{enable} ne '1' && $ns->{device_path} eq $dev)
+                    } sort {
+                    $a <=> $b
+                    } keys $namespaces->%*;
+                nvmet_change($scfg, \@unexport, op => 'remove namespace') if @unexport;
+                # -o at creation proves that this call created the zvol
+                if (nvmet_has_identity($cinv->{volumes}->{$dataset}, $nqn, $nsid, $uuid)) {
+                    my $destroy = ['zfs', 'destroy', '-r', $dataset];
+                    nvmet_change(
+                        $scfg, [[$destroy]],
+                        op => 'destroy zvol',
+                        timeout => $timeout,
+                    );
+                }
+            };
+            log_warn("failed to clean up zvol '$name': $@") if $@;
+            die $error;
+        },
+    );
+}
+
+# Unexports and destroys an owned zvol. A missing zvol is already destroyed.
+my sub nvmet_destroy_volume($scfg, $name) {
+    my $nqn = nvmet_nqn($scfg);
+    my $dataset = nvmet_dataset($scfg, $name);
+    my $dev = "/dev/zvol/$dataset";
+    my $destroy = ['zfs', 'destroy', '-r', $dataset];
+    my $timeout = nvmet_long_timeout();
+
+    nvmet_locked(
+        $scfg,
+        sub {
+            my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1);
+            return if !exists($inv->{types}->{$dataset});
+            my ($nsid, $uuid) = _nvmet_owned_identity($inv, $nqn, $dataset);
+            my ($unexport, $pre) = ([], undef);
+            if ($cfs->{subsystems}->{$nqn}) {
+                _nvmet_plan_export($cfs, $nqn, $nsid, $uuid, $dev); # refusals only
+                ($unexport, $pre) = _nvmet_plan_unexport($cfs, $nqn, $uuid);
+            }
+            my $units = [$unexport->@*, [$destroy]];
+            my $res = nvmet_apply($scfg, $units, op => 'destroy zvol', timeout => $timeout);
+            return if !$res->{rc};
+            my $error = $res->{err};
+
+            # Decide from a fresh read what the failed call changed.
+            my ($finv, $fcfs) = eval { nvmet_read($scfg) };
+            if (my $read_error = $@) {
+                nvmet_rethrow_abandoned($read_error);
+                die "failed to destroy '$dataset': $error\n";
+            }
+            return if !exists($finv->{types}->{$dataset});
+            my $namespaces = nvmet_namespaces($fcfs, $nqn);
+            my ($current) =
+                grep { $namespaces->{$_}->{device_uuid} eq $uuid } keys $namespaces->%*;
+            if (defined($current)) {
+                die "cannot disable namespace '$current': $error\n"
+                    if $namespaces->{$current}->{enable} eq '1';
+                die "cannot remove namespace '$current': $error\n" if !$pre || !$pre->{enabled};
+                my $enable = [nvmet_enable_unit($nqn, $current, $dev)];
+                my $restore = nvmet_apply($scfg, $enable, op => 'enable namespace');
+                die "cannot remove namespace '$current': $error"
+                    . ($restore->{rc} ? "; cannot restore namespace: $restore->{err}" : '') . "\n";
+            }
+
+            # The namespace is gone and the zvol is still there, typically busy.
+            for my $attempt (1 .. 5) {
+                _sleep(1);
+                my $retry =
+                    nvmet_apply($scfg, [[$destroy]], op => 'destroy zvol', timeout => $timeout);
+                return if !$retry->{rc};
+                $error = $retry->{err};
+            }
+            die "failed to destroy '$dataset': $error\n" if !$pre;
+            eval {
+                my (undef, $rcfs) = nvmet_read($scfg);
+                my ($units) = _nvmet_plan_export($rcfs, $nqn, $nsid, $uuid, $dev);
+                nvmet_change($scfg, $units, op => 'restore namespace');
+            };
+            die "failed to destroy '$dataset': $error; namespace restoration also failed: $@"
+                if $@;
+            die "failed to destroy '$dataset': $error; namespace restored\n";
+        },
+    );
+    return;
+}
+
+# Rolls an owned zvol back and always exports it again with the same identity.
+my sub nvmet_rollback_volume($scfg, $name, $snap) {
+    my $snapshot = nvmet_snapshot($scfg, $name, $snap);
+    my $nqn = nvmet_nqn($scfg);
+    my $dataset = nvmet_dataset($scfg, $name);
+    my $dev = "/dev/zvol/$dataset";
+
+    my $timeout = nvmet_long_timeout();
+
+    nvmet_locked(
+        $scfg,
+        sub {
+            my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1);
+            my ($nsid, $uuid) = _nvmet_owned_identity($inv, $nqn, $dataset);
+            die "NVMe subsystem does not exist\n" if !$cfs->{subsystems}->{$nqn};
+            _nvmet_plan_export($cfs, $nqn, $nsid, $uuid, $dev); # refusals only
+            my ($unexport, $pre) = _nvmet_plan_unexport($cfs, $nqn, $uuid);
+            my $set_identity = ['zfs', 'set', nvmet_identity($nqn, $nsid, $uuid), $dataset];
+            my $rollback = ['zfs', 'rollback', $snapshot];
+            my $res = nvmet_apply(
+                $scfg, [$unexport->@*, [$rollback], [$set_identity]],
+                op => 'rollback zvol',
+                timeout => $timeout,
+            );
+            my $error;
+            $error = "NVMe target operation 'rollback zvol' failed: $res->{err}" if $res->{rc};
+
+            # Always export the volume again, from a fresh read if anything failed.
+            eval {
+                my ($rinv, $rcfs) =
+                    $error ? nvmet_read($scfg) : ($inv, nvmet_without_namespace($cfs, $nqn, $pre));
+                nvmet_wait_device($scfg, $dev);
+                my @units;
+                push @units, [$set_identity]
+                    if !nvmet_has_identity($rinv->{volumes}->{$dataset}, $nqn, $nsid, $uuid);
+                my ($export) = _nvmet_plan_export($rcfs, $nqn, $nsid, $uuid, $dev);
+                push @units, $export->@*;
+                nvmet_change($scfg, \@units, op => 'restore namespace', timeout => $timeout);
+            };
+            chomp(my $restore_error = $@);
+
+            die "rollback failed: $error; namespace restoration failed: $restore_error\n"
+                if defined($error) && $restore_error;
+            die "namespace restoration failed after rollback: $restore_error\n"
+                if $restore_error;
+            die "$error\n" if defined($error);
+        },
+    );
+    return;
+}
+
+# Renames vm-* to base-*, exports it under the same identity and takes the
+# __base__ snapshot. A failed conversion is undone from a fresh read.
+my sub nvmet_template_volume($scfg, $name) {
+    my $nqn = nvmet_nqn($scfg);
+    my $old = nvmet_dataset($scfg, $name);
+    my $new = _nvmet_template_name($old);
+    my ($old_dev, $new_dev) = ("/dev/zvol/$old", "/dev/zvol/$new");
+    my $timeout = nvmet_long_timeout();
+
+    nvmet_locked(
+        $scfg,
+        sub {
+            my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1);
+            my ($nsid, $uuid) = _nvmet_owned_identity($inv, $nqn, $old);
+            die "template zvol '$new' already exists\n" if exists($inv->{types}->{$new});
+            die "NVMe subsystem does not exist\n" if !$cfs->{subsystems}->{$nqn};
+            _nvmet_plan_export($cfs, $nqn, $nsid, $uuid, $old_dev); # refusals only
+            my ($unexport) = _nvmet_plan_unexport($cfs, $nqn, $uuid);
+
+            eval {
+                my $rename = [$unexport->@*, [['zfs', 'rename', $old, $new]]];
+                nvmet_change($scfg, $rename, op => 'rename zvol', timeout => $timeout);
+                nvmet_wait_device($scfg, $new_dev);
+                my $build = nvmet_build_unit($nqn, $nsid, $uuid, $new_dev);
+                my $template = [$build, [['zfs', 'snapshot', "$new\@__base__"]]];
+                nvmet_change($scfg, $template, op => 'create template', timeout => $timeout);
+            };
+            my $error = $@ or return;
+            nvmet_rethrow_abandoned($error);
+            chomp($error);
+
+            eval {
+                my ($rinv, $rcfs) = nvmet_read($scfg);
+                if (exists($rinv->{types}->{$new})) {
+                    my ($units) = _nvmet_plan_unexport($rcfs, $nqn, $uuid);
+                    my $back = [$units->@*, [['zfs', 'rename', $new, $old]]];
+                    nvmet_change($scfg, $back, op => 'rename zvol back', timeout => $timeout);
+                    (undef, $rcfs) = nvmet_read($scfg);
+                } elsif (!exists($rinv->{types}->{$old})) {
+                    die "zvol '$old' does not exist\n";
+                }
+                nvmet_wait_device($scfg, $old_dev);
+                my ($units) = _nvmet_plan_export($rcfs, $nqn, $nsid, $uuid, $old_dev);
+                nvmet_change($scfg, $units, op => 'restore namespace');
+            };
+            die "template conversion failed: $error; restoration failed: $@" if $@;
+            die "template conversion failed: $error\n";
+        },
+    );
+    return;
+}
+
+# Grows an owned zvol and lets an exported namespace pick up the new size.
+my sub nvmet_resize_volume($scfg, $name, $size) {
+    my $nqn = nvmet_nqn($scfg);
+    my $dataset = nvmet_dataset($scfg, $name);
+    die "internal error: invalid zvol size\n" if $size !~ $RE_UNSIGNED_INTEGER;
+
+    # The volume is checked before it changes; under the lock, the namespace
+    # read here is still the one to revalidate after the resize. It is
+    # revalidated in the same call, so whenever the zvol grows, also after a
+    # lost reply, the hosts see the new size. A namespace that is not enabled
+    # reads the new size when it is enabled.
+    nvmet_locked(
+        $scfg,
+        sub {
+            my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1);
+            my (undef, $uuid) = _nvmet_owned_identity($inv, $nqn, $dataset);
+            my $namespaces = nvmet_namespaces($cfs, $nqn);
+            my @matches =
+                grep { $namespaces->{$_}->{device_uuid} eq $uuid } keys $namespaces->%*;
+            die "duplicate namespace UUID '$uuid'\n" if @matches > 1;
+            my @resize = (['zfs', 'set', "volsize=${size}k", $dataset]);
+            push @resize, nvmet_write(nvmet_ns_path($nqn, $matches[0]) . '/revalidate_size', 1)
+                if @matches && $namespaces->{ $matches[0] }->{enable} eq '1';
+            nvmet_change(
+                $scfg, [\@resize],
+                op => 'resize zvol',
+                timeout => nvmet_long_timeout(),
+            );
+        },
+    );
+    return;
+}
+
+# Removes the subsystem once it has no volumes and namespaces, then the hosts
+# that no subsystem links any more. Ports are never removed.
+sub _nvmet_delete_target($scfg) {
+    my $nqn = nvmet_nqn($scfg);
+    my $hostnqns = parse_nvme_host_nqns($scfg->{'nvme-host-nqns'}, 1) // [];
+
+    nvmet_locked(
+        $scfg,
+        sub {
+            my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1);
+            my ($units, $candidates) = _nvmet_plan_delete_target($inv, $cfs, $nqn, $hostnqns);
+            my $fresh = $cfs;
+            if ($units->@*) {
+                my $res = nvmet_apply($scfg, $units, op => 'remove NVMe subsystem');
+                (undef, $fresh) = eval { nvmet_read($scfg) };
+                my $read_error = $@;
+                nvmet_rethrow_abandoned($read_error) if $read_error;
+                die "cannot remove NVMe subsystem '$nqn': $res->{err}\n"
+                    if $res->{rc} && (!$fresh || $fresh->{subsystems}->{$nqn});
+                die $read_error if !$fresh;
+            }
+            my $orphans = _nvmet_plan_orphan_hosts($fresh, $candidates);
+            return if !$orphans->@*;
+            my $res = nvmet_apply($scfg, $orphans, op => 'remove orphan NVMe hosts');
+            log_warn("could not remove orphan NVMe host(s): $res->{err}") if $res->{rc};
+        },
+    );
+    return;
+}
+
+# ---------------------------------------------------------------------------
+# Configuration
+# ---------------------------------------------------------------------------
+
+sub verify_nvme_nqn($value, $noerr = undef) {
+
+    if (
+        length($value) > 223
+        || $value !~ $RE_NQN
+    ) {
+        return undef if $noerr;
+        die "value is not a valid NVMe qualified name\n";
+    }
+
+    return $value;
+}
+
+sub parse_nvme_portals($value, $noerr = undef) {
+    my $result = [];
+    my $seen = {};
+
+    for my $entry (split(/,/, $value // '')) {
+        $entry = trim($entry);
+        my ($address, $port, $family);
+        if ($entry =~ $RE_IPV6_PORTAL) {
+            ($address, $port, $family) = ($+{address}, $+{port} // 4420, 'ipv6');
+        } elsif ($entry =~ $RE_IPV4_PORTAL) {
+            ($address, $port, $family) = ($+{address}, $+{port} // 4420, 'ipv4');
+        } else {
+            return undef if $noerr;
+            die "invalid NVMe/TCP portal '$entry'\n";
+        }
+
+        if (
+            !PVE::JSONSchema::pve_verify_ip($address, 1)
+            || ($family eq 'ipv4' && index($address, ':') >= 0)
+            || ($family eq 'ipv6' && index($address, ':') < 0)
+            || $port < 1
+            || $port > 65535
+        ) {
+            return undef if $noerr;
+            die "invalid NVMe/TCP portal '$entry'\n";
+        }
+        $address = nvmet_canonical_address($address) if $family eq 'ipv6';
+        # one listener: compare like _assert_unique_target
+        my $id = join(
+            "\0",
+            nvmet_listener_address({ family => $family, address => $address }),
+            int($port),
+        );
+        if ($seen->{$id}++) {
+            return undef if $noerr;
+            die "duplicate NVMe/TCP portal '$entry'\n";
+        }
+        push $result->@*,
+            {
+                address => $address,
+                port => int($port),
+                family => $family,
+            };
+        if (scalar($result->@*) > $max_paths) {
+            return undef if $noerr;
+            die "at most $max_paths NVMe/TCP portals are supported\n";
+        }
+    }
+
+    if (!$result->@*) {
+        return undef if $noerr;
+        die "at least one NVMe/TCP portal is required\n";
+    }
+
+    return $result;
+}
+
+my sub verify_nvme_portals($value, $noerr = undef) {
+    return undef if !parse_nvme_portals($value, $noerr);
+    return $value;
+}
+
+sub parse_nvme_host_ifaces($value, $noerr = undef) {
+    my $result = [];
+
+    for my $iface (split(/,/, $value // '')) {
+        $iface = trim($iface);
+        # the kernel never names an interface '.' or '..'
+        if (
+            length($iface) < 1
+            || length($iface) > 15
+            || $iface !~ $RE_HOST_IFACE
+            || $iface eq '.'
+            || $iface eq '..'
+        ) {
+            return undef if $noerr;
+            die "invalid NVMe/TCP host interface '$iface'\n";
+        }
+        push $result->@*, $iface;
+    }
+
+    if (!$result->@*) {
+        return undef if $noerr;
+        die "at least one NVMe/TCP host interface is required\n";
+    }
+
+    return $result;
+}
+
+sub parse_nvme_host_nqns($value, $noerr = undef) {
+    my $result = [];
+    my $seen = {};
+
+    for my $hostnqn (split(/,/, $value // '')) {
+        $hostnqn = trim($hostnqn);
+        if (!verify_nvme_nqn($hostnqn, 1) || $seen->{$hostnqn}++) {
+            return undef if $noerr;
+            die "invalid or duplicate NVMe host NQN '$hostnqn'\n";
+        }
+        push $result->@*, $hostnqn;
+        if (scalar($result->@*) > $max_hosts) {
+            return undef if $noerr;
+            die "at most $max_hosts NVMe host NQNs are supported\n";
+        }
+    }
+
+    if (!$result->@*) {
+        return undef if $noerr;
+        die "at least one NVMe host NQN is required\n";
+    }
+
+    return $result;
+}
+
+my sub verify_nvme_host_ifaces($value, $noerr = undef) {
+    return undef if !parse_nvme_host_ifaces($value, $noerr);
+    return $value;
+}
+
+my sub verify_nvme_host_nqns($value, $noerr = undef) {
+    return undef if !parse_nvme_host_nqns($value, $noerr);
+    return $value;
+}
+
+sub _configured_portals($scfg) {
+    my $portals = parse_nvme_portals($scfg->{'nvme-portals'});
+    my $ifaces = parse_nvme_host_ifaces($scfg->{'nvme-host-ifaces'});
+    die "nvme-host-ifaces must contain one interface for each nvme-portals entry\n"
+        if scalar($ifaces->@*) != scalar($portals->@*);
+
+    for (my $i = 0; $i < scalar($portals->@*); $i++) {
+        $portals->[$i]->{host_iface} = $ifaces->[$i];
+    }
+    return $portals;
+}
+
+sub _local_iface_exists($iface) {
+    return -d "/sys/class/net/$iface";
+}
+
+sub _validate_local_ifaces($portals) {
+    for my $portal ($portals->@*) {
+        my $iface = $portal->{host_iface};
+        die "NVMe/TCP host interface '$iface' does not exist on this node\n"
+            if !_local_iface_exists($iface);
+    }
+}
+
+PVE::JSONSchema::register_format('pve-storage-nvme-nqn', \&verify_nvme_nqn);
+PVE::JSONSchema::register_format('pve-storage-nvme-portals', \&verify_nvme_portals);
+PVE::JSONSchema::register_format('pve-storage-nvme-host-ifaces', \&verify_nvme_host_ifaces);
+PVE::JSONSchema::register_format('pve-storage-nvme-host-nqns', \&verify_nvme_host_nqns);
+
+sub type($class) {
+    return 'zfsnvme';
+}
+
+sub plugindata($class) {
+    return {
+        content => [{ images => 1 }, { images => 1 }],
+        'sensitive-properties' => { 'dhchap-key' => 1 },
+    };
+}
+
+sub properties($class) {
+    return {
+        subsysnqn => {
+            description => "NVMe subsystem qualified name.",
+            type => 'string',
+            format => 'pve-storage-nvme-nqn',
+        },
+        'nvme-portals' => {
+            description =>
+                "Comma-separated NVMe/TCP target IP addresses, at most 16. Put IPv6 addresses in"
+                . " brackets. The default TCP port is 4420.",
+            type => 'string',
+            format => 'pve-storage-nvme-portals',
+            maxLength => 2048,
+        },
+        'nvme-host-ifaces' => {
+            description =>
+                "Comma-separated local interfaces, in portal order. Interface names must be"
+                . " identical on every cluster node.",
+            type => 'string',
+            format => 'pve-storage-nvme-host-ifaces',
+            maxLength => 512,
+        },
+        'nvme-host-nqns' => {
+            description =>
+                "Comma-separated /etc/nvme/hostnqn values for every cluster node allowed to use"
+                . " this storage, at most 64. Host NQNs can only be added.",
+            type => 'string',
+            format => 'pve-storage-nvme-host-nqns',
+            maxLength => 8192,
+        },
+        'dhchap-key' => {
+            description =>
+                "NVMe DH-HMAC-CHAP key in secret representation format (DHHC-1). Required when the"
+                . " storage is created; it cannot be changed.",
+            type => 'string',
+            maxLength => 256,
+        },
+        'nvme-iopolicy' => {
+            description => "Native NVMe multipath I/O policy.",
+            type => 'string',
+            enum => ['numa', 'round-robin', 'queue-depth'],
+            default => 'round-robin',
+        },
+        'nvme-keep-alive-tmo' => {
+            description => "NVMe keep-alive timeout in seconds.",
+            type => 'integer',
+            minimum => 1,
+            maximum => 120,
+            default => 5,
+        },
+        'nvme-reconnect-delay' => {
+            description => "Delay between NVMe reconnect attempts in seconds.",
+            type => 'integer',
+            minimum => 1,
+            maximum => 120,
+            default => 2,
+        },
+        'nvme-ctrl-loss-tmo' => {
+            description => "Time to keep retrying a lost NVMe controller in seconds.",
+            type => 'integer',
+            minimum => -1,
+            maximum => 86400,
+            default => 600,
+        },
+        'nvme-fast-io-fail-tmo' => {
+            description =>
+                "Optional time before failing I/O on a reconnecting NVMe controller. Unset queues"
+                . " I/O until controller loss timeout.",
+            type => 'integer',
+            minimum => 0,
+            maximum => 86400,
+            optional => 1,
+        },
+        'nvme-nr-io-queues' => {
+            description => "Number of NVMe/TCP I/O queues per controller.",
+            type => 'integer',
+            minimum => 1,
+            maximum => 1024,
+            optional => 1,
+        },
+    };
+}
+
+sub options($class) {
+    return {
+        server => { fixed => 1 },
+        subsysnqn => { fixed => 1 },
+        'nvme-portals' => { fixed => 1 },
+        'nvme-host-ifaces' => {},
+        'nvme-host-nqns' => {},
+        pool => { fixed => 1 },
+        blocksize => { fixed => 1 },
+        sparse => { optional => 1 },
+        'dhchap-key' => { optional => 1 },
+        'nvme-iopolicy' => { optional => 1 },
+        'nvme-keep-alive-tmo' => { optional => 1 },
+        'nvme-reconnect-delay' => { optional => 1 },
+        'nvme-ctrl-loss-tmo' => { optional => 1 },
+        'nvme-fast-io-fail-tmo' => { optional => 1 },
+        'nvme-nr-io-queues' => { optional => 1 },
+        nodes => { optional => 1 },
+        disable => { optional => 1 },
+        content => { optional => 1 },
+        bwlimit => { optional => 1 },
+    };
+}
+
+sub _validate_fail_fast_timeout($config, $default_ctrl_loss_tmo = undef) {
+    my $fast = $config->{'nvme-fast-io-fail-tmo'};
+    return if !defined($fast);
+
+    my $ctrl = $config->{'nvme-ctrl-loss-tmo'};
+    $ctrl = $default_ctrl_loss_tmo if !defined($ctrl);
+    return if !defined($ctrl) || $ctrl < 0;
+
+    die "nvme-fast-io-fail-tmo must not exceed nvme-ctrl-loss-tmo\n"
+        if $fast > $ctrl;
+}
+
+# Defaults are not injected here: this also runs when storage.cfg is parsed.
+# Every use site applies the schema default itself.
+sub check_config($class, $section_id, $config, $create, $skip_schema_check = undef) {
+    nvmet_pool($config) if defined($config->{pool});
+    verify_nvme_nqn($config->{subsysnqn}) if defined($config->{subsysnqn});
+    parse_nvme_portals($config->{'nvme-portals'}) if defined($config->{'nvme-portals'});
+    if (defined($config->{'nvme-host-ifaces'}) && defined($config->{'nvme-portals'})) {
+        _configured_portals($config);
+    } elsif (defined($config->{'nvme-host-ifaces'})) {
+        parse_nvme_host_ifaces($config->{'nvme-host-ifaces'});
+    }
+    parse_nvme_host_nqns($config->{'nvme-host-nqns'})
+        if defined($config->{'nvme-host-nqns'});
+    _validate_fail_fast_timeout($config, $create ? 600 : undef);
+    return $class->SUPER::check_config($section_id, $config, $create, $skip_schema_check);
+}
+
+# ---------------------------------------------------------------------------
+# ZFS helpers (same formats as the other ZFS plugins)
+# ---------------------------------------------------------------------------
+
+my sub zfs_request($scfg, $timeout, $method, @args) {
+    $timeout = PVE::RPCEnvironment->is_worker() ? 60 * 60 : 10 if !$timeout;
+    my $change = $method ne 'get' && $method ne 'list';
+    my $res = nvmet_exec(
+        $scfg, [['zfs', $method, @args]],
+        op => "zfs $method",
+        timeout => $timeout,
+        change => $change,
+    );
+    die "zfs error on '" . nvmet_server($scfg) . "': $res->{err}\n" if $res->{rc};
+    return nvmet_output($res);
+}
+
+# The direct children of the pool with PVE volume names. The plugin creates
+# and owns zvols only; a filesystem child holds its name.
+my sub zfs_parse_zvol_list($text, $pool) {
+    my $list = [];
+    for my $line (split /\n/, $text // '') {
+        my ($dataset, $size, $origin, $type) = split(/\s+/, $line);
+        next if !defined($type) || ($type ne 'volume' && $type ne 'filesystem');
+        my @parts = split /\//, $dataset;
+        next if @parts < 2;
+        my $name = pop @parts;
+        next if join('/', @parts) ne $pool;
+        next if $name !~ $RE_ZVOL_OWNER;
+        push $list->@*,
+            {
+                name => $name,
+                owner => $+{owner},
+                type => $type,
+                ($type eq 'volume' ? (size => $size + 0) : ()),
+                ($origin ne '-' ? (origin => $origin) : ()),
+            };
+    }
+    return $list;
+}
+
+# The pool's direct children with PVE volume names, so that a name is never
+# handed out twice; with $owned_only, the zvols owned by this storage's NVMe
+# subsystem.
+my sub zfs_list_zvol($scfg, $owned_only) {
+    my $pool = nvmet_pool($scfg);
+    my $text = zfs_request(
+        $scfg,
+        10,
+        'list',
+        '-o',
+        'name,volsize,origin,type',
+        '-t',
+        'volume,filesystem',
+        '-d1',
+        '-Hp',
+        $pool,
+    );
+    my $list = {};
+    my $prefix = "$pool/";
+    for my $zvol (zfs_parse_zvol_list($text, $pool)->@*) {
+        next if $owned_only && $zvol->{type} ne 'volume';
+        my $parent = $zvol->{origin};
+        # an origin in the pool is named relative to it, like the ZFS plugins do
+        $parent = substr($parent, length($prefix)) if $parent && index($parent, $prefix) == 0;
+        $list->{ $zvol->{name} } = {
+            name => $zvol->{name},
+            size => $zvol->{size},
+            parent => $parent,
+            format => 'raw',
+            vmid => $zvol->{owner},
+        };
+    }
+    return $list if !$owned_only;
+
+    my $properties = zfs_request(
+        $scfg,
+        10,
+        'get',
+        '-H',
+        '-d',
+        '1',
+        '-o',
+        'name,value,source',
+        'proxmox:nvme-subsys',
+        $pool,
+    );
+    my $owned = {};
+    for my $line (split(/\n/, $properties)) {
+        my ($dataset, $nqn, $source) = split(/\t/, $line, 3);
+        next if !defined($source) || ($source ne 'local' && $source ne 'received');
+        next if !defined($nqn) || $nqn ne $scfg->{subsysnqn};
+        next if index($dataset, $prefix) != 0;
+        $owned->{ substr($dataset, length($prefix)) } = 1;
+    }
+    for my $name (keys $list->%*) {
+        delete $list->{$name} if !$owned->{$name};
+    }
+    return $list;
+}
+
+my sub zfs_get_properties($scfg, $properties, $dataset, $timeout = undef) {
+    my $text = zfs_request($scfg, $timeout, 'get', '-o', 'value', '-Hp', $properties, $dataset);
+    my @values = split /\n/, $text;
+    return wantarray ? @values : $values[0];
+}
+
+my sub zfs_get_sorted_snapshot_list($scfg, $name, $sort_params) {
+    return [
+        map { s/^.*\@//r } split /\n/,
+        zfs_request(
+            $scfg,
+            undef,
+            'list',
+            '-H',
+            '-r',
+            '-t',
+            'snapshot',
+            '-o',
+            'name',
+            $sort_params->@*,
+            nvmet_dataset($scfg, $name),
+        ),
+    ];
+}
+
+# ---------------------------------------------------------------------------
+# Storage API: volumes
+# ---------------------------------------------------------------------------
+
+sub parse_volname($class, $volname) {
+    if ($volname =~ $RE_VOLNAME) {
+        my $format = ($+{type} eq 'subvol' || $+{type} eq 'basevol') ? 'subvol' : 'raw';
+        my $is_base = $+{type} eq 'base' || $+{type} eq 'basevol';
+        return ('images', $+{name}, $+{vmid}, $+{base}, $+{base_vmid}, $is_base, $format);
+    }
+    die "unable to parse zfs volume name '$volname'\n";
+}
+
+sub list_images($class, $storeid, $scfg, $vmid = undef, $vollist = undef, $cache = undef) {
+    my $res = [];
+    for my $info (values zfs_list_zvol($scfg, 1)->%*) {
+        my $volname =
+            $info->{parent} && $info->{parent} =~ $RE_BASE_SNAPSHOT
+            ? "$storeid:$+{base}/$info->{name}"
+            : "$storeid:$info->{name}";
+        next
+            if $vollist
+            ? !grep { $_ eq $volname } $vollist->@*
+            : defined($vmid) && $info->{vmid} ne $vmid;
+        $info->{volid} = $volname;
+        push $res->@*, $info;
+    }
+    return $res;
+}
+
+# Names are allocated against every direct child of the pool, owned or not, so
+# a leftover zvol that list_images hides is never handed out again.
+sub find_free_diskname($class, $storeid, $scfg, $vmid, $fmt = undef, $add_fmt_suffix = undef) {
+    my $disk_list = [map { "$storeid:$_" } sort keys zfs_list_zvol($scfg, 0)->%*];
+    return PVE::Storage::Plugin::get_next_vm_diskname(
+        $disk_list, $storeid, $vmid, $fmt, $scfg, $add_fmt_suffix,
+    );
+}
+
+sub status($class, $storeid, $scfg, $cache = undef) {
+    my ($available, $used) =
+        eval { zfs_get_properties($scfg, 'available,used', nvmet_pool($scfg)) };
+    if (my $err = $@) {
+        warn "storage '$storeid': $err";
+        return (0, 0, 0, 0);
+    }
+    if (
+        !defined($available)
+        || !defined($used)
+        || $available !~ $RE_UNSIGNED_INTEGER
+        || $used !~ $RE_UNSIGNED_INTEGER
+    ) {
+        warn "unexpected ZFS pool usage for storage '$storeid'\n";
+        return (0, 0, 0, 0);
+    }
+    return ($available + $used, $available, $used, 1);
+}
+
+sub volume_size_info($class, $scfg, $storeid, $volname, $timeout = undef) {
+    my (undef, $name, undef, $parent, undef, undef, $format) = $class->parse_volname($volname);
+    die "volume_size_info requires a ZFS volume\n" if $format ne 'raw';
+    my ($size, $used) = zfs_get_properties(
+        $scfg, 'volsize,usedbydataset', nvmet_dataset($scfg, $name), $timeout,
+    );
+    die "Could not get zfs volume size\n" if !defined($size) || $size !~ $RE_UNSIGNED_INTEGER;
+    $used = defined($used) && $used =~ $RE_UNSIGNED_INTEGER ? $used + 0 : 0;
+    return wantarray ? ($size + 0, 'raw', $used, $parent) : $size + 0;
+}
+
+sub volume_snapshot($class, $scfg, $storeid, $volname, $snap) {
+    my (undef, $name, undef, undef, undef, undef, $format) = $class->parse_volname($volname);
+    die "volume_snapshot requires a ZFS volume\n" if $format ne 'raw';
+    my $snapshot = nvmet_snapshot($scfg, $name, $snap);
+    nvmet_locked($scfg, sub { zfs_request($scfg, undef, 'snapshot', $snapshot) });
+    return;
+}
+
+sub volume_snapshot_delete($class, $scfg, $storeid, $volname, $snap, $running = undef) {
+    my $name = ($class->parse_volname($volname))[1];
+    my $snapshot = nvmet_snapshot($scfg, $name, $snap);
+    nvmet_locked($scfg, sub { zfs_request($scfg, undef, 'destroy', $snapshot) });
+    return;
+}
+
+sub volume_rollback_is_possible($class, $scfg, $storeid, $volname, $snap, $blockers = undef) {
+    my $name = ($class->parse_volname($volname))[1];
+    nvmet_snapshot($scfg, $name, $snap); # validation only
+    my $found;
+    $blockers //= [];
+    for my $snapshot (zfs_get_sorted_snapshot_list($scfg, $name, ['-s', 'creation'])->@*) {
+        $found = 1 if $snapshot eq $snap;
+        push $blockers->@*, $snapshot if $found && $snapshot ne $snap;
+    }
+    die "can't rollback, snapshot '$snap' does not exist on '${storeid}:${volname}'\n" if !$found;
+    die "can't rollback, '$snap' is not most recent snapshot on '${storeid}:${volname}'\n"
+        if $blockers->@*;
+    return 1;
+}
+
+sub volume_snapshot_rollback($class, $scfg, $storeid, $volname, $snap) {
+    my (undef, $name, undef, undef, undef, undef, $format) = $class->parse_volname($volname);
+    die "snapshot rollback requires a ZFS volume\n" if $format ne 'raw';
+    nvmet_rollback_volume($scfg, $name, $snap);
+}
+
+sub volume_snapshot_info($class, $scfg, $storeid, $volname) {
+    my $name = ($class->parse_volname($volname))[1];
+    my $info = {};
+    my $text = zfs_request(
+        $scfg,
+        undef,
+        'list',
+        '-Hp',
+        '-r',
+        '-t',
+        'snapshot',
+        '-o',
+        'name,guid,creation',
+        nvmet_dataset($scfg, $name),
+    );
+    for my $line (split /\n/, $text) {
+        my ($snapshot, $guid, $creation) = split /\s+/, $line;
+        $snapshot =~ s/^.*\@//;
+        $info->{$snapshot} = { id => $guid, timestamp => $creation };
+    }
+    return $info;
+}
+
+sub free_image($class, $storeid, $scfg, $volname, $is_base = undef, $format = undef) {
+    my (undef, $name, undef, undef, undef, undef, $volformat) = $class->parse_volname($volname);
+    die "free_image requires a ZFS volume\n" if $volformat ne 'raw';
+    nvmet_destroy_volume($scfg, $name);
+    return undef;
+}
+
+sub create_base($class, $storeid, $scfg, $volname) {
+    my (undef, $name, undef, $basename, undef, $is_base, $format) = $class->parse_volname($volname);
+    die "create_base not possible with base image\n" if $is_base;
+    die "create_base requires a ZFS volume\n" if $format ne 'raw';
+    my $newname = $name =~ s/^vm-/base-/r;
+    nvmet_template_volume($scfg, $name);
+    return $basename ? "$basename/$newname" : $newname;
+}
+
+sub clone_image($class, $scfg, $storeid, $volname, $vmid, $snap = undef) {
+    $snap ||= '__base__';
+    my (undef, $basename, undef, undef, undef, $is_base, $format) = $class->parse_volname($volname);
+    die "clone_image only works on base images\n" if !$is_base;
+    die "clone_image requires a ZFS volume\n" if $format ne 'raw';
+    my $origin = nvmet_snapshot($scfg, $basename, $snap);
+    my $name = $class->find_free_diskname($storeid, $scfg, $vmid, $format);
+    nvmet_create_volume($scfg, $name, origin => $origin);
+    return "$basename/$name";
+}
+
+sub alloc_image($class, $storeid, $scfg, $vmid, $fmt, $name, $size) {
+    die "unsupported format '$fmt'" if $fmt ne 'raw';
+    die "illegal name '$name' - should be 'vm-$vmid-*'\n"
+        if $name && index($name, "vm-$vmid-") != 0;
+    my $volname = $name || $class->find_free_diskname($storeid, $scfg, $vmid, $fmt);
+    $size += (1024 - $size % 1024) % 1024;
+    nvmet_create_volume($scfg, $volname, size => $size);
+    return $volname;
+}
+
+sub volume_resize($class, $scfg, $storeid, $volname, $size, $running = undef, $snapname = undef) {
+    # QEMU can resize a host_device once the device itself has grown, but the
+    # plugin cannot yet wait for the local NVMe namespace to report the new
+    # size. Until it can, refuse online resize before changing the zvol, so a
+    # running VM never sees a size different from its configuration.
+    die "online resize is not supported for NVMe/TCP block devices; stop the VM first\n"
+        if $running;
+    die "resizing a snapshot is not supported for $class\n" if $snapname;
+    my $new_size = int($size / 1024);
+    $new_size += (1024 - $new_size % 1024) % 1024;
+    my $name = ($class->parse_volname($volname))[1];
+    nvmet_resize_volume($scfg, $name, $new_size);
+    return $new_size;
+}
+
+sub volume_has_feature(
+    $class, $scfg, $feature, $storeid, $volname,
+    $snapname = undef,
+    $running = undef,
+    $opts = undef,
+) {
+    my $features = {
+        snapshot => { current => 1, snap => 1 },
+        clone => { base => 1 },
+        template => { current => 1 },
+        copy => { base => 1, current => 1 },
+    };
+    my $is_base = ($class->parse_volname($volname))[5];
+    my $key = $snapname ? 'snap' : $is_base ? 'base' : 'current';
+    return $features->{$feature}->{$key};
+}
+
+# No stream transfers or renames; the storage has no path, so the base class
+# offers no export or import format either.
+sub volume_export($class, @args) {
+    die "ZFS stream export is not supported for NVMe/TCP storage\n";
+}
+
+sub volume_import($class, @args) {
+    die "ZFS stream import is not supported for NVMe/TCP storage\n";
+}
+
+sub rename_volume($class, @args) {
+    die "renaming volumes is not supported for NVMe/TCP storage\n";
+}
+
+sub rename_snapshot($class, @args) {
+    die "rename_snapshot is not supported for $class";
+}
+
+# ---------------------------------------------------------------------------
+# Secrets and configuration hooks
+# ---------------------------------------------------------------------------
+
+# Each storage needs its own NVMe/TCP listener (address family, address and
+# port), and a disabled storage keeps its listener. nvmet answers a connect to
+# a listener that does not publish the subsystem yet with DNR, and the host
+# then deletes the controller whatever its loss timeout: after a target
+# restart, the storage published second would lose its paths. Storages that
+# reach the same data address must spell the server alike, as it names the
+# target lock.
+sub _assert_unique_target($storeid, $scfg, $cfg = undef) {
+    $cfg //= PVE::Storage::config();
+    my (%listeners, %addresses);
+    for my $portal (parse_nvme_portals($scfg->{'nvme-portals'})->@*) {
+        my $address = join("\0", nvmet_listener_address($portal));
+        $listeners{"$address\0$portal->{port}"} = 1;
+        $addresses{$address} = 1;
+    }
+    my $ids = $cfg->{ids} // {};
+    for my $other_id (sort keys $ids->%*) {
+        next if $other_id eq $storeid;
+        my $other = $ids->{$other_id};
+        next if ($other->{type} // '') ne 'zfsnvme';
+
+        die "NVMe subsystem NQN is already used by storage '$other_id'\n"
+            if ($other->{subsysnqn} // '') eq $scfg->{subsysnqn};
+        my $other_server = $other->{server} // '';
+        # a storage whose portals do not parse cannot be activated
+        for my $portal ((parse_nvme_portals($other->{'nvme-portals'}, 1) // [])->@*) {
+            my $address = join("\0", nvmet_listener_address($portal));
+            die "NVMe/TCP portal '$portal->{address}' port $portal->{port} is already used by"
+                . " storage '$other_id'; give each storage its own address or port\n"
+                if $listeners{"$address\0$portal->{port}"};
+            die "storage '$other_id' reaches target address '$portal->{address}' through server"
+                . " '$other_server'; use the same server value for both storages\n"
+                if $addresses{$address} && $other_server ne $scfg->{server};
+        }
+        next if $other_server ne $scfg->{server};
+        my ($pool, $other_pool) = ($scfg->{pool}, $other->{pool} // '');
+        die "ZFS pool '$pool' on '$scfg->{server}' is already used by storage '$other_id'\n"
+            if $other_pool eq $pool;
+        die "ZFS pool '$pool' on '$scfg->{server}' overlaps the pool of storage '$other_id'\n"
+            if index("$other_pool/", "$pool/") == 0 || index("$pool/", "$other_pool/") == 0;
+    }
+}
+
+# nvmet keeps one DH-HMAC-CHAP key per host NQN, so storages on the same target
+# whose host NQNs overlap must use the same key.
+sub _assert_shared_host_key($storeid, $scfg, $key, $cfg = undef) {
+    $cfg //= PVE::Storage::config();
+    my %hosts = map { $_ => 1 } (parse_nvme_host_nqns($scfg->{'nvme-host-nqns'}, 1) // [])->@*;
+    my $ids = $cfg->{ids} // {};
+    for my $other_id (sort keys $ids->%*) {
+        next if $other_id eq $storeid;
+        my $other = $ids->{$other_id};
+        next if ($other->{type} // '') ne 'zfsnvme';
+        next if ($other->{server} // '') ne ($scfg->{server} // '');
+        my $other_hosts = parse_nvme_host_nqns($other->{'nvme-host-nqns'}, 1) // [];
+        next if !grep { $hosts{$_} } $other_hosts->@*;
+        my $other_key = file_read_firstline(secret_path($other_id));
+        die "storage '$other_id' uses a different DH-HMAC-CHAP key for the same NVMe host NQNs"
+            . " on '$scfg->{server}'\n"
+            if defined($other_key) && $other_key ne $key;
+    }
+}
+
+# DHHC-1:<hash>:<base64 of key and its little-endian CRC-32>: as generated by
+# `nvme gen-dhchap-key`. Messages never contain the key.
+sub _validate_secret($key) {
+    die "missing NVMe DH-HMAC-CHAP key\n" if !defined($key) || $key eq '';
+    my $invalid = "invalid NVMe DH-HMAC-CHAP key representation\n";
+    die $invalid if $key !~ $RE_DHCHAP_KEY;
+    my ($hash, $encoded) = @+{qw(hash secret)};
+    die $invalid if length($encoded) % 4;
+    my $decoded = decode_base64($encoded);
+    my $length = length($decoded) - 4;
+    my %lengths = ('00' => [32, 48, 64], '01' => [32], '02' => [48], '03' => [64]);
+    die $invalid if !grep { $_ == $length } $lengths{$hash}->@*;
+    die $invalid if pack('V', crc32(substr($decoded, 0, $length))) ne substr($decoded, $length);
+    return $key;
+}
+
+my sub set_secret($storeid, $key) {
+    _validate_secret($key);
+    make_path($secret_dir, { mode => 0700 });
+    file_set_contents(secret_path($storeid), "$key\n", 0600);
+}
+
+my sub get_secret($storeid) {
+    my $key = file_read_firstline(secret_path($storeid));
+    return _validate_secret($key);
+}
+
+# Removes the key file of a storage (a test seam).
+sub _unlink_file($path) {
+    return unlink($path);
+}
+
+my sub delete_secret($storeid) {
+    _unlink_file(secret_path($storeid));
+}
+
+sub on_add_hook($class, $storeid, $scfg, %sensitive) {
+    _configured_portals($scfg);
+    parse_nvme_host_nqns($scfg->{'nvme-host-nqns'});
+    _assert_unique_target($storeid, $scfg);
+    my $key = _validate_secret($sensitive{'dhchap-key'});
+    _assert_shared_host_key($storeid, $scfg, $key);
+    set_secret($storeid, $key);
+    return;
+}
+
+sub on_update_hook_full($class, $storeid, $scfg, $update, $delete = undef, $sensitive = undef) {
+    $sensitive //= {};
+    my %prospective = ($scfg->%*, $update->%*);
+    delete @prospective{ $delete->@* } if $delete;
+    verify_nvme_nqn($prospective{subsysnqn});
+    _configured_portals(\%prospective);
+    my $new_hostnqns = parse_nvme_host_nqns($prospective{'nvme-host-nqns'});
+    my %new_hosts = map { $_ => 1 } $new_hostnqns->@*;
+    for my $old_hostnqn (parse_nvme_host_nqns($scfg->{'nvme-host-nqns'})->@*) {
+        die "removing NVMe host NQN '$old_hostnqn' is not supported; move every volume off"
+            . " the storage and recreate it with a new subsystem NQN and a new DH-HMAC-CHAP"
+            . " key\n"
+            if !$new_hosts{$old_hostnqn};
+    }
+    _validate_fail_fast_timeout(\%prospective, 600);
+    _assert_unique_target($storeid, \%prospective);
+
+    my $old_key = file_read_firstline(secret_path($storeid));
+    my $key = exists($sensitive->{'dhchap-key'}) ? $sensitive->{'dhchap-key'} : $old_key;
+    _validate_secret($key);
+
+    # nvmet stores authentication on the global Host NQN object, which can be
+    # shared by multiple subsystems. Replacing an active key in place can make
+    # unrelated storages unrecoverable on their next reconnect. Until PVE can
+    # coordinate a rolling rotation on every node, fail before mutating the
+    # cluster-wide secret.
+    die "NVMe DH-HMAC-CHAP key rotation is not supported; create a new storage with a new"
+        . " subsystem NQN\n"
+        if defined($old_key) && $key ne $old_key;
+    _assert_shared_host_key($storeid, \%prospective, $key);
+
+    set_secret($storeid, $key) if exists($sensitive->{'dhchap-key'});
+    return;
+}
+
+# ---------------------------------------------------------------------------
+# Local NVMe host side
+# ---------------------------------------------------------------------------
+
+my %slow_path_last; # storeid => time of the last slow-path activation
+
+# Probes of the local host (test seams).
+sub _block_device($path) {
+    return -b $path;
+}
+
+sub _link_target($path) {
+    return readlink($path);
+}
+
+sub _fabrics_device() {
+    return '/dev/nvme-fabrics';
+}
+
+# Runs $code in a child and gives up on it after $timeout seconds. Returns the
+# child's result; the child's errors pass through as exceptions, a timeout
+# becomes one. The bound is not hard: as with PVE::Tools::run_command, a
+# child stuck in the kernel (a write or delete in D state) is only reaped
+# once the kernel returns, and a timeout does not mean that the operation did
+# not happen.
+#
+# A stopped task (TERM, also HUP, INT and QUIT) stays pending in the parent
+# while the child runs, and the child inherits the mask: a connect that ends
+# before PVE escalates the stop to KILL, 5 seconds after the TERM, creates
+# the controller and restricts its secret attributes without interruption,
+# and run_fork_with_timeout, whose own handlers would consume the signal when
+# the child delivers a result, never sees it. Once the child is reaped and the
+# mask is restored, the handler of the worker dies with the task marker, also
+# over a result or an error of the child. ALRM, the timeout, is not blocked.
+# The KILL ends the parent and the child together; a controller that the
+# child had created but not restricted yet keeps its secret attributes
+# world-readable until the next activation restricts them.
+#
+# run_fork_with_timeout is called in list context on purpose: only there does
+# run_with_timeout report a timeout as its second return value; in scalar
+# context the timeout would be a warning and an undef result.
+my sub run_bounded($timeout, $what, $code) {
+    my $previous = POSIX::SigSet->new();
+    sigprocmask(SIG_BLOCK, POSIX::SigSet->new(SIGHUP, SIGINT, SIGQUIT, SIGTERM), $previous)
+        or die "cannot block signals: $!\n";
+    my $outer_warn = $SIG{__WARN__};
+    my ($res, $timed_out) = eval {
+        # Only a child that leaves no result makes run_fork_with_timeout warn
+        # in the parent; that case is an exception below. The child keeps the
+        # handler of the caller, so its own warnings reach the task log.
+        local $SIG{__WARN__} = sub($message) { };
+        PVE::Tools::run_fork_with_timeout(
+            $timeout,
+            sub {
+                local $SIG{__WARN__} = $outer_warn;
+                return $code->();
+            },
+        );
+    };
+    my $error = $@;
+    sigprocmask(SIG_SETMASK, $previous) or die "cannot restore signals: $!\n";
+    if ($error) {
+        rethrow_task_interrupt($error);
+        die $error;
+    }
+    die "$what did not complete within $timeout seconds\n" if $timed_out;
+    die "$what left no result\n" if !defined($res);
+    return $res;
+}
+
+# The connect options of one controller, as the fabrics device takes them,
+# and their names. The values are validated configuration and host identities
+# checked by the caller; the kernel reads each one up to the next ','.
+sub _fabrics_options($scfg, $portal, $hostnqn, $hostid, $key) {
+    my @options = (
+        [transport => 'tcp'],
+        [traddr => $portal->{address}],
+        [trsvcid => $portal->{port}],
+        [host_iface => $portal->{host_iface}],
+        [nqn => $scfg->{subsysnqn}],
+        [hostnqn => $hostnqn],
+        [hostid => $hostid],
+        [dhchap_secret => $key],
+        [keep_alive_tmo => $scfg->{'nvme-keep-alive-tmo'} // 5],
+        [reconnect_delay => $scfg->{'nvme-reconnect-delay'} // 2],
+        [ctrl_loss_tmo => $scfg->{'nvme-ctrl-loss-tmo'} // 600],
+    );
+    # unlike unset, which leaves fast I/O failure off, 0 is a real timeout
+    push @options, [fast_io_fail_tmo => $scfg->{'nvme-fast-io-fail-tmo'}]
+        if defined($scfg->{'nvme-fast-io-fail-tmo'});
+    push @options, [nr_io_queues => $scfg->{'nvme-nr-io-queues'}]
+        if defined($scfg->{'nvme-nr-io-queues'});
+    for my $option (@options) {
+        die "internal error: invalid NVMe connect option '$option->[0]'\n"
+            if !defined($option->[1]) || $option->[1] !~ $RE_FABRICS_VALUE;
+    }
+    return (join(',', map { "$_->[0]=$_->[1]" } @options), [map { $_->[0] } @options]);
+}
+
+# The child of a connect (a test seam): creates exactly one controller with
+# one write of $options to the fabrics device and returns its instance. The
+# write is never repeated, because the kernel may have created the controller
+# even when it reports an error. Errors carry the errno text of the failed
+# operation and never the options, which hold the key.
+sub _fabrics_connect($options, $names) {
+    my $device = _fabrics_device();
+    my $probe;
+    if (!sysopen($probe, $device, O_RDONLY | O_NOFOLLOW)) {
+        die "'$device' does not exist; load the nvme-tcp kernel module\n" if $! == ENOENT;
+        die "cannot open '$device': $!\n";
+    }
+    my @stat = stat($probe);
+    die "'$device' is not a character device\n" if !@stat || !S_ISCHR($stat[2]);
+    # without a controller, a read lists the supported options as patterns
+    my $supported = '';
+    sysread($probe, $supported, 4096) // die "cannot read '$device': $!\n";
+    close($probe);
+    my %supported = map { s/=.*//r => 1 } split(/,/, $supported =~ s/\n\z//r);
+    for my $name ($names->@*) {
+        die "the running kernel does not support the NVMe connect option '$name'\n"
+            if !$supported{$name};
+    }
+
+    # Not PVE::SysFSTools::file_write: the kernel returns the instance of the
+    # new controller on a read of the file descriptor that wrote the options.
+    sysopen(my $fh, $device, O_RDWR | O_NOFOLLOW) or die "cannot open '$device': $!\n";
+    my $written = _fabrics_write($fh, $options);
+    die "connect failed: $!\n" if !defined($written);
+    die "connect failed: short write\n" if $written != length($options);
+    my $result = '';
+    sysread($fh, $result, 128) // die "cannot read the connect result: $!\n";
+    die "unexpected connect result\n" if $result !~ $RE_FABRICS_RESULT;
+    my $instance = int($+{instance});
+    # The kernel creates the secret attributes of a controller world-readable;
+    # the caller keeps a stopped task from interrupting before this, short of
+    # a KILL.
+    _restrict_attr("/sys/class/nvme/nvme$instance/$_") for qw(dhchap_secret dhchap_ctrl_secret);
+    close($fh);
+    return $instance;
+}
+
+# The write that creates a controller (a test seam).
+sub _fabrics_write($fh, $options) {
+    return syswrite($fh, $options);
+}
+
+# Writes a sysfs attribute and returns the errno text of a failure, or undef.
+# PVE::SysFSTools::file_write returns undef without a warning when the open
+# fails, leaving its errno in $!, and 0 after warning "error writing ...:
+# <errno text>" when the write fails; that warning is not passed on. With
+# $missing_ok, an attribute that does not exist, because its controller went
+# away, is not a failure.
+my sub write_attribute($path, $value, $missing_ok = 0) {
+    my $warning;
+    my $written = do {
+        local $SIG{__WARN__} = sub($message) { $warning = $message };
+        PVE::SysFSTools::file_write($path, $value);
+    };
+    return if $written;
+    return if !defined($warning) && $missing_ok && $! == ENOENT;
+    return "$!" if !defined($warning);
+    return $warning =~ s/\A.*: //sr =~ s/\n\z//r;
+}
+
+# Deletes controllers (the body of a child, a test seam): the deletion can
+# block in the kernel. The error names the controller and the errno text; a
+# controller that is already gone counts as deleted.
+sub _delete_controllers($controllers) {
+    for my $controller ($controllers->@*) {
+        my $path = "/sys/class/nvme/$controller->{name}/delete_controller";
+        my $error = write_attribute($path, '1', 1) // next;
+        die "cannot delete NVMe controller '$controller->{name}': $error\n";
+    }
+    return scalar($controllers->@*);
+}
+
+# Every local controller of the subsystem.
+my sub nvme_controllers($nqn) {
+    my $controllers = [];
+    PVE::File::dir_glob_foreach(
+        '/sys/class/nvme',
+        $RE_NVME_CONTROLLER,
+        sub($entry) {
+            my $base = "/sys/class/nvme/$entry";
+            my $subsys = file_read_firstline("$base/subsysnqn");
+            return if !defined($subsys) || $subsys ne $nqn;
+            my $address = file_read_firstline("$base/address") // '';
+            my $traddr = $address =~ $RE_TRADDR ? $+{value} : undef;
+            my $trsvcid = $address =~ $RE_TRSVCID ? $+{value} : undef;
+            my $host_iface = $address =~ $RE_HOST_IFACE_ADDRESS ? $+{value} : undef;
+            push $controllers->@*,
+                {
+                    name => $entry,
+                    state => file_read_firstline("$base/state") // 'unknown',
+                    traddr => nvmet_canonical_address($traddr),
+                    trsvcid => $trsvcid,
+                    host_iface => $host_iface,
+                };
+        },
+    );
+    return $controllers;
+}
+
+# Controllers by "address:port", with canonical IPv6 addresses. A portal has
+# more than one controller while its path moves to another interface.
+my sub controller_states($nqn) {
+    my $states = {};
+    for my $controller (nvme_controllers($nqn)->@*) {
+        next if !defined($controller->{traddr}) || !defined($controller->{trsvcid});
+        push $states->{"$controller->{traddr}:$controller->{trsvcid}"}->@*, $controller;
+    }
+    return $states;
+}
+
+# The multipath namespace devices of the subsystem and their partitions. One
+# dir_glob_foreach per directory level, because PVE::File::dir_glob_regex
+# returns only the first match.
+my sub namespace_devices($nqn) {
+    my $devices = {};
+
+    PVE::File::dir_glob_foreach(
+        '/sys/class/nvme-subsystem',
+        $RE_NVME_SUBSYSTEM,
+        sub($entry) {
+            my $base = "/sys/class/nvme-subsystem/$entry";
+            my $subsys = file_read_firstline("$base/subsysnqn");
+            return if !defined($subsys) || $subsys ne $nqn;
+
+            PVE::File::dir_glob_foreach(
+                $base,
+                $RE_NVME_NAMESPACE,
+                sub($device) {
+                    $devices->{"/dev/$device"} = $device if _block_device("/dev/$device");
+                    PVE::File::dir_glob_foreach(
+                        "/sys/class/block/$device",
+                        $RE_NVME_PARTITION,
+                        sub($partition) {
+                            $devices->{"/dev/$partition"} = $partition
+                                if _block_device("/dev/$partition");
+                        },
+                    );
+                },
+            );
+        },
+    );
+    return $devices;
+}
+
+sub _namespace_openers($nqn) {
+    my $devices = namespace_devices($nqn);
+    return [] if !$devices->%*;
+
+    my %openers;
+    PVE::File::dir_glob_foreach(
+        '/proc',
+        '[0-9]+',
+        sub($pid) {
+            PVE::File::dir_glob_foreach(
+                "/proc/$pid/fd",
+                '[0-9]+',
+                sub($fd) {
+                    my $target = _link_target("/proc/$pid/fd/$fd");
+                    return if !defined($target) || !exists($devices->{$target});
+                    my $comm = eval { file_read_firstline("/proc/$pid/comm") } // 'unknown';
+                    $openers{"$pid:$target"} = "$comm (PID $pid, $target)";
+                },
+            );
+        },
+    );
+
+    # Kernel consumers such as device-mapper do not necessarily keep a userspace
+    # file descriptor open, but expose their dependency in the holders directory.
+    for my $path (keys $devices->%*) {
+        my $device = $devices->{$path};
+        my $holders = "/sys/class/block/$device/holders";
+        PVE::File::dir_glob_foreach(
+            $holders,
+            '[^\\.].*',
+            sub($holder) {
+                $openers{"holder:$device:$holder"} = "$path held by $holder";
+            },
+        );
+    }
+
+    return [sort values %openers];
+}
+
+my sub portal_reachable($portal) {
+    return PVE::Network::tcp_ping($portal->{address}, $portal->{port}, 2) // 0;
+}
+
+# The number of portals with a live controller on their configured interface,
+# or, with $any_iface, on any interface: such a path carries I/O, even while
+# it waits to be moved to its configured interface.
+sub _live_portal_count($states, $portals, $any_iface = 0) {
+    my $live = 0;
+    for my $portal ($portals->@*) {
+        my $id = "$portal->{address}:$portal->{port}";
+        $live++ if grep {
+            $_->{state} eq 'live'
+                && ($any_iface || ($_->{host_iface} // '') eq $portal->{host_iface})
+        } ($states->{$id} // [])->@*;
+    }
+    return $live;
+}
+
+# Creates the controller of a portal in a child, giving up after 10 seconds.
+sub _connect_portal($scfg, $portal, $hostnqn, $hostid, $key) {
+    my ($options, $names) = _fabrics_options($scfg, $portal, $hostnqn, $hostid, $key);
+    return run_bounded(10, 'NVMe/TCP connect', sub { return _fabrics_connect($options, $names) });
+}
+
+my sub set_iopolicy($nqn, $policy) {
+    PVE::File::dir_glob_foreach(
+        '/sys/class/nvme-subsystem',
+        $RE_NVME_SUBSYSTEM,
+        sub($entry) {
+            my $base = "/sys/class/nvme-subsystem/$entry";
+            my $subsys = file_read_firstline("$base/subsysnqn");
+            return if !defined($subsys) || $subsys ne $nqn;
+            my $error = write_attribute("$base/iopolicy", "$policy\n");
+            die "unable to set NVMe multipath policy: $error\n" if defined($error);
+        },
+    );
+}
+
+# Applies the reconnect timeouts to connected controllers, which otherwise
+# keep the values of their connect. sysfs shows "off" for -1, and shows the
+# loss timeout rounded up to a multiple of the reconnect delay.
+my sub set_tunables($scfg) {
+    my $delay = $scfg->{'nvme-reconnect-delay'} // 2;
+    my $loss = $scfg->{'nvme-ctrl-loss-tmo'} // 600;
+    my $fast = $scfg->{'nvme-fast-io-fail-tmo'};
+    my $shown_loss = $loss < 0 ? 'off' : int(($loss + $delay - 1) / $delay) * $delay;
+    my $shown_fast = defined($fast) && $fast >= 0 ? $fast : 'off';
+
+    for my $controller (nvme_controllers($scfg->{subsysnqn})->@*) {
+        my $base = "/sys/class/nvme/$controller->{name}";
+        my $current_delay = file_read_firstline("$base/reconnect_delay") // next;
+        my $write = sub($attr, $value) {
+            my $error = write_attribute("$base/$attr", "$value\n") // return;
+            log_warn("cannot set $attr of NVMe controller '$controller->{name}': $error");
+        };
+        my $delay_changed = $current_delay ne "$delay";
+        $write->('reconnect_delay', $delay) if $delay_changed;
+        # the loss timeout is stored as a number of reconnects of the delay
+        $write->('ctrl_loss_tmo', $loss < 0 ? -1 : $loss)
+            if $delay_changed
+            || (file_read_firstline("$base/ctrl_loss_tmo") // '') ne "$shown_loss";
+        $write->('fast_io_fail_tmo', $shown_fast eq 'off' ? -1 : $fast)
+            if (file_read_firstline("$base/fast_io_fail_tmo") // '') ne "$shown_fast";
+    }
+}
+
+# Attempts to make a controller attribute readable by root only (a test seam).
+sub _restrict_attr($path) {
+    my @stat = stat($path);
+    if (!@stat) {
+        log_warn("cannot stat NVMe controller attribute '$path': $!") if $! != ENOENT;
+        return;
+    }
+    return if !($stat[2] & 077);
+    chmod(0600, $path) or log_warn("cannot restrict permissions of '$path': $!");
+}
+
+# The kernel creates the DH-HMAC-CHAP secret attributes world-readable.
+my sub restrict_secret_attrs($nqn) {
+    for my $controller (nvme_controllers($nqn)->@*) {
+        _restrict_attr("/sys/class/nvme/$controller->{name}/$_")
+            for qw(dhchap_secret dhchap_ctrl_secret);
+    }
+}
+
+# Best effort: a controller that went away since it was listed needs no scan,
+# and one that cannot be rescanned is a warning.
+my sub rescan_namespaces($nqn) {
+    for my $controller (nvme_controllers($nqn)->@*) {
+        next if $controller->{state} ne 'live';
+        my $path = "/sys/class/nvme/$controller->{name}/rescan_controller";
+        my $error = write_attribute($path, "1\n", 1) // next;
+        log_warn("cannot rescan NVMe controller '$controller->{name}': $error");
+    }
+}
+
+# Deletes a controller that is dead or that moved to another interface. The
+# activation reconnects or keeps the path either way, so this only warns.
+my sub disconnect_controller($controller) {
+    my $name = $controller->{name};
+    eval {
+        run_bounded(
+            5,
+            "deleting NVMe controller '$name'",
+            sub { return _delete_controllers([$controller]) },
+        );
+    };
+    if (my $error = $@) {
+        rethrow_task_interrupt($error);
+        log_warn($error);
+    }
+}
+
+# Waits up to 10 seconds for a live controller of the portal on its
+# configured interface.
+my sub wait_for_live_path($nqn, $portal) {
+    for (my $attempt = 0; $attempt < 40; $attempt++) {
+        return 1 if _live_portal_count(controller_states($nqn), [$portal]);
+        _sleep(0.25);
+    }
+    return 0;
+}
+
+# Connects every missing or dead path on its configured interface. A path
+# that is connected on another interface moves make-before-break: its old
+# controller is only disconnected once the new one is live, so changing an
+# interface mapping never takes away a path that carries I/O.
+my sub connect_portals($scfg, $portals, $hostnqn, $hostid, $key) {
+    my $nqn = $scfg->{subsysnqn};
+    for my $portal ($portals->@*) {
+        my $id = "$portal->{address}:$portal->{port}";
+        my $iface = $portal->{host_iface};
+        my @controllers = (controller_states($nqn)->{$id} // [])->@*;
+        my ($configured) = grep { ($_->{host_iface} // '') eq $iface } @controllers;
+        my @moved = grep { ($_->{host_iface} // '') ne $iface } @controllers;
+
+        if (!$configured || $configured->{state} eq 'dead') {
+            disconnect_controller($configured) if $configured;
+            if (!portal_reachable($portal)) {
+                log_warn("NVMe/TCP portal '$id' is unreachable");
+                next;
+            }
+            eval { _connect_portal($scfg, $portal, $hostnqn, $hostid, $key) };
+            my $connect_error = $@;
+            # A failed or timed-out connect can still have created a controller.
+            eval { restrict_secret_attrs($nqn); 1 };
+            if (my $error = $@) {
+                rethrow_task_interrupt($error);
+                log_warn("cannot restrict NVMe controller attributes for '$nqn'");
+            }
+            rethrow_task_interrupt($connect_error) if $connect_error;
+            if ($connect_error) {
+                log_warn("$id: $connect_error");
+                next;
+            }
+        }
+        next if !@moved;
+        if (!wait_for_live_path($nqn, $portal)) {
+            log_warn("NVMe/TCP portal '$id' is not live on '$iface' yet; keeping its path on"
+                . " another interface");
+            next;
+        }
+        disconnect_controller($_) for @moved;
+    }
+}
+
+sub activate_storage($class, $storeid, $scfg, $cache = undef) {
+    $cache //= {};
+    my $nqn = $scfg->{subsysnqn};
+    # First, whatever the rest of the activation does, also for controllers
+    # connected by hand. Also try after each connect attempt below, including
+    # failures that may have created a controller.
+    restrict_secret_attrs($nqn);
+
+    my $multipath = file_read_firstline('/sys/module/nvme_core/parameters/multipath')
+        // die "the NVMe kernel modules are not loaded; load the nvme-tcp kernel module\n";
+    die "native NVMe multipath is disabled in the running kernel\n" if $multipath ne 'Y';
+
+    my $portals = _configured_portals($scfg);
+    verify_nvme_nqn($nqn);
+    _validate_local_ifaces($portals);
+    _assert_unique_target($storeid, $scfg);
+    my $hostnqn = file_read_firstline('/etc/nvme/hostnqn')
+        // die "missing /etc/nvme/hostnqn; the nvme-cli package generates it\n";
+    verify_nvme_nqn($hostnqn);
+    my $hostnqns = parse_nvme_host_nqns($scfg->{'nvme-host-nqns'});
+    die "local NVMe host NQN '$hostnqn' is missing from nvme-host-nqns\n"
+        if !grep { $_ eq $hostnqn } $hostnqns->@*;
+    my $policy = $scfg->{'nvme-iopolicy'} // 'round-robin';
+    my $force_reconcile = delete($cache->{'zfsnvme-force-reconcile'}->{$storeid}) // 0;
+    my $states = controller_states($nqn);
+    my $healthy = _live_portal_count($states, $portals);
+    my $usable = _live_portal_count($states, $portals, 1);
+    my $last = $slow_path_last{$storeid};
+    my $recent = defined($last) && _now() - $last < $slow_path_backoff;
+
+    # This method is called by the periodic storage status loop. Once every
+    # configured path is live, lifecycle operations keep the target converged,
+    # so no remote call is needed. A degraded storage retries the target at
+    # most once a minute: in between, one with a live path (also one waiting
+    # to move to its configured interface) is usable, and one without fails
+    # at once instead of waiting for a path again. Missing namespaces
+    # explicitly force the slow path from activate_volume().
+    if (!$force_reconcile && ($healthy == scalar($portals->@*) || $recent)) {
+        die "no live NVMe/TCP path for storage '$storeid'\n" if !$usable;
+        die "missing NVMe DH-HMAC-CHAP key\n"
+            if (file_read_firstline(secret_path($storeid)) // '') eq '';
+        set_iopolicy($nqn, $policy);
+        set_tunables($scfg);
+        return 1;
+    }
+
+    $slow_path_last{$storeid} = _now();
+    my $hostid = file_read_firstline('/etc/nvme/hostid')
+        // die "missing /etc/nvme/hostid; the nvme-cli package generates it\n";
+    die "invalid NVMe host ID\n"
+        if $hostid !~ $RE_NVMET_UUID || $hostid =~ $RE_ZERO_UUID;
+    my $key = get_secret($storeid);
+    for my $name (qw(keep-alive-tmo reconnect-delay ctrl-loss-tmo fast-io-fail-tmo nr-io-queues)) {
+        my $value = $scfg->{"nvme-$name"};
+        die "invalid NVMe connection parameter '$name'\n"
+            if defined($value) && $value !~ $RE_CONFIG_INT;
+    }
+    _nvmet_activate_target($storeid, $scfg, $portals, $hostnqns, $key);
+    connect_portals($scfg, $portals, $hostnqn, $hostid, $key);
+
+    # Existing controllers can be in the middle of their kernel reconnect delay
+    # after a target restart. Do not create duplicates, but give that recovery
+    # cycle enough time to complete before declaring the storage unavailable.
+    for (my $attempt = 0; $attempt < 60; $attempt++) {
+        $states = controller_states($nqn);
+        last if _live_portal_count($states, $portals, 1);
+        _sleep(0.25);
+    }
+    die "no live NVMe/TCP path for storage '$storeid'\n"
+        if !_live_portal_count($states, $portals, 1);
+    $healthy = _live_portal_count($states, $portals);
+    log_warn("storage '$storeid' is degraded: $healthy/" . scalar($portals->@*) . " paths live")
+        if $healthy < scalar($portals->@*);
+
+    set_iopolicy($nqn, $policy);
+    set_tunables($scfg);
+    return 1;
+}
+
+sub deactivate_storage($class, $storeid, $scfg, $cache = undef) {
+    my $nqn = $scfg->{subsysnqn};
+    my $openers = _namespace_openers($nqn);
+    die "refusing to disconnect NVMe storage '$storeid': namespace in use by "
+        . join(', ', $openers->@*) . "\n"
+        if $openers->@*;
+
+    # One child deletes every controller, bounded by 15 seconds in all. When it
+    # gives up, the controllers that are left stay connected and the
+    # deactivation fails.
+    my $controllers = nvme_controllers($nqn);
+    run_bounded(
+        15,
+        "disconnecting NVMe subsystem '$nqn'",
+        sub { return _delete_controllers($controllers) },
+    ) if $controllers->@*;
+    return 1;
+}
+
+# Removing a storage definition never destroys data and never depends on the
+# target, like for every other storage type. Target cleanup is best effort
+# and only happens once the storage owns no volume. Other nodes keep their
+# connections until they are disconnected there.
+sub on_delete_hook($class, $storeid, $scfg) {
+    # First: an activation elsewhere that waits for the target lock then
+    # refuses instead of restoring what is removed below.
+    delete_secret($storeid);
+    eval { $class->deactivate_storage($storeid, $scfg) };
+    if (my $err = $@) {
+        rethrow_task_interrupt($err);
+        chomp($err);
+        log_warn("not disconnecting NVMe storage '$storeid': $err");
+    }
+    eval { _nvmet_delete_target($scfg) };
+    if (my $err = $@) {
+        rethrow_task_interrupt($err);
+        chomp($err);
+        log_warn("keeping the NVMe target configuration of storage '$storeid': $err");
+    }
+    return;
+}
+
+sub path($class, $scfg, $volname, $storeid, $snapname = undef) {
+    die "direct access to snapshots not implemented\n" if defined($snapname);
+    my ($vtype, $name, $vmid) = $class->parse_volname($volname);
+    my $uuid = _nvmet_volume_uuid($scfg, $name);
+    my $path = "/dev/disk/by-id/nvme-uuid.$uuid";
+    return ($path, $vmid, $vtype);
+}
+
+sub qemu_blockdev_options(
+    $class, $scfg, $storeid, $volname,
+    $machine_version = undef,
+    $options = undef,
+) {
+    die "direct access to snapshots not implemented\n" if $options->{'snapshot-name'};
+    my ($path) = $class->path($scfg, $volname, $storeid);
+    return { driver => 'host_device', filename => $path };
+}
+
+# Whether the local device behind a namespace link is the namespace (NSID,
+# UUID) of this subsystem, so a UUID claimed by another subsystem is never
+# handed to a guest.
+sub _nvmet_local_namespace_ok($path, $nqn, $nsid, $uuid) {
+    my $link = _link_target($path) // return 0;
+    my $device = (split m{/}, $link)[-1];
+    return 0 if !defined($device) || $device !~ $RE_NVME_NAMESPACE;
+    my $sys = "/sys/block/$device";
+    return 0 if lc(file_read_firstline("$sys/uuid") // '') ne lc($uuid);
+    return 0 if (file_read_firstline("$sys/nsid") // '') ne "$nsid";
+    return 0 if (file_read_firstline("$sys/device/subsysnqn") // '') ne $nqn;
+    return 1;
+}
+
+sub activate_volume(
+    $class, $storeid, $scfg, $volname,
+    $snapname = undef,
+    $cache = undef,
+    $hints = undef,
+) {
+    die "unable to activate snapshot from remote zfs storage\n" if $snapname;
+    my (undef, $name) = $class->parse_volname($volname);
+    my $nqn = nvmet_nqn($scfg);
+    my $dataset = nvmet_dataset($scfg, $name);
+
+    my ($inv, $cfs) = nvmet_read($scfg);
+    my ($nsid, $uuid) = _nvmet_owned_identity($inv, $nqn, $dataset);
+    my $state = _nvmet_export_state($cfs, $nqn, $nsid, $uuid, "/dev/zvol/$dataset");
+    my $path = "/dev/disk/by-id/nvme-uuid.$uuid";
+    my $ok = sub {
+        return _block_device($path) && _nvmet_local_namespace_ok($path, $nqn, $nsid, $uuid);
+    };
+    return 1 if $state eq 'present' && $ok->();
+
+    $cache //= {};
+    $cache->{'zfsnvme-force-reconcile'}->{$storeid} = 1;
+    $class->activate_storage($storeid, $scfg, $cache);
+    for (my $attempt = 0; $attempt < 40 && !_block_device($path); $attempt++) {
+        # a host can miss the change notice; rescanning is harmless
+        rescan_namespaces($nqn) if $attempt % 8 == 0;
+        _sleep(0.25);
+    }
+    die "NVMe namespace for '$volname' did not appear\n" if !_block_device($path);
+    die "NVMe namespace for '$volname' has an unexpected identity\n" if !$ok->();
+    return 1;
+}
+
+sub deactivate_volume(
+    $class, $storeid, $scfg, $volname, $snapname = undef, $cache = undef,
+) {
+    die "unable to deactivate snapshot from remote zfs storage\n" if $snapname;
+    return 1;
+}
+
+1;



  reply	other threads:[~2026-10-06  8:54 UTC|newest]

Thread overview: 7+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-05  0:26 [PATCH storage v3 0/4] add ZFS over NVMe/TCP storage plugin Joaquin Varela
2026-10-05  0:26 ` Joaquin Varela [this message]
2026-10-05  0:26 ` [PATCH storage v3 2/4] test: add zfsnvme plugin tests Joaquin Varela
2026-10-05  0:26 ` [PATCH storage v3 3/4] zfsnvme: fence target commands of abandoned transactions Joaquin Varela
2026-10-05  0:26 ` [PATCH storage v3 4/4] zfsnvme: wait up to 30 seconds for the shared storage lock Joaquin Varela
2026-10-05  0:26 ` [PATCH docs v3] storage: document ZFS over NVMe/TCP Joaquin Varela
2026-10-05  0:26 ` [PATCH manager v3] ui: storage: add ZFS over NVMe/TCP editor Joaquin Varela

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261005002609.571-2-joaquinvarela@neatech.ar \
    --to=joaquinvarela@neatech.ar \
    --cc=pve-devel@lists.proxmox.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.
Service provided by Proxmox Server Solutions GmbH | Privacy | Legal