From: Joaquin Varela <joaquinvarela@neatech.ar>
To: pve-devel@lists.proxmox.com
Subject: [PATCH storage v3 1/4] zfsnvme: add ZFS over NVMe/TCP storage plugin
Date: Sun, 4 Oct 2026 21:26:04 -0300 [thread overview]
Message-ID: <20261005002609.571-2-joaquinvarela@neatech.ar> (raw)
In-Reply-To: <20261005002609.571-1-joaquinvarela@neatech.ar>
Add the shared storage type 'zfsnvme': ZFS zvols on a remote Linux
target, exported with the kernel NVMe target (nvmet, configured through
configfs) over NVMe/TCP, and consumed by the cluster nodes with native
NVMe multipath and DH-HMAC-CHAP authentication.
The ZFS dataset is the source of truth for a namespace: the subsystem
NQN, NSID and UUID of every volume are stored in the ZFS user properties
'proxmox:nvme-subsys', 'proxmox:nvme-nsid' and 'proxmox:nvme-uuid', and
the pool keeps the highest NSID handed out in 'proxmox:nvme-last-nsid'.
The nodes find a namespace by its UUID. Everything else on the target
(nvmet_tcp module, configfs mount, subsystem, namespaces, ports, host
ACLs) is derived state that an activation rebuilds after a target
restart.
All parsing, validation and planning happen locally. The target state
is read with one 'zfs get' and one 'find' over configfs; other reads
are 'zfs get' and 'zfs list', 'test -b' while a new zvol appears, and
'cat /proc/mounts' to decide whether configfs needs to be mounted. The
target runs quoted simple commands over SSH, joined with '&&' where a
step depends on the previous one: an export starts with 'test -b' on
its zvol, and each publish link repeats two read-only checks right
before it links the subsystem to a port. There are no scripts, loops
or variables on the target. As with ZFS over iSCSI, the SSH key is
/etc/pve/priv/zfs/<server>_id_rsa. Target changes are serialized
cluster-wide by a pmxcfs domain lock per target server, and an
operation whose SSH connection broke is reported as unknown and
repaired by the next activation rather than compensated from a read
that could overtake it.
Each storage needs its own NVMe/TCP listener (address family, address
and port) on the target. nvmet answers a connect to a listener that
does not publish the subsystem yet with DNR, and the host then deletes
the controller whatever its loss timeout, so after a target restart the
storage published second on a shared listener would lose its paths.
Adding, updating and activating a storage therefore refuse a portal
whose listener another zfsnvme storage uses, disabled ones included,
and a target address that another storage reaches through a different
'server' value, since that value names the target lock. A data address
thus identifies one target for all zfsnvme storages of the cluster.
The DH-HMAC-CHAP key is a sensitive property, stored like the other
storage secrets in /etc/pve/priv/storage/<storeid>.nvme-dhchap. It
reaches the target through stdin, and the plugin restricts the
dhchap_key and dhchap_ctrl_key attributes there to 0600 whenever it
writes the key. On the nodes, controllers are created with exactly one
write to /dev/nvme-fabrics in a bounded child (run_fork_with_timeout),
so the key never appears on a command line, and the secret attributes of
a new controller are restricted before the child returns. Rescans and
disconnects are sysfs writes; nvme-cli is not run.
Keep-alive, reconnect delay, controller loss and optional fast I/O
failure timeouts are configurable, with defaults that prefer waiting for
the target over failing guest I/O (ctrl_loss_tmo 600, fast_io_fail_tmo
unset).
Outside the plugin: register the type in PVE::Storage and the Makefile,
add it to @SHARED_STORAGE in PVE::Storage::Plugin, let 'pvesm add' and
'pvesm set' read --dhchap-key from a file, and depend on nvme-cli, which
creates /etc/nvme/hostnqn and /etc/nvme/hostid and remains the
administration tool, like the client packages of the other shared types.
Signed-off-by: Joaquin Varela <joaquinvarela@neatech.ar>
---
debian/control | 1 +
src/PVE/CLI/pvesm.pm | 21 +-
src/PVE/Storage.pm | 2 +
src/PVE/Storage/Makefile | 1 +
src/PVE/Storage/Plugin.pm | 2 +-
src/PVE/Storage/ZFSNVMePlugin.pm | 3177 ++++++++++++++++++++++++++++++
6 files changed, 3201 insertions(+), 3 deletions(-)
create mode 100644 src/PVE/Storage/ZFSNVMePlugin.pm
diff --git a/debian/control b/debian/control
index 850cd57c..ab46e6b6 100644
--- a/debian/control
+++ b/debian/control
@@ -43,6 +43,7 @@ Depends: bzip2,
lvm2,
lzop,
nfs-common,
+ nvme-cli,
proxmox-backup-client (>= 2.1.10~),
proxmox-backup-file-restore,
pve-cluster (>= 5.0-32),
diff --git a/src/PVE/CLI/pvesm.pm b/src/PVE/CLI/pvesm.pm
index 1aaa9f3c..3c68ffb5 100755
--- a/src/PVE/CLI/pvesm.pm
+++ b/src/PVE/CLI/pvesm.pm
@@ -78,12 +78,29 @@ sub param_mapping {
},
};
+ my $dhchap_key_map = {
+ name => 'dhchap-key',
+ desc => 'a file containing the NVMe DH-HMAC-CHAP key',
+ func => sub {
+ my ($value) = @_;
+ # never repeat a key passed in place of the file name
+ die "dhchap-key expects the path of a file containing the key\n"
+ if $value =~ /^DHHC-1:/i || -d $value;
+ my ($key) = split(/\n/, PVE::Tools::file_get_contents($value), 2);
+ $key = PVE::Tools::trim($key // '');
+ die "DH-HMAC-CHAP key file '$value' is empty\n" if $key eq '';
+ return $key;
+ },
+ };
+
my $mapping = {
'cifsscan' => [$password_map],
'cifs' => [$password_map],
'pbs' => [$password_map],
- 'create' => [$password_map, $enc_key_map, $master_key_map, $keyring_map],
- 'update' => [$password_map, $enc_key_map, $master_key_map, $keyring_map],
+ 'create' =>
+ [$password_map, $enc_key_map, $master_key_map, $keyring_map, $dhchap_key_map],
+ 'update' =>
+ [$password_map, $enc_key_map, $master_key_map, $keyring_map, $dhchap_key_map],
};
return $mapping->{$name};
}
diff --git a/src/PVE/Storage.pm b/src/PVE/Storage.pm
index fc3db812..2e5a8031 100755
--- a/src/PVE/Storage.pm
+++ b/src/PVE/Storage.pm
@@ -36,6 +36,7 @@ use PVE::Storage::CephFSPlugin;
use PVE::Storage::ISCSIDirectPlugin;
use PVE::Storage::ZFSPoolPlugin;
use PVE::Storage::ZFSPlugin;
+use PVE::Storage::ZFSNVMePlugin;
use PVE::Storage::PBSPlugin;
use PVE::Storage::BTRFSPlugin;
use PVE::Storage::ESXiPlugin;
@@ -61,6 +62,7 @@ PVE::Storage::CephFSPlugin->register();
PVE::Storage::ISCSIDirectPlugin->register();
PVE::Storage::ZFSPoolPlugin->register();
PVE::Storage::ZFSPlugin->register();
+PVE::Storage::ZFSNVMePlugin->register();
PVE::Storage::PBSPlugin->register();
PVE::Storage::BTRFSPlugin->register();
PVE::Storage::ESXiPlugin->register();
diff --git a/src/PVE/Storage/Makefile b/src/PVE/Storage/Makefile
index a67dc25f..d1cbfe29 100644
--- a/src/PVE/Storage/Makefile
+++ b/src/PVE/Storage/Makefile
@@ -11,6 +11,7 @@ SOURCES= \
ISCSIDirectPlugin.pm \
ZFSPoolPlugin.pm \
ZFSPlugin.pm \
+ ZFSNVMePlugin.pm \
PBSPlugin.pm \
BTRFSPlugin.pm \
LvmThinPlugin.pm \
diff --git a/src/PVE/Storage/Plugin.pm b/src/PVE/Storage/Plugin.pm
index a9e17513..a5aedc2b 100644
--- a/src/PVE/Storage/Plugin.pm
+++ b/src/PVE/Storage/Plugin.pm
@@ -35,7 +35,7 @@ our @COMMON_TAR_FLAGS = qw(
);
our @SHARED_STORAGE = (
- 'iscsi', 'nfs', 'cifs', 'rbd', 'cephfs', 'iscsidirect', 'zfs', 'drbd', 'pbs',
+ 'iscsi', 'nfs', 'cifs', 'rbd', 'cephfs', 'iscsidirect', 'zfs', 'zfsnvme', 'drbd', 'pbs',
);
our $QCOW2_PREALLOCATION = {
diff --git a/src/PVE/Storage/ZFSNVMePlugin.pm b/src/PVE/Storage/ZFSNVMePlugin.pm
new file mode 100644
index 00000000..127680fb
--- /dev/null
+++ b/src/PVE/Storage/ZFSNVMePlugin.pm
@@ -0,0 +1,3177 @@
+package PVE::Storage::ZFSNVMePlugin;
+
+use v5.36;
+
+use Compress::Zlib qw(crc32);
+use Digest::SHA qw(sha256_hex);
+use Errno qw(ENOENT);
+use Fcntl qw(O_NOFOLLOW O_RDONLY O_RDWR S_ISCHR);
+use File::Path qw(make_path);
+use List::Util qw(max);
+use MIME::Base64 qw(decode_base64);
+use POSIX qw(SIG_BLOCK SIG_SETMASK SIGHUP SIGINT SIGQUIT SIGTERM sigprocmask);
+use Socket qw(AF_INET6 inet_ntop inet_pton);
+use Time::HiRes;
+
+use PVE::Cluster;
+use PVE::File;
+use PVE::JSONSchema;
+use PVE::Network;
+use PVE::RESTEnvironment qw(log_warn);
+use PVE::RPCEnvironment;
+use PVE::SysFSTools;
+use PVE::Tools qw(file_read_firstline file_set_contents run_command trim);
+
+use base qw(PVE::Storage::Plugin);
+
+# ZFS zvols on a remote Linux target, exported through the kernel NVMe target
+# (nvmet, configured through configfs) over NVMe/TCP and consumed with native
+# NVMe multipath.
+#
+# The ZFS dataset is the source of truth for a namespace: its owner NQN, NSID
+# and UUID are ZFS user properties. Everything else on the target is derived
+# state that the plugin rebuilds after a target restart.
+#
+# Storage parsing, validation and planning happen locally. SSH carries quoted
+# commands, sometimes joined with `&&`, to read or change ZFS/configfs state.
+# A pmxcfs domain lock per target serializes the changes within the cluster,
+# held like the storage lock of the other shared storage types.
+
+# ---------------------------------------------------------------------------
+# Constants and regular expressions
+# ---------------------------------------------------------------------------
+
+# LogLevel=ERROR keeps a login banner out of the first stderr line, which is
+# the error text of a failed call.
+my @ssh_cmd = (
+ '/usr/bin/ssh', '-o', 'BatchMode=yes', '-o', 'ConnectTimeout=10', '-o', 'LogLevel=ERROR',
+);
+my $id_rsa_path = '/etc/pve/priv/zfs';
+my $secret_dir = '/etc/pve/priv/storage';
+my $max_paths = 16;
+my $max_hosts = 64;
+my $slow_path_backoff = 60;
+
+my $nvmet_root = '/sys/kernel/config/nvmet';
+my $nvmet_model = 'Proxmox ZFS NVMe';
+my $nvmet_marker = 'ZFSNVME-CONFIGFS';
+my $nvmet_max_nsid = 0xfffffffe;
+my $nvmet_max_command = 65536;
+# Seconds to wait for the target lock while another node changes the target;
+# a destroy or rollback can hold it for several seconds.
+my $nvmet_lock_wait = 30;
+# On the pool dataset: the highest NSID handed out for a volume of the pool.
+my $nvmet_last_nsid = 'proxmox:nvme-last-nsid';
+
+my @nvmet_zfs_props =
+ ('type', 'proxmox:nvme-subsys', 'proxmox:nvme-nsid', 'proxmox:nvme-uuid', $nvmet_last_nsid);
+my %nvmet_zfs_keys = (
+ type => 'type',
+ 'proxmox:nvme-subsys' => 'nqn',
+ 'proxmox:nvme-nsid' => 'nsid',
+ 'proxmox:nvme-uuid' => 'uuid',
+ $nvmet_last_nsid => 'last_nsid',
+);
+my @nvmet_cfs_attrs = qw(
+ enable device_path device_uuid buffered_io
+ attr_model attr_serial attr_allow_any_host
+ addr_trtype addr_adrfam addr_traddr addr_trsvcid
+);
+# Every command the plugin runs on the target, besides the `dd | tee` pipe of
+# a key write. The configfs read runs find through env to set its locale.
+my %nvmet_commands = map { $_ => 1 } qw(
+ cat chmod env grep ln mkdir modprobe mount printf rm rmdir test zfs
+);
+
+my $RE_NQN = qr{
+ \A
+ nqn \.
+ [A-Za-z0-9] [A-Za-z0-9.-]*
+ :
+ [A-Za-z0-9] [A-Za-z0-9._:-]*
+ \z
+}nxx;
+my $RE_IPV4_PORTAL = qr{
+ \A
+ (?<address> [^:]+)
+ (?: : (?<port> [0-9]+))?
+ \z
+}nxx;
+my $RE_IPV6_PORTAL = qr{
+ \A
+ \[ (?<address> [^\]]+) \]
+ (?: : (?<port> [0-9]+))?
+ \z
+}nxx;
+# A canonical IPv4-mapped IPv6 address.
+my $RE_IPV4_MAPPED = qr{\A ::ffff: (?<address> [0-9.]+) \z}nxx;
+my $RE_HOST_IFACE = qr{\A [A-Za-z0-9_.-]+ \z}nxx;
+my $RE_DHCHAP_KEY = qr{
+ \A DHHC-1 : (?<hash> 0[0-3]) : (?<secret> [A-Za-z0-9+/]+ ={0,2}) : \z
+}nxx;
+# A stopped PVE task: the worker's signal handler dies with the first text,
+# run_command adds the command context, and run_fork_with_timeout reports
+# the signal in its own words.
+my $RE_TASK_INTERRUPT = qr{
+ \A (?: (?: command \x20 ' .* ' \x20 failed: \x20 )? received \x20 interrupt
+ | interrupted \x20 by \x20 unexpected \x20 signal ) \n \z
+}nsxx;
+my $RE_COMMAND_TIMEOUT = qr{
+ \A (?: command \x20 ' .* ' \x20 failed: \x20 )? got \x20 timeout \n \z
+}nsxx;
+my $RE_COMMAND_EXIT =
+ qr{\A command \x20 ' .* ' \x20 failed: \x20 exit \x20 code \x20 (?<code> [0-9]+) \n \z}nsxx;
+my $RE_FABRICS_VALUE = qr{\A [^\s,\x00]+ \z}nxx;
+my $RE_FABRICS_RESULT = qr{\A instance=(?<instance> [0-9]+) ,cntlid= [0-9]+ \n \z}nxx;
+my $RE_ZERO_UUID = qr{\A 0{8} (?: -0{4}){3} -0{12} \z}nxx;
+my $RE_CONFIG_INT = qr{\A -?[0-9]+ \z}nxx;
+my $RE_NVME_CONTROLLER = qr{\A nvme [0-9]+ \z}nxx;
+my $RE_NVME_SUBSYSTEM = qr{\A nvme-subsys [0-9]+ \z}nxx;
+my $RE_NVME_NAMESPACE = qr{\A nvme [0-9]+ n [0-9]+ \z}nxx;
+my $RE_NVME_PARTITION = qr{\A nvme [0-9]+ n [0-9]+ p [0-9]+ \z}nxx;
+my $RE_TRADDR = qr{(?: \A | ,) traddr=(?<value>[^,]+)}nxx;
+my $RE_TRSVCID = qr{(?: \A | ,) trsvcid=(?<value>[^,]+)}nxx;
+my $RE_HOST_IFACE_ADDRESS = qr{(?: \A | ,) host_iface=(?<value>[^,]+)}nxx;
+my $RE_ZVOL_OWNER = qr{^ (?:vm|base|subvol|basevol)- (?<owner>\d+) - \S+ $}nxx;
+my $RE_VOLNAME = qr{
+ ^
+ (?: (?<base> (?:base|basevol)- (?<base_vmid>\d+) - \S+) /)?
+ (?<name> (?<type>base|basevol|vm|subvol)- (?<vmid>\d+) - \S+)
+ $
+}nxx;
+my $RE_BASE_SNAPSHOT = qr{^ (?<base>\S+) \@__base__ $}nxx;
+my $RE_UNSIGNED_INTEGER = qr{^ (?<value>\d+) $}nxx;
+
+my $RE_NVMET_UUID = qr{\A [0-9a-fA-F]{8} (?: - [0-9a-fA-F]{4}){3} - [0-9a-fA-F]{12} \z}nxx;
+my $RE_NVMET_NSID = qr{\A [1-9] [0-9]* \z}nxx;
+my $RE_NVMET_POOL = qr{\A [A-Za-z0-9] [A-Za-z0-9_.:/-]* \z}nxx;
+my $RE_NVMET_DATASET_NAME = qr{\A [A-Za-z0-9] [A-Za-z0-9_.:-]* \z}nxx;
+my $RE_NVMET_SNAPSHOT = qr{\A [A-Za-z0-9_.:-]+ \z}nxx;
+my $RE_NVMET_TEMPLATE_ZVOL = qr{/ (?<type> vm | base) - (?<suffix> [^/]+) \z}nxx;
+my $RE_NVMET_INHERITED = qr{\A inherited \x20 from \x20 (?<source> .+) \z}nxx;
+# Errors after which the target state is not known: an interrupted task, or a
+# change whose command may still be running on the target.
+my $RE_NVMET_ABANDONED = qr{
+ (?: \A received \x20 interrupt | ; \x20 the \x20 target \x20 state \x20 is \x20 unknown) \n \z
+}nxx;
+
+# Lines of the configfs read (nvmet_configfs_step below): `D <dir>`,
+# `L <link>`, `<file>:<value>` from grep, and `<sha256> <file>` from
+# sha256sum. Attribute names anchor the split, so an NQN containing ':' is
+# never ambiguous.
+my $cfs_root = quotemeta($nvmet_root);
+my $RE_CFS_TOP = qr{\A D \x20 $cfs_root (?: / (?<top> hosts | ports | subsystems))? \z}nxx;
+my $RE_CFS_SUBSYS_DIR = qr{\A D \x20 $cfs_root /subsystems/ (?<nqn> [^/]+) \z}nxx;
+my $RE_CFS_SUBSYS_GROUP = qr{
+ \A D \x20 $cfs_root /subsystems/ [^/]+ / (?: namespaces | allowed_hosts) \z
+}nxx;
+my $RE_CFS_NS_DIR = qr{
+ \A D \x20 $cfs_root /subsystems/ (?<nqn> [^/]+) /namespaces/ (?<nsid> [0-9]+) \z
+}nxx;
+my $RE_CFS_PORT_DIR = qr{\A D \x20 $cfs_root /ports/ (?<port> [0-9]+) \z}nxx;
+my $RE_CFS_PORT_GROUP = qr{\A D \x20 $cfs_root /ports/ [0-9]+ /subsystems \z}nxx;
+my $RE_CFS_HOST_DIR = qr{\A D \x20 $cfs_root /hosts/ (?<host> [^/]+) \z}nxx;
+my $RE_CFS_PORT_LINK = qr{
+ \A L \x20 $cfs_root /ports/ (?<port> [0-9]+) /subsystems/ (?<nqn> [^/]+) \z
+}nxx;
+my $RE_CFS_ACL_LINK = qr{
+ \A L \x20 $cfs_root /subsystems/ (?<nqn> [^/]+) /allowed_hosts/ (?<host> [^/]+) \z
+}nxx;
+my $RE_CFS_OTHER_ENTRY = qr{\A [DL] \x20 $cfs_root /}nxx;
+my $RE_CFS_NS_ATTR = qr{
+ \A $cfs_root /subsystems/ (?<nqn> [^/]+) /namespaces/ (?<nsid> [0-9]+)
+ / (?<attr> enable | device_path | device_uuid | buffered_io) : (?<value> .*) \z
+}nxx;
+my $RE_CFS_SUBSYS_ATTR = qr{
+ \A $cfs_root /subsystems/ (?<nqn> [^/]+)
+ / (?<attr> attr_model | attr_serial | attr_allow_any_host) : (?<value> .*) \z
+}nxx;
+my $RE_CFS_PORT_ATTR = qr{
+ \A $cfs_root /ports/ (?<port> [0-9]+)
+ / (?<attr> addr_trtype | addr_adrfam | addr_traddr | addr_trsvcid) : (?<value> .*) \z
+}nxx;
+my $RE_CFS_KEY_DIGEST = qr{
+ \A (?<sha> [0-9a-f]{64}) \x20 [\x20*] $cfs_root /hosts/ (?<host> [^/]+) /dhchap_key \z
+}nxx;
+my $RE_CFS_OTHER_ATTR = qr{\A $cfs_root /}nxx;
+
+# ---------------------------------------------------------------------------
+# Target addressing and validated names
+# ---------------------------------------------------------------------------
+
+my sub nvmet_server($scfg) {
+ return $scfg->{server};
+}
+
+my sub nvmet_ssh_key($scfg) {
+ return "$id_rsa_path/" . nvmet_server($scfg) . '_id_rsa';
+}
+
+my sub nvmet_lock_id($scfg) {
+ my $server = nvmet_server($scfg) // '';
+ return 'zfsnvme-' . ($server =~ s/[^A-Za-z0-9.-]/_/gr);
+}
+
+my sub secret_path($storeid) {
+ return "$secret_dir/$storeid.nvme-dhchap";
+}
+
+my sub nvmet_nqn($scfg) {
+ my $nqn = $scfg->{subsysnqn};
+ die "invalid NVMe subsystem NQN\n" if !defined($nqn) || !verify_nvme_nqn($nqn, 1);
+ return $nqn;
+}
+
+my sub nvmet_pool($scfg) {
+ my $pool = $scfg->{pool};
+ die "invalid ZFS pool name\n" if !defined($pool) || $pool !~ $RE_NVMET_POOL;
+ return $pool;
+}
+
+my sub nvmet_dataset($scfg, $name) {
+ die "invalid ZFS volume name\n" if !defined($name) || $name !~ $RE_NVMET_DATASET_NAME;
+ return nvmet_pool($scfg) . "/$name";
+}
+
+# zfs destroy reads '%' and ',' in a snapshot name as a range and a list.
+my sub nvmet_snapshot($scfg, $name, $snap) {
+ die "invalid snapshot name\n" if !defined($snap) || $snap !~ $RE_NVMET_SNAPSHOT;
+ return nvmet_dataset($scfg, $name) . "\@$snap";
+}
+
+my sub nvmet_valid_nsid($value) {
+ return defined($value) && $value =~ $RE_NVMET_NSID && $value <= $nvmet_max_nsid;
+}
+
+# IPv6 addresses have many spellings; compare and store only the canonical one.
+my sub nvmet_canonical_address($address) {
+ return $address if !defined($address) || index($address, ':') < 0;
+ my $packed = inet_pton(AF_INET6, $address) // return $address;
+ return inet_ntop(AF_INET6, $packed);
+}
+
+# The address family and address that a portal of parse_nvme_portals reaches:
+# an IPv4-mapped IPv6 address reaches the IPv4 address.
+my sub nvmet_listener_address($portal) {
+ return ('ipv4', $+{address})
+ if $portal->{family} eq 'ipv6' && $portal->{address} =~ $RE_IPV4_MAPPED;
+ return $portal->@{qw(family address)};
+}
+
+# Target-provided names only appear in messages when they are well-formed.
+my sub nvmet_label($name) {
+ return $name =~ $RE_NQN ? " '$name'" : '';
+}
+
+# The other name of a zvol that a template conversion renames: vm-* and base-*.
+my sub nvmet_template_twin($name) {
+ return '' if $name !~ $RE_NVMET_TEMPLATE_ZVOL;
+ my $twin = ($+{type} eq 'vm' ? 'base' : 'vm') . "-$+{suffix}";
+ return substr($name, 0, $-[0]) . "/$twin";
+}
+
+# The timeout of a call that changes ZFS: ZFS waits for transaction groups,
+# which a busy pool can take minutes to sync.
+my sub nvmet_long_timeout() {
+ return PVE::RPCEnvironment->is_worker() ? 60 * 60 : 60;
+}
+
+# ---------------------------------------------------------------------------
+# Remote steps
+#
+# A step is an argv array, a write `{ write => $path, value => $value }` or a
+# key write `{ key => [$path, ...] }` fed from stdin. A unit is a list of steps
+# that must stay in one call.
+# ---------------------------------------------------------------------------
+
+my sub nvmet_subsys_path($nqn) {
+ return "$nvmet_root/subsystems/$nqn";
+}
+
+my sub nvmet_ns_path($nqn, $nsid) {
+ return "$nvmet_root/subsystems/$nqn/namespaces/$nsid";
+}
+
+my sub nvmet_port_path($id) {
+ return "$nvmet_root/ports/$id";
+}
+
+my sub nvmet_host_path($hostnqn) {
+ return "$nvmet_root/hosts/$hostnqn";
+}
+
+my sub nvmet_write($path, $value) {
+ return { write => $path, value => $value };
+}
+
+# Parts of a parsed configfs state; missing parts read as empty.
+my sub nvmet_namespaces($cfs, $nqn) {
+ my $subsys = $cfs->{subsystems}->{$nqn};
+ return ($subsys ? $subsys->{namespaces} : undef) // {};
+}
+
+my sub nvmet_port_linked($cfs, $id, $nqn) {
+ my $port = $cfs->{ports}->{$id};
+ return $port && ($port->{links} // {})->{$nqn};
+}
+
+my sub nvmet_linked_hosts($cfs) {
+ my %linked;
+ for my $subsys (values $cfs->{subsystems}->%*) {
+ $linked{$_} = 1 for keys(($subsys->{acl} // {})->%*);
+ }
+ return \%linked;
+}
+
+# A subsystem whose creation by this plugin stopped after its model was
+# written, before its serial: nobody else writes our model, and it exports,
+# allows and publishes nothing, so completing it takes nothing from anybody.
+# A subsystem with the kernel's default model may belong to another tool.
+my sub nvmet_unfinished_subsystem($cfs, $nqn) {
+ my $subsys = $cfs->{subsystems}->{$nqn};
+ return 0 if !$subsys;
+ return
+ $subsys->{attr_model} eq $nvmet_model
+ && !($subsys->{namespaces} // {})->%*
+ && !($subsys->{acl} // {})->%*
+ && !grep { nvmet_port_linked($cfs, $_, $nqn) } keys $cfs->{ports}->%*;
+}
+
+# Same write order as the kernel requires: the device before enabling it.
+my sub nvmet_build_unit($nqn, $nsid, $uuid, $dev) {
+ my $ns = nvmet_ns_path($nqn, $nsid);
+ return [
+ ['test', '-b', $dev],
+ ['mkdir', $ns],
+ nvmet_write("$ns/device_path", $dev),
+ nvmet_write("$ns/device_uuid", $uuid),
+ nvmet_write("$ns/buffered_io", 0),
+ nvmet_write("$ns/enable", 1),
+ ];
+}
+
+my sub nvmet_enable_unit($nqn, $nsid, $dev) {
+ return [['test', '-b', $dev], nvmet_write(nvmet_ns_path($nqn, $nsid) . '/enable', 1)];
+}
+
+my sub nvmet_unexport_unit($nqn, $nsid, $enabled) {
+ my $ns = nvmet_ns_path($nqn, $nsid);
+ return [($enabled ? (nvmet_write("$ns/enable", 0)) : ()), ['rmdir', $ns]];
+}
+
+my sub nvmet_identity($nqn, $nsid, $uuid) {
+ return ("proxmox:nvme-subsys=$nqn", "proxmox:nvme-nsid=$nsid", "proxmox:nvme-uuid=$uuid");
+}
+
+# The pool and its direct children with their identity properties.
+my sub nvmet_zfs_inventory_step($pool) {
+ return [
+ 'zfs',
+ 'get',
+ '-H',
+ '-p',
+ '-d',
+ '1',
+ '-t',
+ 'filesystem,volume',
+ '-o',
+ 'name,property,value,source',
+ join(',', @nvmet_zfs_props),
+ $pool,
+ ];
+}
+
+# Every configfs object the plugin uses, in one traversal. Keys are never
+# read back; with host NQNs, only their sha256 digests are returned. Only the
+# kernel's own groups of objects the plugin does not use are skipped, so an
+# object of another tool with one of their names is still listed.
+my sub nvmet_configfs_step($hostnqns) {
+ my @names = map { ('-o', '-name', $_) } @nvmet_cfs_attrs;
+ shift @names;
+ my @keys = map { ('-o', '-path', nvmet_host_path($_) . '/dhchap_key') } $hostnqns->@*;
+ shift @keys;
+ return [
+ 'env',
+ 'LC_ALL=C',
+ 'find',
+ $nvmet_root,
+ '(',
+ '-path',
+ "$nvmet_root/subsystems/*/passthru",
+ '-o',
+ '-path',
+ "$nvmet_root/ports/*/ana_groups",
+ '-o',
+ '-path',
+ "$nvmet_root/ports/*/referrals",
+ ')',
+ '-type',
+ 'd',
+ '-prune',
+ '-o',
+ '-type',
+ 'd',
+ '-exec',
+ 'printf',
+ 'D %s\n',
+ '{}',
+ '+',
+ '-o',
+ '-type',
+ 'l',
+ '-exec',
+ 'printf',
+ 'L %s\n',
+ '{}',
+ '+',
+ '-o',
+ '-type',
+ 'f',
+ '(',
+ @names,
+ ')',
+ '-exec',
+ 'grep',
+ '',
+ '/dev/null',
+ '{}',
+ '+',
+ (@keys ? ('-o', '-type', 'f', '(', @keys, ')', '-exec', 'sha256sum', '{}', '+') : ()),
+ ];
+}
+
+# The whole target state: the ZFS inventory, a marker and configfs. Only a
+# locked read loads nvmet first (modprobe).
+my sub nvmet_state_steps($pool, $hostnqns, $modprobe) {
+ return [
+ ($modprobe ? (['modprobe', 'nvmet_tcp']) : ()),
+ nvmet_zfs_inventory_step($pool),
+ ['printf', '%s\n', $nvmet_marker],
+ nvmet_configfs_step($hostnqns),
+ ];
+}
+
+# ---------------------------------------------------------------------------
+# Renderer and runner
+# ---------------------------------------------------------------------------
+
+my sub nvmet_valid_word($word) {
+ return defined($word) && !ref($word) && $word !~ /[\0\r\n]/;
+}
+
+my sub nvmet_valid_path($path) {
+ return nvmet_valid_word($path) && $path =~ m{\A/};
+}
+
+my sub nvmet_quote($word) {
+ return PVE::Tools::shellquote($word);
+}
+
+# Renders steps into one POSIX shell command line: simple commands joined with
+# `&&`, plus `>` for configfs writes and one `dd | tee` pipe for a key read from
+# stdin. There are no loops, variables, conditionals or substitutions. Every
+# word is quoted; all operands are built from validated configuration.
+sub _nvmet_render($steps) {
+ my $invalid = "internal error: invalid NVMe target step\n";
+ die $invalid if ref($steps) ne 'ARRAY' || !$steps->@*;
+
+ my $keys = 0;
+ my @commands;
+ for my $step ($steps->@*) {
+ if (ref($step) eq 'ARRAY') {
+ die $invalid if !$step->@* || grep { !nvmet_valid_word($_) } $step->@*;
+ die $invalid if !$nvmet_commands{ $step->[0] };
+ push @commands, join(' ', map { nvmet_quote($_) } $step->@*);
+ } elsif (ref($step) eq 'HASH' && exists($step->{write})) {
+ die $invalid
+ if keys($step->%*) != 2
+ || !nvmet_valid_path($step->{write})
+ || !nvmet_valid_word($step->{value});
+ push @commands,
+ q{printf '%s\n' }
+ . nvmet_quote($step->{value}) . ' > '
+ . nvmet_quote($step->{write});
+ } elsif (ref($step) eq 'HASH' && exists($step->{key})) {
+ my $paths = $step->{key};
+ die $invalid
+ if keys($step->%*) != 1
+ || ref($paths) ne 'ARRAY'
+ || !$paths->@*
+ || grep { !nvmet_valid_path($_) } $paths->@*;
+ $keys++;
+ die $invalid if $keys > 1;
+ # configfs needs the whole key in one write(2). dd collects its
+ # input into one output block: GNU dd without bs=, busybox dd only
+ # when ibs and obs differ. tee writes it to every host object.
+ push @commands,
+ 'dd ibs=4096 obs=8192 2>/dev/null | tee '
+ . join(' ', map { nvmet_quote($_) } $paths->@*)
+ . ' >/dev/null';
+ } else {
+ die $invalid;
+ }
+ }
+ return join(' && ', @commands);
+}
+
+# Packs whole units into calls of at most $max bytes of rendered command, the
+# single argument sshd runs, below its limit. A rendered command is ASCII
+# (every operand is validated), so its length is its size in bytes.
+sub _nvmet_chunk($units, $max = $nvmet_max_command) {
+ my (@chunks, @current);
+ for my $unit ($units->@*) {
+ my @next = (@current, $unit->@*);
+ if (length(_nvmet_render(\@next)) > $max) {
+ die "internal error: NVMe target command too long\n"
+ if !@current || length(_nvmet_render($unit)) > $max;
+ push @chunks, [@current];
+ @next = $unit->@*;
+ }
+ @current = @next;
+ }
+ push @chunks, [@current] if @current;
+ return \@chunks;
+}
+
+# Never propagate the context run_command adds to the task marker: it quotes
+# the command line.
+my sub rethrow_task_interrupt($error) {
+ die "received interrupt\n" if $error =~ $RE_TASK_INTERRUPT;
+}
+
+# The only code that runs anything on the target. Returns
+# { rc => exit code, or -1 when ssh did not run to its end, out => [lines],
+# err => text } and never dies on a failing command; a stopped task dies with
+# the task marker. %opts: op (label, required), timeout and input (stdin,
+# used for keys).
+sub _nvmet_run($scfg, $steps, %opts) {
+ die "internal error: NVMe target call without label\n" if !defined($opts{op});
+ my $command = _nvmet_render($steps);
+ die "internal error: NVMe target command too long\n" if length($command) > $nvmet_max_command;
+ my $cmd = [@ssh_cmd, '-i', nvmet_ssh_key($scfg), 'root@' . nvmet_server($scfg), $command];
+ my (@out, $err);
+ my $rc = 0;
+ # Not noerr: run_command would then also swallow the interrupt of a stopped
+ # task, which must end the call. The exit code is taken from the exception.
+ eval {
+ run_command(
+ $cmd,
+ timeout => $opts{timeout} // 15,
+ outfunc => sub($line) { push @out, $line },
+ errfunc => sub($line) { $err = $line if !defined($err) && $line ne '' },
+ (defined($opts{input}) ? (input => $opts{input}) : ()),
+ );
+ };
+ if (my $error = $@) {
+ rethrow_task_interrupt($error);
+ return { rc => -1, out => [], err => 'timeout' } if $error =~ $RE_COMMAND_TIMEOUT;
+ return { rc => -1, out => [], err => 'ssh failed' } if $error !~ $RE_COMMAND_EXIT;
+ $rc = $+{code};
+ }
+ return { rc => $rc, out => \@out, err => $err // ($rc ? "exit code $rc" : '') };
+}
+
+# Monotonic clock for the activation backoff and local waits (a test seam).
+sub _now() {
+ return Time::HiRes::clock_gettime(Time::HiRes::CLOCK_MONOTONIC());
+}
+
+# Every local wait of the plugin (a test seam).
+sub _sleep($seconds) {
+ Time::HiRes::sleep($seconds);
+ return;
+}
+
+# ---------------------------------------------------------------------------
+# Parsers (pure)
+# ---------------------------------------------------------------------------
+
+sub _nvmet_split_state($text) {
+ my @parts = split /^\Q$nvmet_marker\E\n/m, $text, -1;
+ die "malformed NVMe target state\n" if @parts != 2;
+ return @parts;
+}
+
+sub _nvmet_property_value($value, $source) {
+ die "malformed ZFS property value\n"
+ if !defined($value)
+ || !defined($source)
+ || $value eq ''
+ || $source eq ''
+ || $value =~ /[\r\n\t]/
+ || $source =~ /[\r\n\t]/;
+ return $value if $source eq 'local' || $source eq 'received';
+ return '-' if $value eq '-' && $source eq '-';
+ if ($source =~ $RE_NVMET_INHERITED) {
+ die "invalid ZFS pool name\n" if $+{source} !~ $RE_NVMET_POOL;
+ return '-';
+ }
+ die "invalid ZFS property source\n";
+}
+
+# Parses the ZFS inventory into { types => { dataset => type }, volumes =>
+# { dataset => { nqn, nsid, uuid } }, last_nsid => highest NSID handed out }.
+# The inventory must be complete: the pool rows and all rows of every
+# dataset, or the read fails. Children whose names PVE never uses (for
+# example with a space) are kept as datasets but never count as owned.
+sub _nvmet_parse_zfs_inventory($text, $pool) {
+ die "invalid ZFS pool name\n" if $pool !~ $RE_NVMET_POOL;
+
+ my (%rows, %foreign);
+ for my $line (split /\n/, $text) {
+ my @fields = split /\t/, $line, -1;
+ die "malformed ZFS inventory row\n" if @fields != 4 || $line =~ /\r/;
+ my ($dataset, $property, $value, $source) = @fields;
+ if ($dataset ne $pool) {
+ die "ZFS dataset is outside configured pool\n" if index($dataset, "$pool/") != 0;
+ my $child = substr($dataset, length($pool) + 1);
+ die "ZFS dataset is outside configured pool\n"
+ if $child eq '' || index($child, '/') >= 0;
+ $foreign{$dataset} = 1 if $child !~ $RE_NVMET_DATASET_NAME;
+ }
+ my $on = $foreign{$dataset} ? '' : " on '$dataset'";
+ my $key = $nvmet_zfs_keys{$property} // die "unexpected ZFS property$on\n";
+ die "duplicate ZFS property '$property'$on\n" if exists($rows{$dataset}->{$key});
+ if ($key eq 'type') {
+ die "invalid ZFS dataset type$on\n"
+ if ($value ne 'filesystem' && $value ne 'volume') || $source ne '-';
+ $rows{$dataset}->{type} = $value;
+ } else {
+ $rows{$dataset}->{$key} = _nvmet_property_value($value, $source);
+ }
+ }
+ die "ZFS dataset inventory is missing the configured pool\n" if !$rows{$pool};
+
+ my (%types, %volumes);
+ for my $dataset (keys %rows) {
+ my $row = $rows{$dataset};
+ die "incomplete ZFS properties" . ($foreign{$dataset} ? '' : " on '$dataset'") . "\n"
+ if keys($row->%*) != scalar(@nvmet_zfs_props);
+ $types{$dataset} = $row->{type};
+ next if $row->{type} ne 'volume';
+ # nvmet shows device_uuid and udev names the by-id link in lower case,
+ # so identities are compared in lower case.
+ $volumes{$dataset} =
+ $foreign{$dataset}
+ ? { nqn => '-', nsid => '-', uuid => '-' }
+ : { nqn => $row->{nqn}, nsid => $row->{nsid}, uuid => lc($row->{uuid}) };
+ }
+ my $last_nsid = $rows{$pool}->{last_nsid};
+ return {
+ types => \%types,
+ volumes => \%volumes,
+ last_nsid => nvmet_valid_nsid($last_nsid) ? $last_nsid : 0,
+ };
+}
+
+# Parses the configfs read into
+# { subsystems => { nqn => { attr_model, attr_serial, attr_allow_any_host,
+# acl => { hostnqn => 1 }, namespaces => { nsid => { enable, device_path,
+# device_uuid, buffered_io } } } },
+# ports => { id => { addr_trtype, addr_adrfam, addr_traddr, addr_trsvcid,
+# links => { nqn => 1 } } },
+# hosts => { hostnqn => { key_sha256 => hex or undef } } }
+# Every object must be complete; unknown objects of newer kernels are ignored.
+sub _nvmet_parse_configfs($text) {
+ my $state = { subsystems => {}, ports => {}, hosts => {} };
+ my (%top, %dir);
+
+ my $subsystem = sub($nqn) {
+ return $state->{subsystems}->{$nqn} //= { acl => {}, namespaces => {} };
+ };
+ my $namespace = sub($nqn, $nsid) {
+ return $subsystem->($nqn)->{namespaces}->{$nsid} //= {};
+ };
+ my $port = sub($id) { return $state->{ports}->{$id} //= { links => {} } };
+ my $host = sub($hostnqn) { return $state->{hosts}->{$hostnqn} //= { key_sha256 => undef } };
+ my $set = sub($object, $attr, $value) {
+ die "duplicate NVMe target configfs attribute\n" if exists($object->{$attr});
+ $object->{$attr} = $value;
+ };
+
+ for my $line (split /\n/, $text) {
+ if ($line =~ $RE_CFS_TOP) {
+ $top{ $+{top} // 'root' } = 1;
+ } elsif ($line =~ $RE_CFS_SUBSYS_DIR) {
+ my $nqn = $+{nqn};
+ $subsystem->($nqn);
+ $dir{"subsystem $nqn"} = 1;
+ } elsif ($line =~ $RE_CFS_SUBSYS_GROUP || $line =~ $RE_CFS_PORT_GROUP) {
+ # default groups, always present with their parent
+ } elsif ($line =~ $RE_CFS_NS_DIR) {
+ my ($nqn, $nsid) = @+{qw(nqn nsid)};
+ $namespace->($nqn, $nsid);
+ $dir{"namespace $nqn/$nsid"} = 1;
+ } elsif ($line =~ $RE_CFS_PORT_DIR) {
+ my $id = $+{port};
+ $port->($id);
+ $dir{"port $id"} = 1;
+ } elsif ($line =~ $RE_CFS_HOST_DIR) {
+ my $hostnqn = $+{host};
+ $host->($hostnqn);
+ $dir{"host $hostnqn"} = 1;
+ } elsif ($line =~ $RE_CFS_PORT_LINK) {
+ my ($id, $nqn) = @+{qw(port nqn)};
+ $port->($id)->{links}->{$nqn} = 1;
+ } elsif ($line =~ $RE_CFS_ACL_LINK) {
+ my ($nqn, $hostnqn) = @+{qw(nqn host)};
+ $subsystem->($nqn)->{acl}->{$hostnqn} = 1;
+ } elsif ($line =~ $RE_CFS_OTHER_ENTRY) {
+ # objects of a shape the plugin does not use
+ } elsif ($line =~ $RE_CFS_NS_ATTR) {
+ my ($nqn, $nsid, $attr, $value) = @+{qw(nqn nsid attr value)};
+ $set->($namespace->($nqn, $nsid), $attr, $value);
+ } elsif ($line =~ $RE_CFS_SUBSYS_ATTR) {
+ my ($nqn, $attr, $value) = @+{qw(nqn attr value)};
+ $set->($subsystem->($nqn), $attr, $value);
+ } elsif ($line =~ $RE_CFS_PORT_ATTR) {
+ my ($id, $attr, $value) = @+{qw(port attr value)};
+ $set->($port->($id), $attr, $value);
+ } elsif ($line =~ $RE_CFS_KEY_DIGEST) {
+ my ($sha, $hostnqn) = @+{qw(sha host)};
+ my $object = $host->($hostnqn);
+ die "duplicate DH-HMAC-CHAP key digest" . nvmet_label($hostnqn) . "\n"
+ if defined($object->{key_sha256});
+ $object->{key_sha256} = $sha;
+ } elsif ($line =~ $RE_CFS_OTHER_ATTR) {
+ # attribute of an object shape the plugin does not use
+ } else {
+ die "unexpected NVMe target configfs line\n";
+ }
+ }
+
+ for my $name (qw(root hosts ports subsystems)) {
+ die "NVMe target configfs is unavailable\n" if !$top{$name};
+ }
+ for my $nqn (keys $state->{subsystems}->%*) {
+ my $subsys = $state->{subsystems}->{$nqn};
+ die "incomplete NVMe subsystem" . nvmet_label($nqn) . "\n"
+ if !$dir{"subsystem $nqn"}
+ || grep { !defined($subsys->{$_}) } qw(attr_model attr_serial attr_allow_any_host);
+ for my $nsid (keys $subsys->{namespaces}->%*) {
+ my $ns = $subsys->{namespaces}->{$nsid};
+ die "incomplete NVMe namespace '$nsid' of subsystem" . nvmet_label($nqn) . "\n"
+ if !$dir{"namespace $nqn/$nsid"}
+ || grep { !defined($ns->{$_}) } qw(enable device_path device_uuid buffered_io);
+ }
+ }
+ for my $id (keys $state->{ports}->%*) {
+ my $object = $state->{ports}->{$id};
+ my @attrs = qw(addr_trtype addr_adrfam addr_traddr addr_trsvcid);
+ die "incomplete NVMe port '$id'\n"
+ if !$dir{"port $id"} || grep { !defined($object->{$_}) } @attrs;
+ }
+ for my $hostnqn (keys $state->{hosts}->%*) {
+ die "DH-HMAC-CHAP key digest for unknown NVMe host" . nvmet_label($hostnqn) . "\n"
+ if !$dir{"host $hostnqn"};
+ }
+ return $state;
+}
+
+sub _nvmet_parse_mounts($text) {
+ for my $line (split /\n/, $text) {
+ my (undef, $mountpoint, $type) = split / /, $line;
+ return 1 if ($mountpoint // '') eq '/sys/kernel/config' && ($type // '') eq 'configfs';
+ }
+ return 0;
+}
+
+# ---------------------------------------------------------------------------
+# Planners (pure)
+# ---------------------------------------------------------------------------
+
+sub _nvmet_serial($nqn) {
+ return 'PVEZFS' . substr(sha256_hex($nqn), 0, 14);
+}
+
+sub _nvmet_template_name($dataset) {
+ die "only VM zvols can become templates\n"
+ if $dataset !~ $RE_NVMET_TEMPLATE_ZVOL || $+{type} ne 'vm';
+ return nvmet_template_twin($dataset);
+}
+
+# The durable identity (NSID, UUID) of a volume owned by $nqn. Refuses a
+# volume that another owned volume duplicates anywhere in the pool.
+sub _nvmet_owned_identity($inv, $nqn, $dataset) {
+ my $type = $inv->{types}->{$dataset} // die "ZFS volume '$dataset' does not exist\n";
+ die "'$dataset' is not a ZFS volume\n" if $type ne 'volume';
+ my $row = $inv->{volumes}->{$dataset};
+ die "ZFS volume '$dataset' is not owned by NVMe subsystem '$nqn'\n" if $row->{nqn} ne $nqn;
+ die "invalid NSID on '$dataset'\n" if !nvmet_valid_nsid($row->{nsid});
+ die "invalid namespace UUID on '$dataset'\n" if $row->{uuid} !~ $RE_NVMET_UUID;
+ for my $other (keys $inv->{volumes}->%*) {
+ next if $other eq $dataset;
+ my $identity = $inv->{volumes}->{$other};
+ next if $identity->{nqn} ne $nqn;
+ die "duplicate NVMe identity on '$dataset'\n"
+ if $identity->{nsid} eq $row->{nsid} || $identity->{uuid} eq $row->{uuid};
+ }
+ return ($row->{nsid}, $row->{uuid});
+}
+
+# { nsid => [uuid, device] } for every volume owned by $nqn, after validating
+# all of them, so a bad identity refuses the whole plan before any change.
+sub _nvmet_desired_namespaces($inv, $nqn) {
+ my (%desired, %uuids);
+ for my $dataset (sort keys $inv->{volumes}->%*) {
+ my $row = $inv->{volumes}->{$dataset};
+ next if $row->{nqn} ne $nqn;
+ die "invalid NSID on '$dataset'\n" if !nvmet_valid_nsid($row->{nsid});
+ die "invalid namespace UUID on '$dataset'\n" if $row->{uuid} !~ $RE_NVMET_UUID;
+ die "duplicate NSID '$row->{nsid}'\n" if $desired{ $row->{nsid} };
+ die "duplicate namespace UUID '$row->{uuid}'\n" if $uuids{ lc($row->{uuid}) }++;
+ $desired{ $row->{nsid} } = [$row->{uuid}, "/dev/zvol/$dataset"];
+ }
+ return \%desired;
+}
+
+# NSIDs grow monotonically, also past the last NSID handed out in the pool,
+# so a host that missed a namespace removal never sees a new volume at an old
+# NSID. Only after the last NSID is used does the allocation fall back to the
+# lowest free one.
+sub _nvmet_allocate_nsid($inv, $cfs, $nqn) {
+ my %used;
+ for my $row (values $inv->{volumes}->%*) {
+ $used{ $row->{nsid} } = 1 if $row->{nqn} eq $nqn && nvmet_valid_nsid($row->{nsid});
+ }
+ $used{$_} = 1 for keys nvmet_namespaces($cfs, $nqn)->%*;
+
+ my $highest = max(0, $inv->{last_nsid} // 0, keys %used);
+ return $highest + 1 if $highest < $nvmet_max_nsid;
+ my $nsid = 1;
+ $nsid++ while $used{$nsid};
+ die "no free namespace ID\n" if $nsid > $nvmet_max_nsid;
+ return $nsid;
+}
+
+# Export state of (NSID, UUID, device) in the subsystem: present, disabled,
+# absent, incomplete (a disabled namespace with another identity, for example
+# one being built), stale (enabled on the other template name of the zvol) or
+# nosubsys. Dies on a conflict.
+sub _nvmet_export_state($cfs, $nqn, $nsid, $uuid, $dev) {
+ return 'nosubsys' if !$cfs->{subsystems}->{$nqn};
+ my $namespaces = nvmet_namespaces($cfs, $nqn);
+ for my $id (keys $namespaces->%*) {
+ next if $id eq $nsid;
+ die "namespace UUID '$uuid' is already in use\n"
+ if $namespaces->{$id}->{device_uuid} eq $uuid;
+ }
+ my $ns = $namespaces->{$nsid} // return 'absent';
+ if ($ns->{device_uuid} eq $uuid && $ns->{device_path} eq $dev) {
+ return $ns->{enable} eq '1' ? 'present' : 'disabled';
+ }
+ # A template conversion or its undo that ssh gave up on can rename the
+ # zvol on the target after an activation exported it under its old name.
+ my $twin = nvmet_template_twin($dev);
+ return 'stale'
+ if $ns->{enable} eq '1'
+ && $ns->{device_uuid} eq $uuid
+ && $twin ne ''
+ && $ns->{device_path} eq $twin;
+ die "NVMe namespace ID '$nsid' has a different identity; refusing to replace it\n"
+ if $ns->{enable} ne '0';
+ return 'incomplete';
+}
+
+# Units that export (NSID, UUID, device) and the state they start from:
+# absent, present, disabled, reclaim or stale. A disabled namespace with
+# another identity, or a stale one, is rebuilt, which is only correct under
+# the target lock. A stale namespace belongs to a template, or to a volume
+# whose template conversion failed, so no guest uses it.
+sub _nvmet_plan_export($cfs, $nqn, $nsid, $uuid, $dev) {
+ die "NVMe subsystem does not exist\n" if !$cfs->{subsystems}->{$nqn};
+ my $state = _nvmet_export_state($cfs, $nqn, $nsid, $uuid, $dev);
+ return ([nvmet_build_unit($nqn, $nsid, $uuid, $dev)], 'absent') if $state eq 'absent';
+ return ([], 'present') if $state eq 'present';
+ return ([nvmet_enable_unit($nqn, $nsid, $dev)], 'disabled') if $state eq 'disabled';
+ my $unexport = nvmet_unexport_unit($nqn, $nsid, $state eq 'stale');
+ my $build = nvmet_build_unit($nqn, $nsid, $uuid, $dev);
+ return ([[$unexport->@*, $build->@*]], $state eq 'stale' ? 'stale' : 'reclaim');
+}
+
+# Units that remove the namespace exporting $uuid, and its previous state.
+sub _nvmet_plan_unexport($cfs, $nqn, $uuid) {
+ my $namespaces = nvmet_namespaces($cfs, $nqn);
+ my @matches = grep { $namespaces->{$_}->{device_uuid} eq $uuid }
+ sort { $a <=> $b } keys $namespaces->%*;
+ die "duplicate namespace UUID '$uuid'\n" if @matches > 1;
+ return ([], undef) if !@matches;
+ my $enabled = $namespaces->{ $matches[0] }->{enable} eq '1';
+ return (
+ [nvmet_unexport_unit($nqn, $matches[0], $enabled)],
+ { nsid => $matches[0], enabled => $enabled },
+ );
+}
+
+# The port serving a portal ({ family, address, port } of parse_nvme_portals).
+# Among complete, matching ports, the one already linked to $nqn wins, then
+# the lowest id. Incomplete ports never match.
+sub _nvmet_find_port($cfs, $nqn, $portal) {
+ my $wanted = nvmet_canonical_address($portal->{address});
+ my @matches = grep {
+ my $port = $cfs->{ports}->{$_};
+ $port->{addr_trtype} eq 'tcp'
+ && $port->{addr_adrfam} eq $portal->{family}
+ && $port->{addr_traddr} ne ''
+ && $port->{addr_trsvcid} eq "$portal->{port}"
+ && nvmet_canonical_address($port->{addr_traddr}) eq $wanted
+ } sort {
+ $a <=> $b
+ } keys $cfs->{ports}->%*;
+ my @linked = grep { nvmet_port_linked($cfs, $_, $nqn) } @matches;
+ return $linked[0] // $matches[0];
+}
+
+# Units that bring the target to the configured state, except publishing:
+# objects (subsystem, ports, hosts), keys, ACLs, namespaces in NSID order,
+# then removal of undesired namespaces.
+# $conf: { nqn, portals (of parse_nvme_portals), hostnqns, keysha }
+sub _nvmet_plan_activation($inv, $cfs, $conf) {
+ my ($nqn, $keysha) = $conf->@{qw(nqn keysha)};
+ my $desired = _nvmet_desired_namespaces($inv, $nqn); # refuse before any unit
+ my $subsys_path = nvmet_subsys_path($nqn);
+ my $subsys = $cfs->{subsystems}->{$nqn};
+ my (@objects, @key_hosts, @acls, @namespaces, %port_ids);
+
+ if (!$subsys) {
+ push @objects,
+ [
+ ['mkdir', $subsys_path],
+ nvmet_write("$subsys_path/attr_model", $nvmet_model),
+ nvmet_write("$subsys_path/attr_serial", _nvmet_serial($nqn)),
+ nvmet_write("$subsys_path/attr_allow_any_host", 0),
+ ];
+ } else {
+ my $serial = _nvmet_serial($nqn);
+ my @attrs;
+ if ($subsys->{attr_model} ne $nvmet_model || $subsys->{attr_serial} ne $serial) {
+ die "refusing to take over existing NVMe subsystem '$nqn': it was not created by"
+ . " Proxmox VE; remove it on the target or use another NQN\n"
+ if !nvmet_unfinished_subsystem($cfs, $nqn);
+ push @attrs, nvmet_write("$subsys_path/attr_serial", $serial);
+ }
+ push @attrs, nvmet_write("$subsys_path/attr_allow_any_host", 0)
+ if $subsys->{attr_allow_any_host} ne '0';
+ push @objects, [@attrs] if @attrs;
+ }
+
+ my %taken = map { $_ => 1 } keys $cfs->{ports}->%*;
+ for my $portal ($conf->{portals}->@*) {
+ my $id = _nvmet_find_port($cfs, $nqn, $portal);
+ if (!defined($id)) {
+ $id = 1;
+ $id++ while $taken{$id};
+ $taken{$id} = 1;
+ my $port_path = nvmet_port_path($id);
+ push @objects,
+ [
+ ['mkdir', $port_path],
+ nvmet_write("$port_path/addr_trtype", 'tcp'),
+ nvmet_write("$port_path/addr_adrfam", $portal->{family}),
+ nvmet_write("$port_path/addr_traddr", $portal->{address}),
+ nvmet_write("$port_path/addr_trsvcid", $portal->{port}),
+ ];
+ }
+ $port_ids{$id} = 1;
+ }
+
+ # nvmet keeps the key on the global host object. A key that differs from
+ # ours is only replaced while no subsystem links the host.
+ my $linked = nvmet_linked_hosts($cfs);
+ for my $hostnqn ($conf->{hostnqns}->@*) {
+ my $host_path = nvmet_host_path($hostnqn);
+ if (my $host = $cfs->{hosts}->{$hostnqn}) {
+ my $sha = $host->{key_sha256}
+ // die "NVMe target does not expose a DH-HMAC-CHAP key for host '$hostnqn'\n";
+ if ($sha ne $keysha) {
+ die "refusing to replace an in-use DH-HMAC-CHAP key for '$hostnqn'\n"
+ if $linked->{$hostnqn};
+ push @key_hosts, $host_path;
+ }
+ } else {
+ push @objects, [['mkdir', $host_path]];
+ push @key_hosts, $host_path;
+ }
+ push @acls, [['ln', '-s', $host_path, "$subsys_path/allowed_hosts/$hostnqn"]]
+ if !$subsys || !($subsys->{acl} // {})->{$hostnqn};
+ }
+ # nvmet creates the key attributes world-readable. The chmod before the
+ # write keeps every later open from reading the key; a descriptor opened
+ # while an attribute was still world-readable is not affected by it.
+ my @keys;
+ if (@key_hosts) {
+ my @attrs = map { ("$_/dhchap_key", "$_/dhchap_ctrl_key") } @key_hosts;
+ @keys = ([['chmod', '0600', @attrs], { key => [map { "$_/dhchap_key" } @key_hosts] }]);
+ }
+
+ my $current = { subsystems => { $nqn => $subsys // { acl => {}, namespaces => {} } } };
+ for my $nsid (sort { $a <=> $b } keys $desired->%*) {
+ my ($units) = _nvmet_plan_export($current, $nqn, $nsid, $desired->{$nsid}->@*);
+ push @namespaces, $units->@*;
+ }
+ my $existing = nvmet_namespaces($cfs, $nqn);
+ for my $nsid (sort { $a <=> $b } keys $existing->%*) {
+ next if $desired->{$nsid};
+ push @namespaces, nvmet_unexport_unit($nqn, $nsid, $existing->{$nsid}->{enable} eq '1');
+ }
+
+ return {
+ prepublish => [@objects, @keys, @acls, @namespaces],
+ port_ids => [sort { $a <=> $b } keys %port_ids],
+ };
+}
+
+# One call per port: each link repeats the guards, so it only succeeds while
+# the subsystem denies unknown hosts and every configured host has its ACL.
+# Under the target lock, only a manager of the target outside this cluster
+# can change that after the verify read; the guards stop the publish then.
+# Links to ports that are no longer configured are removed.
+# Returns [{ port => $id, action => 'link' | 'unlink', steps => [...] }, ...].
+sub _nvmet_plan_publish($cfs, $nqn, $port_ids, $hostnqns) {
+ my $subsys_path = nvmet_subsys_path($nqn);
+ my %wanted = map { $_ => 1 } $port_ids->@*;
+ my @guards = (
+ ['grep', '-qx', '0', "$subsys_path/attr_allow_any_host"],
+ map { ['test', '-L', "$subsys_path/allowed_hosts/$_"] } $hostnqns->@*,
+ );
+
+ my @units;
+ for my $id ($port_ids->@*) {
+ next if nvmet_port_linked($cfs, $id, $nqn);
+ my $link = ['ln', '-s', $subsys_path, nvmet_port_path($id) . "/subsystems/$nqn"];
+ push @units, { port => $id, action => 'link', steps => [@guards, $link] };
+ }
+ for my $id (sort { $a <=> $b } keys $cfs->{ports}->%*) {
+ next if $wanted{$id} || !nvmet_port_linked($cfs, $id, $nqn);
+ my $unlink = ['rm', nvmet_port_path($id) . "/subsystems/$nqn"];
+ push @units, { port => $id, action => 'unlink', steps => [$unlink] };
+ }
+ return \@units;
+}
+
+# Teardown units for the subsystem and the hosts that may become orphans.
+sub _nvmet_plan_delete_target($inv, $cfs, $nqn, $hostnqns) {
+ for my $dataset (sort keys $inv->{volumes}->%*) {
+ die "refusing to delete NVMe subsystem '$nqn': owned ZFS volume '$dataset' exists\n"
+ if $inv->{volumes}->{$dataset}->{nqn} eq $nqn;
+ }
+ my $subsys = $cfs->{subsystems}->{$nqn} // return ([], [$hostnqns->@*]);
+ die "refusing to delete foreign NVMe subsystem '$nqn'\n"
+ if $subsys->{attr_model} ne $nvmet_model || $subsys->{attr_serial} ne _nvmet_serial($nqn);
+ die "refusing to delete NVMe subsystem '$nqn': namespaces remain\n"
+ if nvmet_namespaces($cfs, $nqn)->%*;
+
+ my $subsys_path = nvmet_subsys_path($nqn);
+ my @acl = sort keys(($subsys->{acl} // {})->%*);
+ die "refusing to delete NVMe subsystem '$nqn': it allows a malformed host name\n"
+ if grep { $_ !~ $RE_NQN } @acl;
+ my @links = (
+ (
+ map { nvmet_port_path($_) . "/subsystems/$nqn" }
+ grep { nvmet_port_linked($cfs, $_, $nqn) }
+ sort { $a <=> $b } keys $cfs->{ports}->%*
+ ),
+ (map { "$subsys_path/allowed_hosts/$_" } @acl),
+ );
+ # A retry after a teardown that failed at its rmdir finds no ACL left, so
+ # the configured hosts are candidates as well.
+ my %candidates = map { $_ => 1 } @acl, $hostnqns->@*;
+ return (
+ [[(@links ? (['rm', @links]) : ()), ['rmdir', $subsys_path]]], [sort keys %candidates],
+ );
+}
+
+# Candidate hosts that exist and that no subsystem links any more.
+sub _nvmet_plan_orphan_hosts($cfs, $candidates) {
+ my $linked = nvmet_linked_hosts($cfs);
+ my %seen;
+ my @orphans = map { nvmet_host_path($_) }
+ grep { !$seen{$_}++ && $_ =~ $RE_NQN && $cfs->{hosts}->{$_} && !$linked->{$_} }
+ sort $candidates->@*;
+ return @orphans ? [[['rmdir', @orphans]]] : [];
+}
+
+# ---------------------------------------------------------------------------
+# Target lock and remote execution
+# ---------------------------------------------------------------------------
+
+my %nvmet_lock_owner; # lock id => pid of the process holding the domain lock
+
+# Runs $code under the pmxcfs domain lock of the target. Like the pmxcfs
+# lock itself, it is not re-entrant, and a child forked inside is not the
+# owner.
+my sub nvmet_locked($scfg, $code) {
+ my $id = nvmet_lock_id($scfg);
+ die "cluster not quorate - refusing NVMe target changes\n"
+ if !PVE::Cluster::check_cfs_quorum(1);
+ my $res = PVE::Cluster::cfs_lock_domain(
+ $id,
+ $nvmet_lock_wait,
+ sub {
+ local $nvmet_lock_owner{$id} = $$;
+ return $code->();
+ },
+ );
+ die $@ if $@;
+ return $res;
+}
+
+# Runs steps on the target. %opts: op, timeout, input, change (the steps
+# change the target, which needs the lock).
+#
+# ssh does not stop a command on the target when it gives up on it: a change
+# whose connection broke (exit code 255), any call under the lock that ssh
+# did not run to its end (-1: its timeout, or ssh was killed), and a call
+# during which the task was stopped (it dies with the task marker) may still
+# be running there. Its outcome is unknown, so it is never compensated from a
+# read that could come before the rest of it. The next activation repairs
+# from the target state.
+my sub nvmet_exec($scfg, $steps, %opts) {
+ my $locked = ($nvmet_lock_owner{ nvmet_lock_id($scfg) } // 0) == $$;
+ die "internal error: NVMe target change without target lock\n"
+ if $opts{change} && !$locked;
+ my $res = _nvmet_run(
+ $scfg, $steps,
+ op => $opts{op},
+ timeout => $opts{timeout} // 15,
+ (defined($opts{input}) ? (input => $opts{input}) : ()),
+ );
+ die "NVMe target operation '$opts{op}' did not complete ($res->{err});"
+ . " the target state is unknown\n"
+ if ($locked && $res->{rc} == -1)
+ || ($opts{change} && $res->{rc} == 255);
+ return $res;
+}
+
+# An abandoned operation is passed on instead of being compensated.
+my sub nvmet_rethrow_abandoned($error) {
+ die $error if $error =~ $RE_NVMET_ABANDONED;
+}
+
+my sub nvmet_unreachable($res) {
+ return $res->{rc} == 255 || $res->{rc} == -1;
+}
+
+my sub nvmet_check($res, $op) {
+ die "NVMe target operation '$op' failed: $res->{err}\n" if $res->{rc};
+}
+
+my sub nvmet_read_failed($scfg, $res) {
+ die "NVMe target '" . nvmet_server($scfg) . "' is unreachable: $res->{err}\n"
+ if nvmet_unreachable($res);
+ die "cannot read NVMe target state: $res->{err}\n";
+}
+
+my sub nvmet_output($res) {
+ return join('', map { "$_\n" } $res->{out}->@*);
+}
+
+# A read, retried twice after 200 ms unless the target is unreachable: a read
+# without the lock can meet a configfs object that another node removes
+# while find lists it, which makes find fail once.
+my sub nvmet_read_call($scfg, $steps, %opts) {
+ my $res;
+ for my $attempt (0 .. 2) {
+ _sleep(0.2) if $attempt;
+ $res = nvmet_exec($scfg, $steps, %opts, timeout => 15);
+ last if !$res->{rc} || nvmet_unreachable($res);
+ }
+ return $res;
+}
+
+# Reads the target state and returns (ZFS inventory, configfs state).
+# keys => [hostnqns] adds the digests of their keys. modprobe => 1, for locked
+# reads only, loads nvmet first and mounts configfs if that is what failed.
+# With tolerant => 1, a failed or unparsable read returns an empty list,
+# unless the target is unreachable.
+my sub nvmet_read($scfg, %opts) {
+ my $pool = nvmet_pool($scfg);
+ my $hostnqns = $opts{keys} // [];
+ # loading a module that may still be loading is harmless
+ my @modprobe = $opts{modprobe} ? (change => 1) : ();
+ my $read = sub () {
+ return nvmet_read_call(
+ $scfg,
+ nvmet_state_steps($pool, $hostnqns, $opts{modprobe}),
+ op => 'read target state',
+ @modprobe,
+ );
+ };
+ my $res = $read->();
+ if ($res->{rc} && $opts{modprobe} && !nvmet_unreachable($res)) {
+ my $mounts = nvmet_exec($scfg, [['cat', '/proc/mounts']], op => 'read mounts');
+ nvmet_check($mounts, 'read mounts');
+ if (!_nvmet_parse_mounts(nvmet_output($mounts))) {
+ my $mount = ['mount', '-t', 'configfs', 'none', '/sys/kernel/config'];
+ my $mounted = nvmet_exec($scfg, [$mount], op => 'mount configfs', change => 1);
+ nvmet_check($mounted, 'mount configfs');
+ $res = $read->();
+ }
+ }
+ if ($res->{rc}) {
+ return () if $opts{tolerant} && !nvmet_unreachable($res);
+ nvmet_read_failed($scfg, $res);
+ }
+ my @state = eval {
+ my ($zfs, $configfs) = _nvmet_split_state(nvmet_output($res));
+ (_nvmet_parse_zfs_inventory($zfs, $pool), _nvmet_parse_configfs($configfs));
+ };
+ if (my $error = $@) {
+ return () if $opts{tolerant};
+ die $error;
+ }
+ return @state;
+}
+
+my sub nvmet_read_zfs($scfg) {
+ my $pool = nvmet_pool($scfg);
+ my $res = nvmet_read_call($scfg, [nvmet_zfs_inventory_step($pool)], op => 'read ZFS inventory');
+ nvmet_read_failed($scfg, $res) if $res->{rc};
+ return _nvmet_parse_zfs_inventory(nvmet_output($res), $pool);
+}
+
+# Applies units in as few calls as possible and stops at the first failing
+# call, whose result it returns. A chain only saves SSH round trips: every
+# decision was taken locally before, and `&&` only stops at the first
+# failing command. Only the call with the key step gets stdin. The default
+# timeout suits configfs chains; ZFS changes pass their own.
+my sub nvmet_apply($scfg, $units, %opts) {
+ my $input = delete $opts{input};
+ $opts{timeout} //= 30;
+ for my $chunk (_nvmet_chunk($units)->@*) {
+ my $with_key = grep { ref($_) eq 'HASH' && exists($_->{key}) } $chunk->@*;
+ my @input = $with_key ? (input => $input) : ();
+ my $res = nvmet_exec($scfg, $chunk, %opts, @input, change => 1);
+ return $res if $res->{rc};
+ }
+ return { rc => 0, out => [], err => '' };
+}
+
+my sub nvmet_change($scfg, $units, %opts) {
+ nvmet_check(nvmet_apply($scfg, $units, %opts), $opts{op});
+}
+
+my sub nvmet_wait_device($scfg, $dev) {
+ my $deadline = _now() + 10;
+ while (1) {
+ my $res = nvmet_exec($scfg, [['test', '-b', $dev]], op => 'wait for zvol');
+ return if !$res->{rc};
+ nvmet_read_failed($scfg, $res) if nvmet_unreachable($res);
+ last if _now() >= $deadline;
+ _sleep(0.25);
+ }
+ die "zvol '$dev' is not a block device after 10 seconds\n";
+}
+
+my sub nvmet_new_uuid($inv, $cfs) {
+ my %used = map { lc($_->{uuid}) => 1 } values $inv->{volumes}->%*;
+ for my $subsys (values $cfs->{subsystems}->%*) {
+ $used{ lc($_->{device_uuid}) } = 1 for values(($subsys->{namespaces} // {})->%*);
+ }
+ for (1 .. 10) {
+ my $uuid = lc(file_read_firstline('/proc/sys/kernel/random/uuid') // '');
+ return $uuid if $uuid =~ $RE_NVMET_UUID && !$used{$uuid};
+ }
+ die "cannot generate an unused namespace UUID\n";
+}
+
+my sub nvmet_has_identity($row, $nqn, $nsid, $uuid) {
+ return $row && $row->{nqn} eq $nqn && $row->{nsid} eq "$nsid" && $row->{uuid} eq $uuid;
+}
+
+# The configfs state after a successful unexport of namespace $pre->{nsid}.
+my sub nvmet_without_namespace($cfs, $nqn, $pre) {
+ return $cfs if !$pre;
+ my $subsys = $cfs->{subsystems}->{$nqn};
+ my %namespaces = $subsys->{namespaces}->%*;
+ delete $namespaces{ $pre->{nsid} };
+ my $changed = { $subsys->%*, namespaces => \%namespaces };
+ return { $cfs->%*, subsystems => { $cfs->{subsystems}->%*, $nqn => $changed } };
+}
+
+# Links the subsystem to each desired port separately. One port that cannot
+# be bound (for example, its address is missing on the target) only warns
+# while another desired port serves the subsystem.
+my sub nvmet_publish($scfg, $cfs, $nqn, $port_ids, $hostnqns) {
+ my @failed;
+ for my $unit (_nvmet_plan_publish($cfs, $nqn, $port_ids, $hostnqns)->@*) {
+ my $op = "$unit->{action} NVMe subsystem on port $unit->{port}";
+ my $res = nvmet_apply($scfg, [$unit->{steps}], op => $op);
+ push @failed, { $unit->%*, err => $res->{err} } if $res->{rc};
+ }
+ return if !@failed;
+
+ my (undef, $fresh) = eval { nvmet_read($scfg) };
+ nvmet_rethrow_abandoned($@) if $@;
+ my $published = $fresh && grep { nvmet_port_linked($fresh, $_, $nqn) } $port_ids->@*;
+ my ($first) = grep { $_->{action} eq 'link' } @failed;
+ $first //= $failed[0];
+ die "cannot publish NVMe subsystem '$nqn': $first->{err}\n" if !$published;
+ for my $failure (@failed) {
+ my $what = $failure->{action} eq 'link' ? 'publish' : 'unpublish';
+ log_warn("cannot $what NVMe subsystem '$nqn' on port $failure->{port}: $failure->{err}");
+ }
+}
+
+# ---------------------------------------------------------------------------
+# Target flows
+# ---------------------------------------------------------------------------
+
+# Brings the target to the configured state and publishes the subsystem.
+# A converged target costs one read and no lock. Otherwise, under the lock:
+# read, apply, verify with a fresh read, and publish only once the verify
+# read plans nothing more. $portals and $hostnqns are the parsed and validated
+# storage configuration.
+sub _nvmet_activate_target($storeid, $scfg, $portals, $hostnqns, $key) {
+ my $nqn = nvmet_nqn($scfg);
+ my $conf = {
+ nqn => $nqn,
+ portals => $portals,
+ hostnqns => $hostnqns,
+ keysha => sha256_hex("$key\n"),
+ };
+ my $converged = sub($inv, $cfs) {
+ my $plan = _nvmet_plan_activation($inv, $cfs, $conf);
+ return !$plan->{prepublish}->@*
+ && !_nvmet_plan_publish($cfs, $nqn, $plan->{port_ids}, $hostnqns)->@*;
+ };
+
+ # Without the lock, a refusal or failed read only means "take the lock".
+ my @state = nvmet_read($scfg, keys => $hostnqns, tolerant => 1);
+ return if @state && eval { $converged->(@state) };
+
+ nvmet_locked(
+ $scfg,
+ sub {
+ # Removing the storage deletes its key before it removes the target
+ # configuration, so an activation that waited for the lock meanwhile
+ # does not restore it.
+ die "storage '$storeid' is being removed or its key changed\n"
+ if (file_read_firstline(secret_path($storeid)) // '') ne $key;
+ my ($inv, $cfs) = nvmet_read($scfg, keys => $hostnqns, modprobe => 1);
+ my $plan = _nvmet_plan_activation($inv, $cfs, $conf);
+ my $error;
+ for my $round (1 .. 2) {
+ last if !$plan->{prepublish}->@*;
+ my $apply = nvmet_apply(
+ $scfg, $plan->{prepublish},
+ op => 'configure NVMe target',
+ input => "$key\n",
+ );
+ $error //= $apply->{err} if $apply->{rc};
+ ($inv, $cfs) = nvmet_read($scfg, keys => $hostnqns);
+ $plan = _nvmet_plan_activation($inv, $cfs, $conf);
+ }
+ die "NVMe target did not converge: " . ($error // 'state mismatch') . "\n"
+ if $plan->{prepublish}->@*;
+ nvmet_publish($scfg, $cfs, $nqn, $plan->{port_ids}, $hostnqns);
+ },
+ );
+ return;
+}
+
+# The namespace UUID of an owned volume, from one read-only ZFS query.
+sub _nvmet_volume_uuid($scfg, $name) {
+ my $dataset = nvmet_dataset($scfg, $name);
+ my $inv = nvmet_read_zfs($scfg);
+ my (undef, $uuid) = _nvmet_owned_identity($inv, nvmet_nqn($scfg), $dataset);
+ return $uuid;
+}
+
+# Creates (size => KiB) or clones (origin => pool/base@snap) a zvol with its
+# identity set at creation, then exports it. A failure removes only what this
+# call created, decided from a fresh read.
+my sub nvmet_create_volume($scfg, $name, %opts) {
+ my $nqn = nvmet_nqn($scfg);
+ my $dataset = nvmet_dataset($scfg, $name);
+ my $dev = "/dev/zvol/$dataset";
+ die "internal error: invalid zvol creation request\n"
+ if defined($opts{size}) == defined($opts{origin})
+ || (defined($opts{size}) && $opts{size} !~ $RE_UNSIGNED_INTEGER);
+
+ return nvmet_locked(
+ $scfg,
+ sub {
+ my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1);
+ die "zvol '$dataset' already exists\n" if exists($inv->{types}->{$dataset});
+ die "NVMe subsystem does not exist; activate the storage first\n"
+ if !$cfs->{subsystems}->{$nqn};
+ my $nsid = _nvmet_allocate_nsid($inv, $cfs, $nqn);
+ my $uuid = nvmet_new_uuid($inv, $cfs);
+ my (undef, $state) = _nvmet_plan_export($cfs, $nqn, $nsid, $uuid, $dev);
+ die "internal error: namespace ID '$nsid' is in use\n" if $state ne 'absent';
+
+ # Record the NSID on the pool before the zvol exists: a creation that
+ # completes on the target only after this call gave up on it must
+ # never share its NSID with a later volume.
+ my $timeout = nvmet_long_timeout();
+ if ($nsid > $inv->{last_nsid}) {
+ my $reserve = ['zfs', 'set', "$nvmet_last_nsid=$nsid", nvmet_pool($scfg)];
+ nvmet_change(
+ $scfg, [[$reserve]],
+ op => 'reserve namespace ID',
+ timeout => $timeout,
+ );
+ }
+
+ my @identity = map { ('-o', $_) } nvmet_identity($nqn, $nsid, $uuid);
+ my @create_opts = (
+ ($scfg->{sparse} ? ('-s') : ()),
+ ($scfg->{blocksize} ? ('-b', $scfg->{blocksize}) : ()),
+ );
+ my $create =
+ defined($opts{origin})
+ ? ['zfs', 'clone', @identity, $opts{origin}, $dataset]
+ : ['zfs', 'create', @create_opts, @identity, '-V', "$opts{size}k", $dataset];
+ my $res = nvmet_apply(
+ $scfg, [[$create]],
+ op => 'create zvol',
+ timeout => $timeout,
+ );
+ if ($res->{rc}) {
+ # A definite command failure may still have created the zvol.
+ # Only a zvol with this call's identity counts. Ambiguous SSH
+ # outcomes have already aborted the transaction in nvmet_exec.
+ my $check = eval { nvmet_read_zfs($scfg) };
+ nvmet_rethrow_abandoned($@) if $@;
+ my $row = $check ? $check->{volumes}->{$dataset} : undef;
+ die "NVMe target operation 'create zvol' failed: $res->{err}\n"
+ if !nvmet_has_identity($row, $nqn, $nsid, $uuid);
+ }
+
+ my $verified;
+ eval {
+ nvmet_wait_device($scfg, $dev);
+ my $build_unit = nvmet_build_unit($nqn, $nsid, $uuid, $dev);
+ my $build = nvmet_apply($scfg, [$build_unit], op => 'export namespace');
+ $verified = [nvmet_read($scfg)];
+ my ($vinv, $vcfs) = $verified->@*;
+ my ($vnsid, $vuuid) = _nvmet_owned_identity($vinv, $nqn, $dataset);
+ die "NVMe identity of zvol '$dataset' changed\n"
+ if $vnsid ne $nsid || $vuuid ne $uuid;
+ if (_nvmet_export_state($vcfs, $nqn, $nsid, $uuid, $dev) ne 'present') {
+ nvmet_check($build, 'export namespace');
+ die "NVMe namespace '$nsid' was not exported\n";
+ }
+ };
+ my $error = $@ or return $uuid;
+ nvmet_rethrow_abandoned($error);
+
+ eval {
+ my ($cinv, $ccfs) = $verified ? $verified->@* : nvmet_read($scfg);
+ my $namespaces = nvmet_namespaces($ccfs, $nqn);
+ # The build writes the UUID before it enables the namespace, so an
+ # enabled namespace without our UUID was not built by this call.
+ my @unexport =
+ map { nvmet_unexport_unit($nqn, $_, $namespaces->{$_}->{enable} eq '1') }
+ grep {
+ my $ns = $namespaces->{$_};
+ $ns->{device_uuid} eq $uuid
+ || ($ns->{enable} ne '1' && $ns->{device_path} eq $dev)
+ } sort {
+ $a <=> $b
+ } keys $namespaces->%*;
+ nvmet_change($scfg, \@unexport, op => 'remove namespace') if @unexport;
+ # -o at creation proves that this call created the zvol
+ if (nvmet_has_identity($cinv->{volumes}->{$dataset}, $nqn, $nsid, $uuid)) {
+ my $destroy = ['zfs', 'destroy', '-r', $dataset];
+ nvmet_change(
+ $scfg, [[$destroy]],
+ op => 'destroy zvol',
+ timeout => $timeout,
+ );
+ }
+ };
+ log_warn("failed to clean up zvol '$name': $@") if $@;
+ die $error;
+ },
+ );
+}
+
+# Unexports and destroys an owned zvol. A missing zvol is already destroyed.
+my sub nvmet_destroy_volume($scfg, $name) {
+ my $nqn = nvmet_nqn($scfg);
+ my $dataset = nvmet_dataset($scfg, $name);
+ my $dev = "/dev/zvol/$dataset";
+ my $destroy = ['zfs', 'destroy', '-r', $dataset];
+ my $timeout = nvmet_long_timeout();
+
+ nvmet_locked(
+ $scfg,
+ sub {
+ my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1);
+ return if !exists($inv->{types}->{$dataset});
+ my ($nsid, $uuid) = _nvmet_owned_identity($inv, $nqn, $dataset);
+ my ($unexport, $pre) = ([], undef);
+ if ($cfs->{subsystems}->{$nqn}) {
+ _nvmet_plan_export($cfs, $nqn, $nsid, $uuid, $dev); # refusals only
+ ($unexport, $pre) = _nvmet_plan_unexport($cfs, $nqn, $uuid);
+ }
+ my $units = [$unexport->@*, [$destroy]];
+ my $res = nvmet_apply($scfg, $units, op => 'destroy zvol', timeout => $timeout);
+ return if !$res->{rc};
+ my $error = $res->{err};
+
+ # Decide from a fresh read what the failed call changed.
+ my ($finv, $fcfs) = eval { nvmet_read($scfg) };
+ if (my $read_error = $@) {
+ nvmet_rethrow_abandoned($read_error);
+ die "failed to destroy '$dataset': $error\n";
+ }
+ return if !exists($finv->{types}->{$dataset});
+ my $namespaces = nvmet_namespaces($fcfs, $nqn);
+ my ($current) =
+ grep { $namespaces->{$_}->{device_uuid} eq $uuid } keys $namespaces->%*;
+ if (defined($current)) {
+ die "cannot disable namespace '$current': $error\n"
+ if $namespaces->{$current}->{enable} eq '1';
+ die "cannot remove namespace '$current': $error\n" if !$pre || !$pre->{enabled};
+ my $enable = [nvmet_enable_unit($nqn, $current, $dev)];
+ my $restore = nvmet_apply($scfg, $enable, op => 'enable namespace');
+ die "cannot remove namespace '$current': $error"
+ . ($restore->{rc} ? "; cannot restore namespace: $restore->{err}" : '') . "\n";
+ }
+
+ # The namespace is gone and the zvol is still there, typically busy.
+ for my $attempt (1 .. 5) {
+ _sleep(1);
+ my $retry =
+ nvmet_apply($scfg, [[$destroy]], op => 'destroy zvol', timeout => $timeout);
+ return if !$retry->{rc};
+ $error = $retry->{err};
+ }
+ die "failed to destroy '$dataset': $error\n" if !$pre;
+ eval {
+ my (undef, $rcfs) = nvmet_read($scfg);
+ my ($units) = _nvmet_plan_export($rcfs, $nqn, $nsid, $uuid, $dev);
+ nvmet_change($scfg, $units, op => 'restore namespace');
+ };
+ die "failed to destroy '$dataset': $error; namespace restoration also failed: $@"
+ if $@;
+ die "failed to destroy '$dataset': $error; namespace restored\n";
+ },
+ );
+ return;
+}
+
+# Rolls an owned zvol back and always exports it again with the same identity.
+my sub nvmet_rollback_volume($scfg, $name, $snap) {
+ my $snapshot = nvmet_snapshot($scfg, $name, $snap);
+ my $nqn = nvmet_nqn($scfg);
+ my $dataset = nvmet_dataset($scfg, $name);
+ my $dev = "/dev/zvol/$dataset";
+
+ my $timeout = nvmet_long_timeout();
+
+ nvmet_locked(
+ $scfg,
+ sub {
+ my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1);
+ my ($nsid, $uuid) = _nvmet_owned_identity($inv, $nqn, $dataset);
+ die "NVMe subsystem does not exist\n" if !$cfs->{subsystems}->{$nqn};
+ _nvmet_plan_export($cfs, $nqn, $nsid, $uuid, $dev); # refusals only
+ my ($unexport, $pre) = _nvmet_plan_unexport($cfs, $nqn, $uuid);
+ my $set_identity = ['zfs', 'set', nvmet_identity($nqn, $nsid, $uuid), $dataset];
+ my $rollback = ['zfs', 'rollback', $snapshot];
+ my $res = nvmet_apply(
+ $scfg, [$unexport->@*, [$rollback], [$set_identity]],
+ op => 'rollback zvol',
+ timeout => $timeout,
+ );
+ my $error;
+ $error = "NVMe target operation 'rollback zvol' failed: $res->{err}" if $res->{rc};
+
+ # Always export the volume again, from a fresh read if anything failed.
+ eval {
+ my ($rinv, $rcfs) =
+ $error ? nvmet_read($scfg) : ($inv, nvmet_without_namespace($cfs, $nqn, $pre));
+ nvmet_wait_device($scfg, $dev);
+ my @units;
+ push @units, [$set_identity]
+ if !nvmet_has_identity($rinv->{volumes}->{$dataset}, $nqn, $nsid, $uuid);
+ my ($export) = _nvmet_plan_export($rcfs, $nqn, $nsid, $uuid, $dev);
+ push @units, $export->@*;
+ nvmet_change($scfg, \@units, op => 'restore namespace', timeout => $timeout);
+ };
+ chomp(my $restore_error = $@);
+
+ die "rollback failed: $error; namespace restoration failed: $restore_error\n"
+ if defined($error) && $restore_error;
+ die "namespace restoration failed after rollback: $restore_error\n"
+ if $restore_error;
+ die "$error\n" if defined($error);
+ },
+ );
+ return;
+}
+
+# Renames vm-* to base-*, exports it under the same identity and takes the
+# __base__ snapshot. A failed conversion is undone from a fresh read.
+my sub nvmet_template_volume($scfg, $name) {
+ my $nqn = nvmet_nqn($scfg);
+ my $old = nvmet_dataset($scfg, $name);
+ my $new = _nvmet_template_name($old);
+ my ($old_dev, $new_dev) = ("/dev/zvol/$old", "/dev/zvol/$new");
+ my $timeout = nvmet_long_timeout();
+
+ nvmet_locked(
+ $scfg,
+ sub {
+ my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1);
+ my ($nsid, $uuid) = _nvmet_owned_identity($inv, $nqn, $old);
+ die "template zvol '$new' already exists\n" if exists($inv->{types}->{$new});
+ die "NVMe subsystem does not exist\n" if !$cfs->{subsystems}->{$nqn};
+ _nvmet_plan_export($cfs, $nqn, $nsid, $uuid, $old_dev); # refusals only
+ my ($unexport) = _nvmet_plan_unexport($cfs, $nqn, $uuid);
+
+ eval {
+ my $rename = [$unexport->@*, [['zfs', 'rename', $old, $new]]];
+ nvmet_change($scfg, $rename, op => 'rename zvol', timeout => $timeout);
+ nvmet_wait_device($scfg, $new_dev);
+ my $build = nvmet_build_unit($nqn, $nsid, $uuid, $new_dev);
+ my $template = [$build, [['zfs', 'snapshot', "$new\@__base__"]]];
+ nvmet_change($scfg, $template, op => 'create template', timeout => $timeout);
+ };
+ my $error = $@ or return;
+ nvmet_rethrow_abandoned($error);
+ chomp($error);
+
+ eval {
+ my ($rinv, $rcfs) = nvmet_read($scfg);
+ if (exists($rinv->{types}->{$new})) {
+ my ($units) = _nvmet_plan_unexport($rcfs, $nqn, $uuid);
+ my $back = [$units->@*, [['zfs', 'rename', $new, $old]]];
+ nvmet_change($scfg, $back, op => 'rename zvol back', timeout => $timeout);
+ (undef, $rcfs) = nvmet_read($scfg);
+ } elsif (!exists($rinv->{types}->{$old})) {
+ die "zvol '$old' does not exist\n";
+ }
+ nvmet_wait_device($scfg, $old_dev);
+ my ($units) = _nvmet_plan_export($rcfs, $nqn, $nsid, $uuid, $old_dev);
+ nvmet_change($scfg, $units, op => 'restore namespace');
+ };
+ die "template conversion failed: $error; restoration failed: $@" if $@;
+ die "template conversion failed: $error\n";
+ },
+ );
+ return;
+}
+
+# Grows an owned zvol and lets an exported namespace pick up the new size.
+my sub nvmet_resize_volume($scfg, $name, $size) {
+ my $nqn = nvmet_nqn($scfg);
+ my $dataset = nvmet_dataset($scfg, $name);
+ die "internal error: invalid zvol size\n" if $size !~ $RE_UNSIGNED_INTEGER;
+
+ # The volume is checked before it changes; under the lock, the namespace
+ # read here is still the one to revalidate after the resize. It is
+ # revalidated in the same call, so whenever the zvol grows, also after a
+ # lost reply, the hosts see the new size. A namespace that is not enabled
+ # reads the new size when it is enabled.
+ nvmet_locked(
+ $scfg,
+ sub {
+ my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1);
+ my (undef, $uuid) = _nvmet_owned_identity($inv, $nqn, $dataset);
+ my $namespaces = nvmet_namespaces($cfs, $nqn);
+ my @matches =
+ grep { $namespaces->{$_}->{device_uuid} eq $uuid } keys $namespaces->%*;
+ die "duplicate namespace UUID '$uuid'\n" if @matches > 1;
+ my @resize = (['zfs', 'set', "volsize=${size}k", $dataset]);
+ push @resize, nvmet_write(nvmet_ns_path($nqn, $matches[0]) . '/revalidate_size', 1)
+ if @matches && $namespaces->{ $matches[0] }->{enable} eq '1';
+ nvmet_change(
+ $scfg, [\@resize],
+ op => 'resize zvol',
+ timeout => nvmet_long_timeout(),
+ );
+ },
+ );
+ return;
+}
+
+# Removes the subsystem once it has no volumes and namespaces, then the hosts
+# that no subsystem links any more. Ports are never removed.
+sub _nvmet_delete_target($scfg) {
+ my $nqn = nvmet_nqn($scfg);
+ my $hostnqns = parse_nvme_host_nqns($scfg->{'nvme-host-nqns'}, 1) // [];
+
+ nvmet_locked(
+ $scfg,
+ sub {
+ my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1);
+ my ($units, $candidates) = _nvmet_plan_delete_target($inv, $cfs, $nqn, $hostnqns);
+ my $fresh = $cfs;
+ if ($units->@*) {
+ my $res = nvmet_apply($scfg, $units, op => 'remove NVMe subsystem');
+ (undef, $fresh) = eval { nvmet_read($scfg) };
+ my $read_error = $@;
+ nvmet_rethrow_abandoned($read_error) if $read_error;
+ die "cannot remove NVMe subsystem '$nqn': $res->{err}\n"
+ if $res->{rc} && (!$fresh || $fresh->{subsystems}->{$nqn});
+ die $read_error if !$fresh;
+ }
+ my $orphans = _nvmet_plan_orphan_hosts($fresh, $candidates);
+ return if !$orphans->@*;
+ my $res = nvmet_apply($scfg, $orphans, op => 'remove orphan NVMe hosts');
+ log_warn("could not remove orphan NVMe host(s): $res->{err}") if $res->{rc};
+ },
+ );
+ return;
+}
+
+# ---------------------------------------------------------------------------
+# Configuration
+# ---------------------------------------------------------------------------
+
+sub verify_nvme_nqn($value, $noerr = undef) {
+
+ if (
+ length($value) > 223
+ || $value !~ $RE_NQN
+ ) {
+ return undef if $noerr;
+ die "value is not a valid NVMe qualified name\n";
+ }
+
+ return $value;
+}
+
+sub parse_nvme_portals($value, $noerr = undef) {
+ my $result = [];
+ my $seen = {};
+
+ for my $entry (split(/,/, $value // '')) {
+ $entry = trim($entry);
+ my ($address, $port, $family);
+ if ($entry =~ $RE_IPV6_PORTAL) {
+ ($address, $port, $family) = ($+{address}, $+{port} // 4420, 'ipv6');
+ } elsif ($entry =~ $RE_IPV4_PORTAL) {
+ ($address, $port, $family) = ($+{address}, $+{port} // 4420, 'ipv4');
+ } else {
+ return undef if $noerr;
+ die "invalid NVMe/TCP portal '$entry'\n";
+ }
+
+ if (
+ !PVE::JSONSchema::pve_verify_ip($address, 1)
+ || ($family eq 'ipv4' && index($address, ':') >= 0)
+ || ($family eq 'ipv6' && index($address, ':') < 0)
+ || $port < 1
+ || $port > 65535
+ ) {
+ return undef if $noerr;
+ die "invalid NVMe/TCP portal '$entry'\n";
+ }
+ $address = nvmet_canonical_address($address) if $family eq 'ipv6';
+ # one listener: compare like _assert_unique_target
+ my $id = join(
+ "\0",
+ nvmet_listener_address({ family => $family, address => $address }),
+ int($port),
+ );
+ if ($seen->{$id}++) {
+ return undef if $noerr;
+ die "duplicate NVMe/TCP portal '$entry'\n";
+ }
+ push $result->@*,
+ {
+ address => $address,
+ port => int($port),
+ family => $family,
+ };
+ if (scalar($result->@*) > $max_paths) {
+ return undef if $noerr;
+ die "at most $max_paths NVMe/TCP portals are supported\n";
+ }
+ }
+
+ if (!$result->@*) {
+ return undef if $noerr;
+ die "at least one NVMe/TCP portal is required\n";
+ }
+
+ return $result;
+}
+
+my sub verify_nvme_portals($value, $noerr = undef) {
+ return undef if !parse_nvme_portals($value, $noerr);
+ return $value;
+}
+
+sub parse_nvme_host_ifaces($value, $noerr = undef) {
+ my $result = [];
+
+ for my $iface (split(/,/, $value // '')) {
+ $iface = trim($iface);
+ # the kernel never names an interface '.' or '..'
+ if (
+ length($iface) < 1
+ || length($iface) > 15
+ || $iface !~ $RE_HOST_IFACE
+ || $iface eq '.'
+ || $iface eq '..'
+ ) {
+ return undef if $noerr;
+ die "invalid NVMe/TCP host interface '$iface'\n";
+ }
+ push $result->@*, $iface;
+ }
+
+ if (!$result->@*) {
+ return undef if $noerr;
+ die "at least one NVMe/TCP host interface is required\n";
+ }
+
+ return $result;
+}
+
+sub parse_nvme_host_nqns($value, $noerr = undef) {
+ my $result = [];
+ my $seen = {};
+
+ for my $hostnqn (split(/,/, $value // '')) {
+ $hostnqn = trim($hostnqn);
+ if (!verify_nvme_nqn($hostnqn, 1) || $seen->{$hostnqn}++) {
+ return undef if $noerr;
+ die "invalid or duplicate NVMe host NQN '$hostnqn'\n";
+ }
+ push $result->@*, $hostnqn;
+ if (scalar($result->@*) > $max_hosts) {
+ return undef if $noerr;
+ die "at most $max_hosts NVMe host NQNs are supported\n";
+ }
+ }
+
+ if (!$result->@*) {
+ return undef if $noerr;
+ die "at least one NVMe host NQN is required\n";
+ }
+
+ return $result;
+}
+
+my sub verify_nvme_host_ifaces($value, $noerr = undef) {
+ return undef if !parse_nvme_host_ifaces($value, $noerr);
+ return $value;
+}
+
+my sub verify_nvme_host_nqns($value, $noerr = undef) {
+ return undef if !parse_nvme_host_nqns($value, $noerr);
+ return $value;
+}
+
+sub _configured_portals($scfg) {
+ my $portals = parse_nvme_portals($scfg->{'nvme-portals'});
+ my $ifaces = parse_nvme_host_ifaces($scfg->{'nvme-host-ifaces'});
+ die "nvme-host-ifaces must contain one interface for each nvme-portals entry\n"
+ if scalar($ifaces->@*) != scalar($portals->@*);
+
+ for (my $i = 0; $i < scalar($portals->@*); $i++) {
+ $portals->[$i]->{host_iface} = $ifaces->[$i];
+ }
+ return $portals;
+}
+
+sub _local_iface_exists($iface) {
+ return -d "/sys/class/net/$iface";
+}
+
+sub _validate_local_ifaces($portals) {
+ for my $portal ($portals->@*) {
+ my $iface = $portal->{host_iface};
+ die "NVMe/TCP host interface '$iface' does not exist on this node\n"
+ if !_local_iface_exists($iface);
+ }
+}
+
+PVE::JSONSchema::register_format('pve-storage-nvme-nqn', \&verify_nvme_nqn);
+PVE::JSONSchema::register_format('pve-storage-nvme-portals', \&verify_nvme_portals);
+PVE::JSONSchema::register_format('pve-storage-nvme-host-ifaces', \&verify_nvme_host_ifaces);
+PVE::JSONSchema::register_format('pve-storage-nvme-host-nqns', \&verify_nvme_host_nqns);
+
+sub type($class) {
+ return 'zfsnvme';
+}
+
+sub plugindata($class) {
+ return {
+ content => [{ images => 1 }, { images => 1 }],
+ 'sensitive-properties' => { 'dhchap-key' => 1 },
+ };
+}
+
+sub properties($class) {
+ return {
+ subsysnqn => {
+ description => "NVMe subsystem qualified name.",
+ type => 'string',
+ format => 'pve-storage-nvme-nqn',
+ },
+ 'nvme-portals' => {
+ description =>
+ "Comma-separated NVMe/TCP target IP addresses, at most 16. Put IPv6 addresses in"
+ . " brackets. The default TCP port is 4420.",
+ type => 'string',
+ format => 'pve-storage-nvme-portals',
+ maxLength => 2048,
+ },
+ 'nvme-host-ifaces' => {
+ description =>
+ "Comma-separated local interfaces, in portal order. Interface names must be"
+ . " identical on every cluster node.",
+ type => 'string',
+ format => 'pve-storage-nvme-host-ifaces',
+ maxLength => 512,
+ },
+ 'nvme-host-nqns' => {
+ description =>
+ "Comma-separated /etc/nvme/hostnqn values for every cluster node allowed to use"
+ . " this storage, at most 64. Host NQNs can only be added.",
+ type => 'string',
+ format => 'pve-storage-nvme-host-nqns',
+ maxLength => 8192,
+ },
+ 'dhchap-key' => {
+ description =>
+ "NVMe DH-HMAC-CHAP key in secret representation format (DHHC-1). Required when the"
+ . " storage is created; it cannot be changed.",
+ type => 'string',
+ maxLength => 256,
+ },
+ 'nvme-iopolicy' => {
+ description => "Native NVMe multipath I/O policy.",
+ type => 'string',
+ enum => ['numa', 'round-robin', 'queue-depth'],
+ default => 'round-robin',
+ },
+ 'nvme-keep-alive-tmo' => {
+ description => "NVMe keep-alive timeout in seconds.",
+ type => 'integer',
+ minimum => 1,
+ maximum => 120,
+ default => 5,
+ },
+ 'nvme-reconnect-delay' => {
+ description => "Delay between NVMe reconnect attempts in seconds.",
+ type => 'integer',
+ minimum => 1,
+ maximum => 120,
+ default => 2,
+ },
+ 'nvme-ctrl-loss-tmo' => {
+ description => "Time to keep retrying a lost NVMe controller in seconds.",
+ type => 'integer',
+ minimum => -1,
+ maximum => 86400,
+ default => 600,
+ },
+ 'nvme-fast-io-fail-tmo' => {
+ description =>
+ "Optional time before failing I/O on a reconnecting NVMe controller. Unset queues"
+ . " I/O until controller loss timeout.",
+ type => 'integer',
+ minimum => 0,
+ maximum => 86400,
+ optional => 1,
+ },
+ 'nvme-nr-io-queues' => {
+ description => "Number of NVMe/TCP I/O queues per controller.",
+ type => 'integer',
+ minimum => 1,
+ maximum => 1024,
+ optional => 1,
+ },
+ };
+}
+
+sub options($class) {
+ return {
+ server => { fixed => 1 },
+ subsysnqn => { fixed => 1 },
+ 'nvme-portals' => { fixed => 1 },
+ 'nvme-host-ifaces' => {},
+ 'nvme-host-nqns' => {},
+ pool => { fixed => 1 },
+ blocksize => { fixed => 1 },
+ sparse => { optional => 1 },
+ 'dhchap-key' => { optional => 1 },
+ 'nvme-iopolicy' => { optional => 1 },
+ 'nvme-keep-alive-tmo' => { optional => 1 },
+ 'nvme-reconnect-delay' => { optional => 1 },
+ 'nvme-ctrl-loss-tmo' => { optional => 1 },
+ 'nvme-fast-io-fail-tmo' => { optional => 1 },
+ 'nvme-nr-io-queues' => { optional => 1 },
+ nodes => { optional => 1 },
+ disable => { optional => 1 },
+ content => { optional => 1 },
+ bwlimit => { optional => 1 },
+ };
+}
+
+sub _validate_fail_fast_timeout($config, $default_ctrl_loss_tmo = undef) {
+ my $fast = $config->{'nvme-fast-io-fail-tmo'};
+ return if !defined($fast);
+
+ my $ctrl = $config->{'nvme-ctrl-loss-tmo'};
+ $ctrl = $default_ctrl_loss_tmo if !defined($ctrl);
+ return if !defined($ctrl) || $ctrl < 0;
+
+ die "nvme-fast-io-fail-tmo must not exceed nvme-ctrl-loss-tmo\n"
+ if $fast > $ctrl;
+}
+
+# Defaults are not injected here: this also runs when storage.cfg is parsed.
+# Every use site applies the schema default itself.
+sub check_config($class, $section_id, $config, $create, $skip_schema_check = undef) {
+ nvmet_pool($config) if defined($config->{pool});
+ verify_nvme_nqn($config->{subsysnqn}) if defined($config->{subsysnqn});
+ parse_nvme_portals($config->{'nvme-portals'}) if defined($config->{'nvme-portals'});
+ if (defined($config->{'nvme-host-ifaces'}) && defined($config->{'nvme-portals'})) {
+ _configured_portals($config);
+ } elsif (defined($config->{'nvme-host-ifaces'})) {
+ parse_nvme_host_ifaces($config->{'nvme-host-ifaces'});
+ }
+ parse_nvme_host_nqns($config->{'nvme-host-nqns'})
+ if defined($config->{'nvme-host-nqns'});
+ _validate_fail_fast_timeout($config, $create ? 600 : undef);
+ return $class->SUPER::check_config($section_id, $config, $create, $skip_schema_check);
+}
+
+# ---------------------------------------------------------------------------
+# ZFS helpers (same formats as the other ZFS plugins)
+# ---------------------------------------------------------------------------
+
+my sub zfs_request($scfg, $timeout, $method, @args) {
+ $timeout = PVE::RPCEnvironment->is_worker() ? 60 * 60 : 10 if !$timeout;
+ my $change = $method ne 'get' && $method ne 'list';
+ my $res = nvmet_exec(
+ $scfg, [['zfs', $method, @args]],
+ op => "zfs $method",
+ timeout => $timeout,
+ change => $change,
+ );
+ die "zfs error on '" . nvmet_server($scfg) . "': $res->{err}\n" if $res->{rc};
+ return nvmet_output($res);
+}
+
+# The direct children of the pool with PVE volume names. The plugin creates
+# and owns zvols only; a filesystem child holds its name.
+my sub zfs_parse_zvol_list($text, $pool) {
+ my $list = [];
+ for my $line (split /\n/, $text // '') {
+ my ($dataset, $size, $origin, $type) = split(/\s+/, $line);
+ next if !defined($type) || ($type ne 'volume' && $type ne 'filesystem');
+ my @parts = split /\//, $dataset;
+ next if @parts < 2;
+ my $name = pop @parts;
+ next if join('/', @parts) ne $pool;
+ next if $name !~ $RE_ZVOL_OWNER;
+ push $list->@*,
+ {
+ name => $name,
+ owner => $+{owner},
+ type => $type,
+ ($type eq 'volume' ? (size => $size + 0) : ()),
+ ($origin ne '-' ? (origin => $origin) : ()),
+ };
+ }
+ return $list;
+}
+
+# The pool's direct children with PVE volume names, so that a name is never
+# handed out twice; with $owned_only, the zvols owned by this storage's NVMe
+# subsystem.
+my sub zfs_list_zvol($scfg, $owned_only) {
+ my $pool = nvmet_pool($scfg);
+ my $text = zfs_request(
+ $scfg,
+ 10,
+ 'list',
+ '-o',
+ 'name,volsize,origin,type',
+ '-t',
+ 'volume,filesystem',
+ '-d1',
+ '-Hp',
+ $pool,
+ );
+ my $list = {};
+ my $prefix = "$pool/";
+ for my $zvol (zfs_parse_zvol_list($text, $pool)->@*) {
+ next if $owned_only && $zvol->{type} ne 'volume';
+ my $parent = $zvol->{origin};
+ # an origin in the pool is named relative to it, like the ZFS plugins do
+ $parent = substr($parent, length($prefix)) if $parent && index($parent, $prefix) == 0;
+ $list->{ $zvol->{name} } = {
+ name => $zvol->{name},
+ size => $zvol->{size},
+ parent => $parent,
+ format => 'raw',
+ vmid => $zvol->{owner},
+ };
+ }
+ return $list if !$owned_only;
+
+ my $properties = zfs_request(
+ $scfg,
+ 10,
+ 'get',
+ '-H',
+ '-d',
+ '1',
+ '-o',
+ 'name,value,source',
+ 'proxmox:nvme-subsys',
+ $pool,
+ );
+ my $owned = {};
+ for my $line (split(/\n/, $properties)) {
+ my ($dataset, $nqn, $source) = split(/\t/, $line, 3);
+ next if !defined($source) || ($source ne 'local' && $source ne 'received');
+ next if !defined($nqn) || $nqn ne $scfg->{subsysnqn};
+ next if index($dataset, $prefix) != 0;
+ $owned->{ substr($dataset, length($prefix)) } = 1;
+ }
+ for my $name (keys $list->%*) {
+ delete $list->{$name} if !$owned->{$name};
+ }
+ return $list;
+}
+
+my sub zfs_get_properties($scfg, $properties, $dataset, $timeout = undef) {
+ my $text = zfs_request($scfg, $timeout, 'get', '-o', 'value', '-Hp', $properties, $dataset);
+ my @values = split /\n/, $text;
+ return wantarray ? @values : $values[0];
+}
+
+my sub zfs_get_sorted_snapshot_list($scfg, $name, $sort_params) {
+ return [
+ map { s/^.*\@//r } split /\n/,
+ zfs_request(
+ $scfg,
+ undef,
+ 'list',
+ '-H',
+ '-r',
+ '-t',
+ 'snapshot',
+ '-o',
+ 'name',
+ $sort_params->@*,
+ nvmet_dataset($scfg, $name),
+ ),
+ ];
+}
+
+# ---------------------------------------------------------------------------
+# Storage API: volumes
+# ---------------------------------------------------------------------------
+
+sub parse_volname($class, $volname) {
+ if ($volname =~ $RE_VOLNAME) {
+ my $format = ($+{type} eq 'subvol' || $+{type} eq 'basevol') ? 'subvol' : 'raw';
+ my $is_base = $+{type} eq 'base' || $+{type} eq 'basevol';
+ return ('images', $+{name}, $+{vmid}, $+{base}, $+{base_vmid}, $is_base, $format);
+ }
+ die "unable to parse zfs volume name '$volname'\n";
+}
+
+sub list_images($class, $storeid, $scfg, $vmid = undef, $vollist = undef, $cache = undef) {
+ my $res = [];
+ for my $info (values zfs_list_zvol($scfg, 1)->%*) {
+ my $volname =
+ $info->{parent} && $info->{parent} =~ $RE_BASE_SNAPSHOT
+ ? "$storeid:$+{base}/$info->{name}"
+ : "$storeid:$info->{name}";
+ next
+ if $vollist
+ ? !grep { $_ eq $volname } $vollist->@*
+ : defined($vmid) && $info->{vmid} ne $vmid;
+ $info->{volid} = $volname;
+ push $res->@*, $info;
+ }
+ return $res;
+}
+
+# Names are allocated against every direct child of the pool, owned or not, so
+# a leftover zvol that list_images hides is never handed out again.
+sub find_free_diskname($class, $storeid, $scfg, $vmid, $fmt = undef, $add_fmt_suffix = undef) {
+ my $disk_list = [map { "$storeid:$_" } sort keys zfs_list_zvol($scfg, 0)->%*];
+ return PVE::Storage::Plugin::get_next_vm_diskname(
+ $disk_list, $storeid, $vmid, $fmt, $scfg, $add_fmt_suffix,
+ );
+}
+
+sub status($class, $storeid, $scfg, $cache = undef) {
+ my ($available, $used) =
+ eval { zfs_get_properties($scfg, 'available,used', nvmet_pool($scfg)) };
+ if (my $err = $@) {
+ warn "storage '$storeid': $err";
+ return (0, 0, 0, 0);
+ }
+ if (
+ !defined($available)
+ || !defined($used)
+ || $available !~ $RE_UNSIGNED_INTEGER
+ || $used !~ $RE_UNSIGNED_INTEGER
+ ) {
+ warn "unexpected ZFS pool usage for storage '$storeid'\n";
+ return (0, 0, 0, 0);
+ }
+ return ($available + $used, $available, $used, 1);
+}
+
+sub volume_size_info($class, $scfg, $storeid, $volname, $timeout = undef) {
+ my (undef, $name, undef, $parent, undef, undef, $format) = $class->parse_volname($volname);
+ die "volume_size_info requires a ZFS volume\n" if $format ne 'raw';
+ my ($size, $used) = zfs_get_properties(
+ $scfg, 'volsize,usedbydataset', nvmet_dataset($scfg, $name), $timeout,
+ );
+ die "Could not get zfs volume size\n" if !defined($size) || $size !~ $RE_UNSIGNED_INTEGER;
+ $used = defined($used) && $used =~ $RE_UNSIGNED_INTEGER ? $used + 0 : 0;
+ return wantarray ? ($size + 0, 'raw', $used, $parent) : $size + 0;
+}
+
+sub volume_snapshot($class, $scfg, $storeid, $volname, $snap) {
+ my (undef, $name, undef, undef, undef, undef, $format) = $class->parse_volname($volname);
+ die "volume_snapshot requires a ZFS volume\n" if $format ne 'raw';
+ my $snapshot = nvmet_snapshot($scfg, $name, $snap);
+ nvmet_locked($scfg, sub { zfs_request($scfg, undef, 'snapshot', $snapshot) });
+ return;
+}
+
+sub volume_snapshot_delete($class, $scfg, $storeid, $volname, $snap, $running = undef) {
+ my $name = ($class->parse_volname($volname))[1];
+ my $snapshot = nvmet_snapshot($scfg, $name, $snap);
+ nvmet_locked($scfg, sub { zfs_request($scfg, undef, 'destroy', $snapshot) });
+ return;
+}
+
+sub volume_rollback_is_possible($class, $scfg, $storeid, $volname, $snap, $blockers = undef) {
+ my $name = ($class->parse_volname($volname))[1];
+ nvmet_snapshot($scfg, $name, $snap); # validation only
+ my $found;
+ $blockers //= [];
+ for my $snapshot (zfs_get_sorted_snapshot_list($scfg, $name, ['-s', 'creation'])->@*) {
+ $found = 1 if $snapshot eq $snap;
+ push $blockers->@*, $snapshot if $found && $snapshot ne $snap;
+ }
+ die "can't rollback, snapshot '$snap' does not exist on '${storeid}:${volname}'\n" if !$found;
+ die "can't rollback, '$snap' is not most recent snapshot on '${storeid}:${volname}'\n"
+ if $blockers->@*;
+ return 1;
+}
+
+sub volume_snapshot_rollback($class, $scfg, $storeid, $volname, $snap) {
+ my (undef, $name, undef, undef, undef, undef, $format) = $class->parse_volname($volname);
+ die "snapshot rollback requires a ZFS volume\n" if $format ne 'raw';
+ nvmet_rollback_volume($scfg, $name, $snap);
+}
+
+sub volume_snapshot_info($class, $scfg, $storeid, $volname) {
+ my $name = ($class->parse_volname($volname))[1];
+ my $info = {};
+ my $text = zfs_request(
+ $scfg,
+ undef,
+ 'list',
+ '-Hp',
+ '-r',
+ '-t',
+ 'snapshot',
+ '-o',
+ 'name,guid,creation',
+ nvmet_dataset($scfg, $name),
+ );
+ for my $line (split /\n/, $text) {
+ my ($snapshot, $guid, $creation) = split /\s+/, $line;
+ $snapshot =~ s/^.*\@//;
+ $info->{$snapshot} = { id => $guid, timestamp => $creation };
+ }
+ return $info;
+}
+
+sub free_image($class, $storeid, $scfg, $volname, $is_base = undef, $format = undef) {
+ my (undef, $name, undef, undef, undef, undef, $volformat) = $class->parse_volname($volname);
+ die "free_image requires a ZFS volume\n" if $volformat ne 'raw';
+ nvmet_destroy_volume($scfg, $name);
+ return undef;
+}
+
+sub create_base($class, $storeid, $scfg, $volname) {
+ my (undef, $name, undef, $basename, undef, $is_base, $format) = $class->parse_volname($volname);
+ die "create_base not possible with base image\n" if $is_base;
+ die "create_base requires a ZFS volume\n" if $format ne 'raw';
+ my $newname = $name =~ s/^vm-/base-/r;
+ nvmet_template_volume($scfg, $name);
+ return $basename ? "$basename/$newname" : $newname;
+}
+
+sub clone_image($class, $scfg, $storeid, $volname, $vmid, $snap = undef) {
+ $snap ||= '__base__';
+ my (undef, $basename, undef, undef, undef, $is_base, $format) = $class->parse_volname($volname);
+ die "clone_image only works on base images\n" if !$is_base;
+ die "clone_image requires a ZFS volume\n" if $format ne 'raw';
+ my $origin = nvmet_snapshot($scfg, $basename, $snap);
+ my $name = $class->find_free_diskname($storeid, $scfg, $vmid, $format);
+ nvmet_create_volume($scfg, $name, origin => $origin);
+ return "$basename/$name";
+}
+
+sub alloc_image($class, $storeid, $scfg, $vmid, $fmt, $name, $size) {
+ die "unsupported format '$fmt'" if $fmt ne 'raw';
+ die "illegal name '$name' - should be 'vm-$vmid-*'\n"
+ if $name && index($name, "vm-$vmid-") != 0;
+ my $volname = $name || $class->find_free_diskname($storeid, $scfg, $vmid, $fmt);
+ $size += (1024 - $size % 1024) % 1024;
+ nvmet_create_volume($scfg, $volname, size => $size);
+ return $volname;
+}
+
+sub volume_resize($class, $scfg, $storeid, $volname, $size, $running = undef, $snapname = undef) {
+ # QEMU can resize a host_device once the device itself has grown, but the
+ # plugin cannot yet wait for the local NVMe namespace to report the new
+ # size. Until it can, refuse online resize before changing the zvol, so a
+ # running VM never sees a size different from its configuration.
+ die "online resize is not supported for NVMe/TCP block devices; stop the VM first\n"
+ if $running;
+ die "resizing a snapshot is not supported for $class\n" if $snapname;
+ my $new_size = int($size / 1024);
+ $new_size += (1024 - $new_size % 1024) % 1024;
+ my $name = ($class->parse_volname($volname))[1];
+ nvmet_resize_volume($scfg, $name, $new_size);
+ return $new_size;
+}
+
+sub volume_has_feature(
+ $class, $scfg, $feature, $storeid, $volname,
+ $snapname = undef,
+ $running = undef,
+ $opts = undef,
+) {
+ my $features = {
+ snapshot => { current => 1, snap => 1 },
+ clone => { base => 1 },
+ template => { current => 1 },
+ copy => { base => 1, current => 1 },
+ };
+ my $is_base = ($class->parse_volname($volname))[5];
+ my $key = $snapname ? 'snap' : $is_base ? 'base' : 'current';
+ return $features->{$feature}->{$key};
+}
+
+# No stream transfers or renames; the storage has no path, so the base class
+# offers no export or import format either.
+sub volume_export($class, @args) {
+ die "ZFS stream export is not supported for NVMe/TCP storage\n";
+}
+
+sub volume_import($class, @args) {
+ die "ZFS stream import is not supported for NVMe/TCP storage\n";
+}
+
+sub rename_volume($class, @args) {
+ die "renaming volumes is not supported for NVMe/TCP storage\n";
+}
+
+sub rename_snapshot($class, @args) {
+ die "rename_snapshot is not supported for $class";
+}
+
+# ---------------------------------------------------------------------------
+# Secrets and configuration hooks
+# ---------------------------------------------------------------------------
+
+# Each storage needs its own NVMe/TCP listener (address family, address and
+# port), and a disabled storage keeps its listener. nvmet answers a connect to
+# a listener that does not publish the subsystem yet with DNR, and the host
+# then deletes the controller whatever its loss timeout: after a target
+# restart, the storage published second would lose its paths. Storages that
+# reach the same data address must spell the server alike, as it names the
+# target lock.
+sub _assert_unique_target($storeid, $scfg, $cfg = undef) {
+ $cfg //= PVE::Storage::config();
+ my (%listeners, %addresses);
+ for my $portal (parse_nvme_portals($scfg->{'nvme-portals'})->@*) {
+ my $address = join("\0", nvmet_listener_address($portal));
+ $listeners{"$address\0$portal->{port}"} = 1;
+ $addresses{$address} = 1;
+ }
+ my $ids = $cfg->{ids} // {};
+ for my $other_id (sort keys $ids->%*) {
+ next if $other_id eq $storeid;
+ my $other = $ids->{$other_id};
+ next if ($other->{type} // '') ne 'zfsnvme';
+
+ die "NVMe subsystem NQN is already used by storage '$other_id'\n"
+ if ($other->{subsysnqn} // '') eq $scfg->{subsysnqn};
+ my $other_server = $other->{server} // '';
+ # a storage whose portals do not parse cannot be activated
+ for my $portal ((parse_nvme_portals($other->{'nvme-portals'}, 1) // [])->@*) {
+ my $address = join("\0", nvmet_listener_address($portal));
+ die "NVMe/TCP portal '$portal->{address}' port $portal->{port} is already used by"
+ . " storage '$other_id'; give each storage its own address or port\n"
+ if $listeners{"$address\0$portal->{port}"};
+ die "storage '$other_id' reaches target address '$portal->{address}' through server"
+ . " '$other_server'; use the same server value for both storages\n"
+ if $addresses{$address} && $other_server ne $scfg->{server};
+ }
+ next if $other_server ne $scfg->{server};
+ my ($pool, $other_pool) = ($scfg->{pool}, $other->{pool} // '');
+ die "ZFS pool '$pool' on '$scfg->{server}' is already used by storage '$other_id'\n"
+ if $other_pool eq $pool;
+ die "ZFS pool '$pool' on '$scfg->{server}' overlaps the pool of storage '$other_id'\n"
+ if index("$other_pool/", "$pool/") == 0 || index("$pool/", "$other_pool/") == 0;
+ }
+}
+
+# nvmet keeps one DH-HMAC-CHAP key per host NQN, so storages on the same target
+# whose host NQNs overlap must use the same key.
+sub _assert_shared_host_key($storeid, $scfg, $key, $cfg = undef) {
+ $cfg //= PVE::Storage::config();
+ my %hosts = map { $_ => 1 } (parse_nvme_host_nqns($scfg->{'nvme-host-nqns'}, 1) // [])->@*;
+ my $ids = $cfg->{ids} // {};
+ for my $other_id (sort keys $ids->%*) {
+ next if $other_id eq $storeid;
+ my $other = $ids->{$other_id};
+ next if ($other->{type} // '') ne 'zfsnvme';
+ next if ($other->{server} // '') ne ($scfg->{server} // '');
+ my $other_hosts = parse_nvme_host_nqns($other->{'nvme-host-nqns'}, 1) // [];
+ next if !grep { $hosts{$_} } $other_hosts->@*;
+ my $other_key = file_read_firstline(secret_path($other_id));
+ die "storage '$other_id' uses a different DH-HMAC-CHAP key for the same NVMe host NQNs"
+ . " on '$scfg->{server}'\n"
+ if defined($other_key) && $other_key ne $key;
+ }
+}
+
+# DHHC-1:<hash>:<base64 of key and its little-endian CRC-32>: as generated by
+# `nvme gen-dhchap-key`. Messages never contain the key.
+sub _validate_secret($key) {
+ die "missing NVMe DH-HMAC-CHAP key\n" if !defined($key) || $key eq '';
+ my $invalid = "invalid NVMe DH-HMAC-CHAP key representation\n";
+ die $invalid if $key !~ $RE_DHCHAP_KEY;
+ my ($hash, $encoded) = @+{qw(hash secret)};
+ die $invalid if length($encoded) % 4;
+ my $decoded = decode_base64($encoded);
+ my $length = length($decoded) - 4;
+ my %lengths = ('00' => [32, 48, 64], '01' => [32], '02' => [48], '03' => [64]);
+ die $invalid if !grep { $_ == $length } $lengths{$hash}->@*;
+ die $invalid if pack('V', crc32(substr($decoded, 0, $length))) ne substr($decoded, $length);
+ return $key;
+}
+
+my sub set_secret($storeid, $key) {
+ _validate_secret($key);
+ make_path($secret_dir, { mode => 0700 });
+ file_set_contents(secret_path($storeid), "$key\n", 0600);
+}
+
+my sub get_secret($storeid) {
+ my $key = file_read_firstline(secret_path($storeid));
+ return _validate_secret($key);
+}
+
+# Removes the key file of a storage (a test seam).
+sub _unlink_file($path) {
+ return unlink($path);
+}
+
+my sub delete_secret($storeid) {
+ _unlink_file(secret_path($storeid));
+}
+
+sub on_add_hook($class, $storeid, $scfg, %sensitive) {
+ _configured_portals($scfg);
+ parse_nvme_host_nqns($scfg->{'nvme-host-nqns'});
+ _assert_unique_target($storeid, $scfg);
+ my $key = _validate_secret($sensitive{'dhchap-key'});
+ _assert_shared_host_key($storeid, $scfg, $key);
+ set_secret($storeid, $key);
+ return;
+}
+
+sub on_update_hook_full($class, $storeid, $scfg, $update, $delete = undef, $sensitive = undef) {
+ $sensitive //= {};
+ my %prospective = ($scfg->%*, $update->%*);
+ delete @prospective{ $delete->@* } if $delete;
+ verify_nvme_nqn($prospective{subsysnqn});
+ _configured_portals(\%prospective);
+ my $new_hostnqns = parse_nvme_host_nqns($prospective{'nvme-host-nqns'});
+ my %new_hosts = map { $_ => 1 } $new_hostnqns->@*;
+ for my $old_hostnqn (parse_nvme_host_nqns($scfg->{'nvme-host-nqns'})->@*) {
+ die "removing NVMe host NQN '$old_hostnqn' is not supported; move every volume off"
+ . " the storage and recreate it with a new subsystem NQN and a new DH-HMAC-CHAP"
+ . " key\n"
+ if !$new_hosts{$old_hostnqn};
+ }
+ _validate_fail_fast_timeout(\%prospective, 600);
+ _assert_unique_target($storeid, \%prospective);
+
+ my $old_key = file_read_firstline(secret_path($storeid));
+ my $key = exists($sensitive->{'dhchap-key'}) ? $sensitive->{'dhchap-key'} : $old_key;
+ _validate_secret($key);
+
+ # nvmet stores authentication on the global Host NQN object, which can be
+ # shared by multiple subsystems. Replacing an active key in place can make
+ # unrelated storages unrecoverable on their next reconnect. Until PVE can
+ # coordinate a rolling rotation on every node, fail before mutating the
+ # cluster-wide secret.
+ die "NVMe DH-HMAC-CHAP key rotation is not supported; create a new storage with a new"
+ . " subsystem NQN\n"
+ if defined($old_key) && $key ne $old_key;
+ _assert_shared_host_key($storeid, \%prospective, $key);
+
+ set_secret($storeid, $key) if exists($sensitive->{'dhchap-key'});
+ return;
+}
+
+# ---------------------------------------------------------------------------
+# Local NVMe host side
+# ---------------------------------------------------------------------------
+
+my %slow_path_last; # storeid => time of the last slow-path activation
+
+# Probes of the local host (test seams).
+sub _block_device($path) {
+ return -b $path;
+}
+
+sub _link_target($path) {
+ return readlink($path);
+}
+
+sub _fabrics_device() {
+ return '/dev/nvme-fabrics';
+}
+
+# Runs $code in a child and gives up on it after $timeout seconds. Returns the
+# child's result; the child's errors pass through as exceptions, a timeout
+# becomes one. The bound is not hard: as with PVE::Tools::run_command, a
+# child stuck in the kernel (a write or delete in D state) is only reaped
+# once the kernel returns, and a timeout does not mean that the operation did
+# not happen.
+#
+# A stopped task (TERM, also HUP, INT and QUIT) stays pending in the parent
+# while the child runs, and the child inherits the mask: a connect that ends
+# before PVE escalates the stop to KILL, 5 seconds after the TERM, creates
+# the controller and restricts its secret attributes without interruption,
+# and run_fork_with_timeout, whose own handlers would consume the signal when
+# the child delivers a result, never sees it. Once the child is reaped and the
+# mask is restored, the handler of the worker dies with the task marker, also
+# over a result or an error of the child. ALRM, the timeout, is not blocked.
+# The KILL ends the parent and the child together; a controller that the
+# child had created but not restricted yet keeps its secret attributes
+# world-readable until the next activation restricts them.
+#
+# run_fork_with_timeout is called in list context on purpose: only there does
+# run_with_timeout report a timeout as its second return value; in scalar
+# context the timeout would be a warning and an undef result.
+my sub run_bounded($timeout, $what, $code) {
+ my $previous = POSIX::SigSet->new();
+ sigprocmask(SIG_BLOCK, POSIX::SigSet->new(SIGHUP, SIGINT, SIGQUIT, SIGTERM), $previous)
+ or die "cannot block signals: $!\n";
+ my $outer_warn = $SIG{__WARN__};
+ my ($res, $timed_out) = eval {
+ # Only a child that leaves no result makes run_fork_with_timeout warn
+ # in the parent; that case is an exception below. The child keeps the
+ # handler of the caller, so its own warnings reach the task log.
+ local $SIG{__WARN__} = sub($message) { };
+ PVE::Tools::run_fork_with_timeout(
+ $timeout,
+ sub {
+ local $SIG{__WARN__} = $outer_warn;
+ return $code->();
+ },
+ );
+ };
+ my $error = $@;
+ sigprocmask(SIG_SETMASK, $previous) or die "cannot restore signals: $!\n";
+ if ($error) {
+ rethrow_task_interrupt($error);
+ die $error;
+ }
+ die "$what did not complete within $timeout seconds\n" if $timed_out;
+ die "$what left no result\n" if !defined($res);
+ return $res;
+}
+
+# The connect options of one controller, as the fabrics device takes them,
+# and their names. The values are validated configuration and host identities
+# checked by the caller; the kernel reads each one up to the next ','.
+sub _fabrics_options($scfg, $portal, $hostnqn, $hostid, $key) {
+ my @options = (
+ [transport => 'tcp'],
+ [traddr => $portal->{address}],
+ [trsvcid => $portal->{port}],
+ [host_iface => $portal->{host_iface}],
+ [nqn => $scfg->{subsysnqn}],
+ [hostnqn => $hostnqn],
+ [hostid => $hostid],
+ [dhchap_secret => $key],
+ [keep_alive_tmo => $scfg->{'nvme-keep-alive-tmo'} // 5],
+ [reconnect_delay => $scfg->{'nvme-reconnect-delay'} // 2],
+ [ctrl_loss_tmo => $scfg->{'nvme-ctrl-loss-tmo'} // 600],
+ );
+ # unlike unset, which leaves fast I/O failure off, 0 is a real timeout
+ push @options, [fast_io_fail_tmo => $scfg->{'nvme-fast-io-fail-tmo'}]
+ if defined($scfg->{'nvme-fast-io-fail-tmo'});
+ push @options, [nr_io_queues => $scfg->{'nvme-nr-io-queues'}]
+ if defined($scfg->{'nvme-nr-io-queues'});
+ for my $option (@options) {
+ die "internal error: invalid NVMe connect option '$option->[0]'\n"
+ if !defined($option->[1]) || $option->[1] !~ $RE_FABRICS_VALUE;
+ }
+ return (join(',', map { "$_->[0]=$_->[1]" } @options), [map { $_->[0] } @options]);
+}
+
+# The child of a connect (a test seam): creates exactly one controller with
+# one write of $options to the fabrics device and returns its instance. The
+# write is never repeated, because the kernel may have created the controller
+# even when it reports an error. Errors carry the errno text of the failed
+# operation and never the options, which hold the key.
+sub _fabrics_connect($options, $names) {
+ my $device = _fabrics_device();
+ my $probe;
+ if (!sysopen($probe, $device, O_RDONLY | O_NOFOLLOW)) {
+ die "'$device' does not exist; load the nvme-tcp kernel module\n" if $! == ENOENT;
+ die "cannot open '$device': $!\n";
+ }
+ my @stat = stat($probe);
+ die "'$device' is not a character device\n" if !@stat || !S_ISCHR($stat[2]);
+ # without a controller, a read lists the supported options as patterns
+ my $supported = '';
+ sysread($probe, $supported, 4096) // die "cannot read '$device': $!\n";
+ close($probe);
+ my %supported = map { s/=.*//r => 1 } split(/,/, $supported =~ s/\n\z//r);
+ for my $name ($names->@*) {
+ die "the running kernel does not support the NVMe connect option '$name'\n"
+ if !$supported{$name};
+ }
+
+ # Not PVE::SysFSTools::file_write: the kernel returns the instance of the
+ # new controller on a read of the file descriptor that wrote the options.
+ sysopen(my $fh, $device, O_RDWR | O_NOFOLLOW) or die "cannot open '$device': $!\n";
+ my $written = _fabrics_write($fh, $options);
+ die "connect failed: $!\n" if !defined($written);
+ die "connect failed: short write\n" if $written != length($options);
+ my $result = '';
+ sysread($fh, $result, 128) // die "cannot read the connect result: $!\n";
+ die "unexpected connect result\n" if $result !~ $RE_FABRICS_RESULT;
+ my $instance = int($+{instance});
+ # The kernel creates the secret attributes of a controller world-readable;
+ # the caller keeps a stopped task from interrupting before this, short of
+ # a KILL.
+ _restrict_attr("/sys/class/nvme/nvme$instance/$_") for qw(dhchap_secret dhchap_ctrl_secret);
+ close($fh);
+ return $instance;
+}
+
+# The write that creates a controller (a test seam).
+sub _fabrics_write($fh, $options) {
+ return syswrite($fh, $options);
+}
+
+# Writes a sysfs attribute and returns the errno text of a failure, or undef.
+# PVE::SysFSTools::file_write returns undef without a warning when the open
+# fails, leaving its errno in $!, and 0 after warning "error writing ...:
+# <errno text>" when the write fails; that warning is not passed on. With
+# $missing_ok, an attribute that does not exist, because its controller went
+# away, is not a failure.
+my sub write_attribute($path, $value, $missing_ok = 0) {
+ my $warning;
+ my $written = do {
+ local $SIG{__WARN__} = sub($message) { $warning = $message };
+ PVE::SysFSTools::file_write($path, $value);
+ };
+ return if $written;
+ return if !defined($warning) && $missing_ok && $! == ENOENT;
+ return "$!" if !defined($warning);
+ return $warning =~ s/\A.*: //sr =~ s/\n\z//r;
+}
+
+# Deletes controllers (the body of a child, a test seam): the deletion can
+# block in the kernel. The error names the controller and the errno text; a
+# controller that is already gone counts as deleted.
+sub _delete_controllers($controllers) {
+ for my $controller ($controllers->@*) {
+ my $path = "/sys/class/nvme/$controller->{name}/delete_controller";
+ my $error = write_attribute($path, '1', 1) // next;
+ die "cannot delete NVMe controller '$controller->{name}': $error\n";
+ }
+ return scalar($controllers->@*);
+}
+
+# Every local controller of the subsystem.
+my sub nvme_controllers($nqn) {
+ my $controllers = [];
+ PVE::File::dir_glob_foreach(
+ '/sys/class/nvme',
+ $RE_NVME_CONTROLLER,
+ sub($entry) {
+ my $base = "/sys/class/nvme/$entry";
+ my $subsys = file_read_firstline("$base/subsysnqn");
+ return if !defined($subsys) || $subsys ne $nqn;
+ my $address = file_read_firstline("$base/address") // '';
+ my $traddr = $address =~ $RE_TRADDR ? $+{value} : undef;
+ my $trsvcid = $address =~ $RE_TRSVCID ? $+{value} : undef;
+ my $host_iface = $address =~ $RE_HOST_IFACE_ADDRESS ? $+{value} : undef;
+ push $controllers->@*,
+ {
+ name => $entry,
+ state => file_read_firstline("$base/state") // 'unknown',
+ traddr => nvmet_canonical_address($traddr),
+ trsvcid => $trsvcid,
+ host_iface => $host_iface,
+ };
+ },
+ );
+ return $controllers;
+}
+
+# Controllers by "address:port", with canonical IPv6 addresses. A portal has
+# more than one controller while its path moves to another interface.
+my sub controller_states($nqn) {
+ my $states = {};
+ for my $controller (nvme_controllers($nqn)->@*) {
+ next if !defined($controller->{traddr}) || !defined($controller->{trsvcid});
+ push $states->{"$controller->{traddr}:$controller->{trsvcid}"}->@*, $controller;
+ }
+ return $states;
+}
+
+# The multipath namespace devices of the subsystem and their partitions. One
+# dir_glob_foreach per directory level, because PVE::File::dir_glob_regex
+# returns only the first match.
+my sub namespace_devices($nqn) {
+ my $devices = {};
+
+ PVE::File::dir_glob_foreach(
+ '/sys/class/nvme-subsystem',
+ $RE_NVME_SUBSYSTEM,
+ sub($entry) {
+ my $base = "/sys/class/nvme-subsystem/$entry";
+ my $subsys = file_read_firstline("$base/subsysnqn");
+ return if !defined($subsys) || $subsys ne $nqn;
+
+ PVE::File::dir_glob_foreach(
+ $base,
+ $RE_NVME_NAMESPACE,
+ sub($device) {
+ $devices->{"/dev/$device"} = $device if _block_device("/dev/$device");
+ PVE::File::dir_glob_foreach(
+ "/sys/class/block/$device",
+ $RE_NVME_PARTITION,
+ sub($partition) {
+ $devices->{"/dev/$partition"} = $partition
+ if _block_device("/dev/$partition");
+ },
+ );
+ },
+ );
+ },
+ );
+ return $devices;
+}
+
+sub _namespace_openers($nqn) {
+ my $devices = namespace_devices($nqn);
+ return [] if !$devices->%*;
+
+ my %openers;
+ PVE::File::dir_glob_foreach(
+ '/proc',
+ '[0-9]+',
+ sub($pid) {
+ PVE::File::dir_glob_foreach(
+ "/proc/$pid/fd",
+ '[0-9]+',
+ sub($fd) {
+ my $target = _link_target("/proc/$pid/fd/$fd");
+ return if !defined($target) || !exists($devices->{$target});
+ my $comm = eval { file_read_firstline("/proc/$pid/comm") } // 'unknown';
+ $openers{"$pid:$target"} = "$comm (PID $pid, $target)";
+ },
+ );
+ },
+ );
+
+ # Kernel consumers such as device-mapper do not necessarily keep a userspace
+ # file descriptor open, but expose their dependency in the holders directory.
+ for my $path (keys $devices->%*) {
+ my $device = $devices->{$path};
+ my $holders = "/sys/class/block/$device/holders";
+ PVE::File::dir_glob_foreach(
+ $holders,
+ '[^\\.].*',
+ sub($holder) {
+ $openers{"holder:$device:$holder"} = "$path held by $holder";
+ },
+ );
+ }
+
+ return [sort values %openers];
+}
+
+my sub portal_reachable($portal) {
+ return PVE::Network::tcp_ping($portal->{address}, $portal->{port}, 2) // 0;
+}
+
+# The number of portals with a live controller on their configured interface,
+# or, with $any_iface, on any interface: such a path carries I/O, even while
+# it waits to be moved to its configured interface.
+sub _live_portal_count($states, $portals, $any_iface = 0) {
+ my $live = 0;
+ for my $portal ($portals->@*) {
+ my $id = "$portal->{address}:$portal->{port}";
+ $live++ if grep {
+ $_->{state} eq 'live'
+ && ($any_iface || ($_->{host_iface} // '') eq $portal->{host_iface})
+ } ($states->{$id} // [])->@*;
+ }
+ return $live;
+}
+
+# Creates the controller of a portal in a child, giving up after 10 seconds.
+sub _connect_portal($scfg, $portal, $hostnqn, $hostid, $key) {
+ my ($options, $names) = _fabrics_options($scfg, $portal, $hostnqn, $hostid, $key);
+ return run_bounded(10, 'NVMe/TCP connect', sub { return _fabrics_connect($options, $names) });
+}
+
+my sub set_iopolicy($nqn, $policy) {
+ PVE::File::dir_glob_foreach(
+ '/sys/class/nvme-subsystem',
+ $RE_NVME_SUBSYSTEM,
+ sub($entry) {
+ my $base = "/sys/class/nvme-subsystem/$entry";
+ my $subsys = file_read_firstline("$base/subsysnqn");
+ return if !defined($subsys) || $subsys ne $nqn;
+ my $error = write_attribute("$base/iopolicy", "$policy\n");
+ die "unable to set NVMe multipath policy: $error\n" if defined($error);
+ },
+ );
+}
+
+# Applies the reconnect timeouts to connected controllers, which otherwise
+# keep the values of their connect. sysfs shows "off" for -1, and shows the
+# loss timeout rounded up to a multiple of the reconnect delay.
+my sub set_tunables($scfg) {
+ my $delay = $scfg->{'nvme-reconnect-delay'} // 2;
+ my $loss = $scfg->{'nvme-ctrl-loss-tmo'} // 600;
+ my $fast = $scfg->{'nvme-fast-io-fail-tmo'};
+ my $shown_loss = $loss < 0 ? 'off' : int(($loss + $delay - 1) / $delay) * $delay;
+ my $shown_fast = defined($fast) && $fast >= 0 ? $fast : 'off';
+
+ for my $controller (nvme_controllers($scfg->{subsysnqn})->@*) {
+ my $base = "/sys/class/nvme/$controller->{name}";
+ my $current_delay = file_read_firstline("$base/reconnect_delay") // next;
+ my $write = sub($attr, $value) {
+ my $error = write_attribute("$base/$attr", "$value\n") // return;
+ log_warn("cannot set $attr of NVMe controller '$controller->{name}': $error");
+ };
+ my $delay_changed = $current_delay ne "$delay";
+ $write->('reconnect_delay', $delay) if $delay_changed;
+ # the loss timeout is stored as a number of reconnects of the delay
+ $write->('ctrl_loss_tmo', $loss < 0 ? -1 : $loss)
+ if $delay_changed
+ || (file_read_firstline("$base/ctrl_loss_tmo") // '') ne "$shown_loss";
+ $write->('fast_io_fail_tmo', $shown_fast eq 'off' ? -1 : $fast)
+ if (file_read_firstline("$base/fast_io_fail_tmo") // '') ne "$shown_fast";
+ }
+}
+
+# Attempts to make a controller attribute readable by root only (a test seam).
+sub _restrict_attr($path) {
+ my @stat = stat($path);
+ if (!@stat) {
+ log_warn("cannot stat NVMe controller attribute '$path': $!") if $! != ENOENT;
+ return;
+ }
+ return if !($stat[2] & 077);
+ chmod(0600, $path) or log_warn("cannot restrict permissions of '$path': $!");
+}
+
+# The kernel creates the DH-HMAC-CHAP secret attributes world-readable.
+my sub restrict_secret_attrs($nqn) {
+ for my $controller (nvme_controllers($nqn)->@*) {
+ _restrict_attr("/sys/class/nvme/$controller->{name}/$_")
+ for qw(dhchap_secret dhchap_ctrl_secret);
+ }
+}
+
+# Best effort: a controller that went away since it was listed needs no scan,
+# and one that cannot be rescanned is a warning.
+my sub rescan_namespaces($nqn) {
+ for my $controller (nvme_controllers($nqn)->@*) {
+ next if $controller->{state} ne 'live';
+ my $path = "/sys/class/nvme/$controller->{name}/rescan_controller";
+ my $error = write_attribute($path, "1\n", 1) // next;
+ log_warn("cannot rescan NVMe controller '$controller->{name}': $error");
+ }
+}
+
+# Deletes a controller that is dead or that moved to another interface. The
+# activation reconnects or keeps the path either way, so this only warns.
+my sub disconnect_controller($controller) {
+ my $name = $controller->{name};
+ eval {
+ run_bounded(
+ 5,
+ "deleting NVMe controller '$name'",
+ sub { return _delete_controllers([$controller]) },
+ );
+ };
+ if (my $error = $@) {
+ rethrow_task_interrupt($error);
+ log_warn($error);
+ }
+}
+
+# Waits up to 10 seconds for a live controller of the portal on its
+# configured interface.
+my sub wait_for_live_path($nqn, $portal) {
+ for (my $attempt = 0; $attempt < 40; $attempt++) {
+ return 1 if _live_portal_count(controller_states($nqn), [$portal]);
+ _sleep(0.25);
+ }
+ return 0;
+}
+
+# Connects every missing or dead path on its configured interface. A path
+# that is connected on another interface moves make-before-break: its old
+# controller is only disconnected once the new one is live, so changing an
+# interface mapping never takes away a path that carries I/O.
+my sub connect_portals($scfg, $portals, $hostnqn, $hostid, $key) {
+ my $nqn = $scfg->{subsysnqn};
+ for my $portal ($portals->@*) {
+ my $id = "$portal->{address}:$portal->{port}";
+ my $iface = $portal->{host_iface};
+ my @controllers = (controller_states($nqn)->{$id} // [])->@*;
+ my ($configured) = grep { ($_->{host_iface} // '') eq $iface } @controllers;
+ my @moved = grep { ($_->{host_iface} // '') ne $iface } @controllers;
+
+ if (!$configured || $configured->{state} eq 'dead') {
+ disconnect_controller($configured) if $configured;
+ if (!portal_reachable($portal)) {
+ log_warn("NVMe/TCP portal '$id' is unreachable");
+ next;
+ }
+ eval { _connect_portal($scfg, $portal, $hostnqn, $hostid, $key) };
+ my $connect_error = $@;
+ # A failed or timed-out connect can still have created a controller.
+ eval { restrict_secret_attrs($nqn); 1 };
+ if (my $error = $@) {
+ rethrow_task_interrupt($error);
+ log_warn("cannot restrict NVMe controller attributes for '$nqn'");
+ }
+ rethrow_task_interrupt($connect_error) if $connect_error;
+ if ($connect_error) {
+ log_warn("$id: $connect_error");
+ next;
+ }
+ }
+ next if !@moved;
+ if (!wait_for_live_path($nqn, $portal)) {
+ log_warn("NVMe/TCP portal '$id' is not live on '$iface' yet; keeping its path on"
+ . " another interface");
+ next;
+ }
+ disconnect_controller($_) for @moved;
+ }
+}
+
+sub activate_storage($class, $storeid, $scfg, $cache = undef) {
+ $cache //= {};
+ my $nqn = $scfg->{subsysnqn};
+ # First, whatever the rest of the activation does, also for controllers
+ # connected by hand. Also try after each connect attempt below, including
+ # failures that may have created a controller.
+ restrict_secret_attrs($nqn);
+
+ my $multipath = file_read_firstline('/sys/module/nvme_core/parameters/multipath')
+ // die "the NVMe kernel modules are not loaded; load the nvme-tcp kernel module\n";
+ die "native NVMe multipath is disabled in the running kernel\n" if $multipath ne 'Y';
+
+ my $portals = _configured_portals($scfg);
+ verify_nvme_nqn($nqn);
+ _validate_local_ifaces($portals);
+ _assert_unique_target($storeid, $scfg);
+ my $hostnqn = file_read_firstline('/etc/nvme/hostnqn')
+ // die "missing /etc/nvme/hostnqn; the nvme-cli package generates it\n";
+ verify_nvme_nqn($hostnqn);
+ my $hostnqns = parse_nvme_host_nqns($scfg->{'nvme-host-nqns'});
+ die "local NVMe host NQN '$hostnqn' is missing from nvme-host-nqns\n"
+ if !grep { $_ eq $hostnqn } $hostnqns->@*;
+ my $policy = $scfg->{'nvme-iopolicy'} // 'round-robin';
+ my $force_reconcile = delete($cache->{'zfsnvme-force-reconcile'}->{$storeid}) // 0;
+ my $states = controller_states($nqn);
+ my $healthy = _live_portal_count($states, $portals);
+ my $usable = _live_portal_count($states, $portals, 1);
+ my $last = $slow_path_last{$storeid};
+ my $recent = defined($last) && _now() - $last < $slow_path_backoff;
+
+ # This method is called by the periodic storage status loop. Once every
+ # configured path is live, lifecycle operations keep the target converged,
+ # so no remote call is needed. A degraded storage retries the target at
+ # most once a minute: in between, one with a live path (also one waiting
+ # to move to its configured interface) is usable, and one without fails
+ # at once instead of waiting for a path again. Missing namespaces
+ # explicitly force the slow path from activate_volume().
+ if (!$force_reconcile && ($healthy == scalar($portals->@*) || $recent)) {
+ die "no live NVMe/TCP path for storage '$storeid'\n" if !$usable;
+ die "missing NVMe DH-HMAC-CHAP key\n"
+ if (file_read_firstline(secret_path($storeid)) // '') eq '';
+ set_iopolicy($nqn, $policy);
+ set_tunables($scfg);
+ return 1;
+ }
+
+ $slow_path_last{$storeid} = _now();
+ my $hostid = file_read_firstline('/etc/nvme/hostid')
+ // die "missing /etc/nvme/hostid; the nvme-cli package generates it\n";
+ die "invalid NVMe host ID\n"
+ if $hostid !~ $RE_NVMET_UUID || $hostid =~ $RE_ZERO_UUID;
+ my $key = get_secret($storeid);
+ for my $name (qw(keep-alive-tmo reconnect-delay ctrl-loss-tmo fast-io-fail-tmo nr-io-queues)) {
+ my $value = $scfg->{"nvme-$name"};
+ die "invalid NVMe connection parameter '$name'\n"
+ if defined($value) && $value !~ $RE_CONFIG_INT;
+ }
+ _nvmet_activate_target($storeid, $scfg, $portals, $hostnqns, $key);
+ connect_portals($scfg, $portals, $hostnqn, $hostid, $key);
+
+ # Existing controllers can be in the middle of their kernel reconnect delay
+ # after a target restart. Do not create duplicates, but give that recovery
+ # cycle enough time to complete before declaring the storage unavailable.
+ for (my $attempt = 0; $attempt < 60; $attempt++) {
+ $states = controller_states($nqn);
+ last if _live_portal_count($states, $portals, 1);
+ _sleep(0.25);
+ }
+ die "no live NVMe/TCP path for storage '$storeid'\n"
+ if !_live_portal_count($states, $portals, 1);
+ $healthy = _live_portal_count($states, $portals);
+ log_warn("storage '$storeid' is degraded: $healthy/" . scalar($portals->@*) . " paths live")
+ if $healthy < scalar($portals->@*);
+
+ set_iopolicy($nqn, $policy);
+ set_tunables($scfg);
+ return 1;
+}
+
+sub deactivate_storage($class, $storeid, $scfg, $cache = undef) {
+ my $nqn = $scfg->{subsysnqn};
+ my $openers = _namespace_openers($nqn);
+ die "refusing to disconnect NVMe storage '$storeid': namespace in use by "
+ . join(', ', $openers->@*) . "\n"
+ if $openers->@*;
+
+ # One child deletes every controller, bounded by 15 seconds in all. When it
+ # gives up, the controllers that are left stay connected and the
+ # deactivation fails.
+ my $controllers = nvme_controllers($nqn);
+ run_bounded(
+ 15,
+ "disconnecting NVMe subsystem '$nqn'",
+ sub { return _delete_controllers($controllers) },
+ ) if $controllers->@*;
+ return 1;
+}
+
+# Removing a storage definition never destroys data and never depends on the
+# target, like for every other storage type. Target cleanup is best effort
+# and only happens once the storage owns no volume. Other nodes keep their
+# connections until they are disconnected there.
+sub on_delete_hook($class, $storeid, $scfg) {
+ # First: an activation elsewhere that waits for the target lock then
+ # refuses instead of restoring what is removed below.
+ delete_secret($storeid);
+ eval { $class->deactivate_storage($storeid, $scfg) };
+ if (my $err = $@) {
+ rethrow_task_interrupt($err);
+ chomp($err);
+ log_warn("not disconnecting NVMe storage '$storeid': $err");
+ }
+ eval { _nvmet_delete_target($scfg) };
+ if (my $err = $@) {
+ rethrow_task_interrupt($err);
+ chomp($err);
+ log_warn("keeping the NVMe target configuration of storage '$storeid': $err");
+ }
+ return;
+}
+
+sub path($class, $scfg, $volname, $storeid, $snapname = undef) {
+ die "direct access to snapshots not implemented\n" if defined($snapname);
+ my ($vtype, $name, $vmid) = $class->parse_volname($volname);
+ my $uuid = _nvmet_volume_uuid($scfg, $name);
+ my $path = "/dev/disk/by-id/nvme-uuid.$uuid";
+ return ($path, $vmid, $vtype);
+}
+
+sub qemu_blockdev_options(
+ $class, $scfg, $storeid, $volname,
+ $machine_version = undef,
+ $options = undef,
+) {
+ die "direct access to snapshots not implemented\n" if $options->{'snapshot-name'};
+ my ($path) = $class->path($scfg, $volname, $storeid);
+ return { driver => 'host_device', filename => $path };
+}
+
+# Whether the local device behind a namespace link is the namespace (NSID,
+# UUID) of this subsystem, so a UUID claimed by another subsystem is never
+# handed to a guest.
+sub _nvmet_local_namespace_ok($path, $nqn, $nsid, $uuid) {
+ my $link = _link_target($path) // return 0;
+ my $device = (split m{/}, $link)[-1];
+ return 0 if !defined($device) || $device !~ $RE_NVME_NAMESPACE;
+ my $sys = "/sys/block/$device";
+ return 0 if lc(file_read_firstline("$sys/uuid") // '') ne lc($uuid);
+ return 0 if (file_read_firstline("$sys/nsid") // '') ne "$nsid";
+ return 0 if (file_read_firstline("$sys/device/subsysnqn") // '') ne $nqn;
+ return 1;
+}
+
+sub activate_volume(
+ $class, $storeid, $scfg, $volname,
+ $snapname = undef,
+ $cache = undef,
+ $hints = undef,
+) {
+ die "unable to activate snapshot from remote zfs storage\n" if $snapname;
+ my (undef, $name) = $class->parse_volname($volname);
+ my $nqn = nvmet_nqn($scfg);
+ my $dataset = nvmet_dataset($scfg, $name);
+
+ my ($inv, $cfs) = nvmet_read($scfg);
+ my ($nsid, $uuid) = _nvmet_owned_identity($inv, $nqn, $dataset);
+ my $state = _nvmet_export_state($cfs, $nqn, $nsid, $uuid, "/dev/zvol/$dataset");
+ my $path = "/dev/disk/by-id/nvme-uuid.$uuid";
+ my $ok = sub {
+ return _block_device($path) && _nvmet_local_namespace_ok($path, $nqn, $nsid, $uuid);
+ };
+ return 1 if $state eq 'present' && $ok->();
+
+ $cache //= {};
+ $cache->{'zfsnvme-force-reconcile'}->{$storeid} = 1;
+ $class->activate_storage($storeid, $scfg, $cache);
+ for (my $attempt = 0; $attempt < 40 && !_block_device($path); $attempt++) {
+ # a host can miss the change notice; rescanning is harmless
+ rescan_namespaces($nqn) if $attempt % 8 == 0;
+ _sleep(0.25);
+ }
+ die "NVMe namespace for '$volname' did not appear\n" if !_block_device($path);
+ die "NVMe namespace for '$volname' has an unexpected identity\n" if !$ok->();
+ return 1;
+}
+
+sub deactivate_volume(
+ $class, $storeid, $scfg, $volname, $snapname = undef, $cache = undef,
+) {
+ die "unable to deactivate snapshot from remote zfs storage\n" if $snapname;
+ return 1;
+}
+
+1;
next prev parent reply other threads:[~2026-10-06 8:54 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-05 0:26 [PATCH storage v3 0/4] add ZFS over NVMe/TCP storage plugin Joaquin Varela
2026-10-05 0:26 ` Joaquin Varela [this message]
2026-10-05 0:26 ` [PATCH storage v3 2/4] test: add zfsnvme plugin tests Joaquin Varela
2026-10-05 0:26 ` [PATCH storage v3 3/4] zfsnvme: fence target commands of abandoned transactions Joaquin Varela
2026-10-05 0:26 ` [PATCH storage v3 4/4] zfsnvme: wait up to 30 seconds for the shared storage lock Joaquin Varela
2026-10-05 0:26 ` [PATCH docs v3] storage: document ZFS over NVMe/TCP Joaquin Varela
2026-10-05 0:26 ` [PATCH manager v3] ui: storage: add ZFS over NVMe/TCP editor Joaquin Varela
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261005002609.571-2-joaquinvarela@neatech.ar \
--to=joaquinvarela@neatech.ar \
--cc=pve-devel@lists.proxmox.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox