* [PATCH storage v3 0/4] add ZFS over NVMe/TCP storage plugin
@ 2026-10-05 0:26 Joaquin Varela
2026-10-05 0:26 ` [PATCH storage v3 1/4] zfsnvme: " Joaquin Varela
` (5 more replies)
0 siblings, 6 replies; 7+ messages in thread
From: Joaquin Varela @ 2026-10-05 0:26 UTC (permalink / raw)
To: pve-devel
Hi,
this is v3 of the zfsnvme backend: ZFS volumes on a remote Linux
target, exported with the kernel's nvmet over NVMe/TCP and used through
native NVMe multipath and DH-HMAC-CHAP on the nodes.
The scope is one remote Linux ZFS/nvmet target and its integration with
the PVE shared-storage API. This series does not implement target high
availability, pool ownership between target machines, or target fencing.
PVE remains responsible for VM placement and node fencing. The namespace
identity and reconnect behavior are intended to allow an independently
managed target to recover, but a multi-head target has not been qualified.
v2: https://lore.proxmox.com/pve-devel/cover.1785636979.git.joaquinvarela@neatech.ar/
Changes since v2, following Max's review:
1. The plugin inherits from PVE::Storage::Plugin directly. The ZFS
helpers it needs are private copies; volume names and their parsing
are the same as in the other ZFS plugins.
2. PVE::Storage::LunCmd::NVMET is gone; the nvmet code lives in the
plugin with an nvmet_ prefix.
3. There is no remote helper anymore. ZFS/configfs parsing,
storage-configuration validation and operation planning happen locally,
in functions that are tested on their own. Optional patch 3/4 adds a
fixed target-side transport guard, described below.
The target state is read in one call: 'zfs get' on the pool and its
direct children, and one 'find' over the nvmet configfs tree that
lists its objects, prints the attributes the plugin uses with 'grep'
and reads host keys only as 'sha256sum' digests. The first such read
of an operation under the target lock starts with
'modprobe nvmet_tcp'. Other reads are 'zfs get' and 'zfs list' like
in the other ZFS plugins, 'test -b' while a new zvol appears, and
'cat /proc/mounts' when that first locked read fails, to decide
whether configfs needs to be mounted. Changes are quoted simple
commands, joined with '&&' so that a step only runs after the
previous one succeeded, plus redirections into configfs attributes
and a 'dd | tee' pipe that writes the key from standard input,
without bash-isms. An export starts with 'test -b' on its zvol, and
each publish link repeats two read-only checks right before it links
the subsystem to a port: that the subsystem denies unknown hosts and
that every configured host has its ACL. A pmxcfs domain lock per
target serializes changes within the cluster. Patch 3/4 adds the only
shell control flow, see below.
4. opendir/readdir and glob were replaced by
PVE::File::dir_glob_foreach, nested where a sysfs walk spans several
directory levels, because dir_glob_regex returns only the first
match; IO::Socket::IP was replaced by PVE::Network::tcp_ping, and
IO::File by PVE::SysFSTools::file_write.
Other changes:
- Each storage needs its own NVMe/TCP listeners (address and port) on
the target. nvmet answers a connect to a listener that does not
publish the subsystem yet with DNR, and the node then deletes the
controller regardless of the controller loss timeout. After a target
restart the plugin restores one storage at a time, so storages that
shared a listener could lose their controllers. Adding, updating and
activating a storage now refuse a portal whose listener another
zfsnvme storage uses, also a disabled one, and a target address that
another storage reaches through a different server value, since the
server value names the lock per target. IPv6 addresses are compared
in canonical form.
- Nodes create each missing path with a single write to
/dev/nvme-fabrics in a short-lived child process
(PVE::Tools::run_fork_with_timeout) instead of running
'nvme connect --config'. With the Debian trixie versions tested here,
nvme-cli 2.13-2 and libnvme 1.13-2, the invocation with an explicit
host NQN and host ID re-issues connects to existing fabrics controllers
of that host, and its sysfs scan replaces the configured DH-HMAC-CHAP
host secret with the secret of an existing controller. On a node with
another NVMe/TCP controller
of the same host NQN, such as a second zfsnvme storage on a different
target, a new path can then be attempted with the other target's key.
The same happens with nvme-cli 2.16 and libnvme 1.16.2; nvme-cli 3.1
reworked 'nvme connect --config' and did not reproduce this behavior
in our tests. The direct write also keeps the key off command lines
and avoids a temporary nvme-cli configuration file. The key still
lives in the PVE private key file and the kernel attributes described
below. It uses sysopen rather than PVE::SysFSTools::file_write,
because the kernel returns the instance of the new controller on a
read of the descriptor that wrote the options. Controllers are
deleted and rescanned through sysfs. nvme-cli stays a dependency
because it creates /etc/nvme/hostnqn and /etc/nvme/hostid.
- The plugin restricts the DH-HMAC-CHAP secret attributes of the node
controllers, and of the target host entries it writes the key to, to
mode 0600. The kernel creates them world-readable, so on the nodes
this is best effort; the docs describe the remaining exposure.
- A correction to the v2 cover letter: the "verified 45.7 GB mixed-I/O
soak" ran for only 741 seconds and was a short checksum exercise; it
is not evidence of cache-independent readback or of crash durability.
The patches:
1/4 adds the backend.
2/4 adds the tests: the plugin against mocked host interfaces, and an
emulated target that checks every command the plugin plans and
replays single and double faults during each operation.
3/4 is optional, and a question for you. ssh does not stop a remote
command when the client gives up, so a delayed command of an
abandoned operation can still change the target after the next
operation started. This patch wraps every call made inside a locked
transaction, reads included, in a small, fixed POSIX sh guard: a
flock on a file below /run on the target, a per-transaction token
that every later command of the transaction checks, and a
compare-and-set of that token when a transaction starts. Read-only
calls outside a transaction do not use it. It excludes stale
control-plane commands; it does not fence hosts or guest I/O. It
adds two SSH round trips (owner observation and claim) to every
locked operation. Without it, the backend relies on the cluster lock
like ZFS over iSCSI does. Would you prefer to keep it, a smaller
variant with only the flock and the token check, or to drop it? 4/4
applies and passes its tests with or without 3/4.
4/4 is also a question: it overrides cluster_lock_storage so that the
shared storage lock waits up to 30 seconds instead of 10 when the
caller passes no timeout, since a target change over SSH can hold
it for several seconds. Explicit timeouts are unchanged. Would you
rather have this in the plugin, as a parameter in the core, or
should I reduce the SSH round trips per operation instead?
The series is much larger than v2: the plugin has about 3,300 lines,
because the local parsing and planning replace the remote helper, and
the tests about 7,800 (two test files and two fixtures). I can split
1/4 into the schema and pvesm, the target planning, the node connect
and the volume operations, and split or trim the test patch, whatever
is easier to review.
Testing: every patch passes 'make test' on its own, unprivileged and
with stock libpve-common-perl, and the Perl files are formatted with
proxmox-perltidy. On a nested two-node PVE 9 cluster (kernel 7.0.14)
with an Ubuntu 24.04 target (kernel 6.8), I ran the checks below with
a package built from the code of this series. The only exception is
the reproduction of the shared listener problem, which deliberately
used the code before the listener check:
- activation and DH-HMAC-CHAP on both nodes and both paths through the
new connect path: one read-write open of /dev/nvme-fabrics per path,
no nvme-cli process;
- allocation, snapshot, rollback, resize, template, linked clone and
free on both nodes;
- a node that already has another NVMe/TCP controller with the same host
NQN and a different key: the plugin connects with its own key, while a
single 'nvme connect --config' in the same situation issued five
connect writes and failed authentication;
- live migration in both directions under guest writes, with readback;
- the loss and restore of one path, and a 15-second loss of both paths,
under guest writes, without guest I/O errors;
- a target reboot under guest writes: the plugin rebuilt the target
configuration and the committed data verified;
- stopping an activation task while a connect was in flight, with stock
libpve-common-perl: no process was left behind and none of the six
cancellations left a controller; a controller created out of band, as
a cancelled connect can leave behind, is adopted by the next
activation;
- the controller loss behavior the docs describe: with a 20-second
controller loss timeout, the controllers and the multipath device were
removed and an open descriptor stayed dead, while -1 with a 5-second
fast I/O fail timeout failed I/O after about 15 seconds and kept the
device, which worked again once the paths were back;
- two storages that share a listener: with the code before the
listener check, after their configuration was removed from the
target, as a target restart does, nvmet rejected the reconnects of
the storage restored second with DNR, both nodes deleted its
controllers, and an open descriptor on its multipath device stayed
dead; with this series, adding the second storage on the same
listener is refused, and with a listener per storage both recovered
from the same removal, in either order, without deleting a
controller, while an open O_DIRECT descriptor kept reading the
written data, with one read blocked for about 5 seconds in each of
the two runs.
This is lab qualification, not hardware qualification.
I also saw Dietmar's RFC "add guided remote storage setup and SAN
visibility", which adds NVMe-oF transport and path status helpers to
Diskmanage. Once it is applied, the plugin can use them instead of its
own sysfs parsing:
https://lore.proxmox.com/pve-devel/20260731102156.3947857-1-dietmar@proxmox.com/
The documentation and the editor follow as replies to this cover
letter, as separate patches for pve-docs and pve-manager.
Joaquin Varela (4):
zfsnvme: add ZFS over NVMe/TCP storage plugin
test: add zfsnvme plugin tests
zfsnvme: fence target commands of abandoned transactions
zfsnvme: wait up to 30 seconds for the shared storage lock
debian/control | 1 +
src/PVE/CLI/pvesm.pm | 21 +-
src/PVE/Storage.pm | 2 +
src/PVE/Storage/Makefile | 1 +
src/PVE/Storage/Plugin.pm | 2 +-
src/PVE/Storage/ZFSNVMePlugin.pm | 3319 +++++++++++
src/test/run_plugin_tests.pl | 2 +
.../zfsnvme_fixtures/configfs_snapshot.txt | 38 +
src/test/zfsnvme_fixtures/zfs_inventory.txt | 15 +
src/test/zfsnvme_target_test.pm | 5299 +++++++++++++++++
src/test/zfsnvme_test.pm | 2429 ++++++++
11 files changed, 11126 insertions(+), 3 deletions(-)
create mode 100644 src/PVE/Storage/ZFSNVMePlugin.pm
create mode 100644 src/test/zfsnvme_fixtures/configfs_snapshot.txt
create mode 100644 src/test/zfsnvme_fixtures/zfs_inventory.txt
create mode 100644 src/test/zfsnvme_target_test.pm
create mode 100644 src/test/zfsnvme_test.pm
^ permalink raw reply [flat|nested] 7+ messages in thread
* [PATCH storage v3 1/4] zfsnvme: add ZFS over NVMe/TCP storage plugin
2026-10-05 0:26 [PATCH storage v3 0/4] add ZFS over NVMe/TCP storage plugin Joaquin Varela
@ 2026-10-05 0:26 ` Joaquin Varela
2026-10-05 0:26 ` [PATCH storage v3 2/4] test: add zfsnvme plugin tests Joaquin Varela
` (4 subsequent siblings)
5 siblings, 0 replies; 7+ messages in thread
From: Joaquin Varela @ 2026-10-05 0:26 UTC (permalink / raw)
To: pve-devel
Add the shared storage type 'zfsnvme': ZFS zvols on a remote Linux
target, exported with the kernel NVMe target (nvmet, configured through
configfs) over NVMe/TCP, and consumed by the cluster nodes with native
NVMe multipath and DH-HMAC-CHAP authentication.
The ZFS dataset is the source of truth for a namespace: the subsystem
NQN, NSID and UUID of every volume are stored in the ZFS user properties
'proxmox:nvme-subsys', 'proxmox:nvme-nsid' and 'proxmox:nvme-uuid', and
the pool keeps the highest NSID handed out in 'proxmox:nvme-last-nsid'.
The nodes find a namespace by its UUID. Everything else on the target
(nvmet_tcp module, configfs mount, subsystem, namespaces, ports, host
ACLs) is derived state that an activation rebuilds after a target
restart.
All parsing, validation and planning happen locally. The target state
is read with one 'zfs get' and one 'find' over configfs; other reads
are 'zfs get' and 'zfs list', 'test -b' while a new zvol appears, and
'cat /proc/mounts' to decide whether configfs needs to be mounted. The
target runs quoted simple commands over SSH, joined with '&&' where a
step depends on the previous one: an export starts with 'test -b' on
its zvol, and each publish link repeats two read-only checks right
before it links the subsystem to a port. There are no scripts, loops
or variables on the target. As with ZFS over iSCSI, the SSH key is
/etc/pve/priv/zfs/<server>_id_rsa. Target changes are serialized
cluster-wide by a pmxcfs domain lock per target server, and an
operation whose SSH connection broke is reported as unknown and
repaired by the next activation rather than compensated from a read
that could overtake it.
Each storage needs its own NVMe/TCP listener (address family, address
and port) on the target. nvmet answers a connect to a listener that
does not publish the subsystem yet with DNR, and the host then deletes
the controller whatever its loss timeout, so after a target restart the
storage published second on a shared listener would lose its paths.
Adding, updating and activating a storage therefore refuse a portal
whose listener another zfsnvme storage uses, disabled ones included,
and a target address that another storage reaches through a different
'server' value, since that value names the target lock. A data address
thus identifies one target for all zfsnvme storages of the cluster.
The DH-HMAC-CHAP key is a sensitive property, stored like the other
storage secrets in /etc/pve/priv/storage/<storeid>.nvme-dhchap. It
reaches the target through stdin, and the plugin restricts the
dhchap_key and dhchap_ctrl_key attributes there to 0600 whenever it
writes the key. On the nodes, controllers are created with exactly one
write to /dev/nvme-fabrics in a bounded child (run_fork_with_timeout),
so the key never appears on a command line, and the secret attributes of
a new controller are restricted before the child returns. Rescans and
disconnects are sysfs writes; nvme-cli is not run.
Keep-alive, reconnect delay, controller loss and optional fast I/O
failure timeouts are configurable, with defaults that prefer waiting for
the target over failing guest I/O (ctrl_loss_tmo 600, fast_io_fail_tmo
unset).
Outside the plugin: register the type in PVE::Storage and the Makefile,
add it to @SHARED_STORAGE in PVE::Storage::Plugin, let 'pvesm add' and
'pvesm set' read --dhchap-key from a file, and depend on nvme-cli, which
creates /etc/nvme/hostnqn and /etc/nvme/hostid and remains the
administration tool, like the client packages of the other shared types.
Signed-off-by: Joaquin Varela <joaquinvarela@neatech.ar>
---
debian/control | 1 +
src/PVE/CLI/pvesm.pm | 21 +-
src/PVE/Storage.pm | 2 +
src/PVE/Storage/Makefile | 1 +
src/PVE/Storage/Plugin.pm | 2 +-
src/PVE/Storage/ZFSNVMePlugin.pm | 3177 ++++++++++++++++++++++++++++++
6 files changed, 3201 insertions(+), 3 deletions(-)
create mode 100644 src/PVE/Storage/ZFSNVMePlugin.pm
diff --git a/debian/control b/debian/control
index 850cd57c..ab46e6b6 100644
--- a/debian/control
+++ b/debian/control
@@ -43,6 +43,7 @@ Depends: bzip2,
lvm2,
lzop,
nfs-common,
+ nvme-cli,
proxmox-backup-client (>= 2.1.10~),
proxmox-backup-file-restore,
pve-cluster (>= 5.0-32),
diff --git a/src/PVE/CLI/pvesm.pm b/src/PVE/CLI/pvesm.pm
index 1aaa9f3c..3c68ffb5 100755
--- a/src/PVE/CLI/pvesm.pm
+++ b/src/PVE/CLI/pvesm.pm
@@ -78,12 +78,29 @@ sub param_mapping {
},
};
+ my $dhchap_key_map = {
+ name => 'dhchap-key',
+ desc => 'a file containing the NVMe DH-HMAC-CHAP key',
+ func => sub {
+ my ($value) = @_;
+ # never repeat a key passed in place of the file name
+ die "dhchap-key expects the path of a file containing the key\n"
+ if $value =~ /^DHHC-1:/i || -d $value;
+ my ($key) = split(/\n/, PVE::Tools::file_get_contents($value), 2);
+ $key = PVE::Tools::trim($key // '');
+ die "DH-HMAC-CHAP key file '$value' is empty\n" if $key eq '';
+ return $key;
+ },
+ };
+
my $mapping = {
'cifsscan' => [$password_map],
'cifs' => [$password_map],
'pbs' => [$password_map],
- 'create' => [$password_map, $enc_key_map, $master_key_map, $keyring_map],
- 'update' => [$password_map, $enc_key_map, $master_key_map, $keyring_map],
+ 'create' =>
+ [$password_map, $enc_key_map, $master_key_map, $keyring_map, $dhchap_key_map],
+ 'update' =>
+ [$password_map, $enc_key_map, $master_key_map, $keyring_map, $dhchap_key_map],
};
return $mapping->{$name};
}
diff --git a/src/PVE/Storage.pm b/src/PVE/Storage.pm
index fc3db812..2e5a8031 100755
--- a/src/PVE/Storage.pm
+++ b/src/PVE/Storage.pm
@@ -36,6 +36,7 @@ use PVE::Storage::CephFSPlugin;
use PVE::Storage::ISCSIDirectPlugin;
use PVE::Storage::ZFSPoolPlugin;
use PVE::Storage::ZFSPlugin;
+use PVE::Storage::ZFSNVMePlugin;
use PVE::Storage::PBSPlugin;
use PVE::Storage::BTRFSPlugin;
use PVE::Storage::ESXiPlugin;
@@ -61,6 +62,7 @@ PVE::Storage::CephFSPlugin->register();
PVE::Storage::ISCSIDirectPlugin->register();
PVE::Storage::ZFSPoolPlugin->register();
PVE::Storage::ZFSPlugin->register();
+PVE::Storage::ZFSNVMePlugin->register();
PVE::Storage::PBSPlugin->register();
PVE::Storage::BTRFSPlugin->register();
PVE::Storage::ESXiPlugin->register();
diff --git a/src/PVE/Storage/Makefile b/src/PVE/Storage/Makefile
index a67dc25f..d1cbfe29 100644
--- a/src/PVE/Storage/Makefile
+++ b/src/PVE/Storage/Makefile
@@ -11,6 +11,7 @@ SOURCES= \
ISCSIDirectPlugin.pm \
ZFSPoolPlugin.pm \
ZFSPlugin.pm \
+ ZFSNVMePlugin.pm \
PBSPlugin.pm \
BTRFSPlugin.pm \
LvmThinPlugin.pm \
diff --git a/src/PVE/Storage/Plugin.pm b/src/PVE/Storage/Plugin.pm
index a9e17513..a5aedc2b 100644
--- a/src/PVE/Storage/Plugin.pm
+++ b/src/PVE/Storage/Plugin.pm
@@ -35,7 +35,7 @@ our @COMMON_TAR_FLAGS = qw(
);
our @SHARED_STORAGE = (
- 'iscsi', 'nfs', 'cifs', 'rbd', 'cephfs', 'iscsidirect', 'zfs', 'drbd', 'pbs',
+ 'iscsi', 'nfs', 'cifs', 'rbd', 'cephfs', 'iscsidirect', 'zfs', 'zfsnvme', 'drbd', 'pbs',
);
our $QCOW2_PREALLOCATION = {
diff --git a/src/PVE/Storage/ZFSNVMePlugin.pm b/src/PVE/Storage/ZFSNVMePlugin.pm
new file mode 100644
index 00000000..127680fb
--- /dev/null
+++ b/src/PVE/Storage/ZFSNVMePlugin.pm
@@ -0,0 +1,3177 @@
+package PVE::Storage::ZFSNVMePlugin;
+
+use v5.36;
+
+use Compress::Zlib qw(crc32);
+use Digest::SHA qw(sha256_hex);
+use Errno qw(ENOENT);
+use Fcntl qw(O_NOFOLLOW O_RDONLY O_RDWR S_ISCHR);
+use File::Path qw(make_path);
+use List::Util qw(max);
+use MIME::Base64 qw(decode_base64);
+use POSIX qw(SIG_BLOCK SIG_SETMASK SIGHUP SIGINT SIGQUIT SIGTERM sigprocmask);
+use Socket qw(AF_INET6 inet_ntop inet_pton);
+use Time::HiRes;
+
+use PVE::Cluster;
+use PVE::File;
+use PVE::JSONSchema;
+use PVE::Network;
+use PVE::RESTEnvironment qw(log_warn);
+use PVE::RPCEnvironment;
+use PVE::SysFSTools;
+use PVE::Tools qw(file_read_firstline file_set_contents run_command trim);
+
+use base qw(PVE::Storage::Plugin);
+
+# ZFS zvols on a remote Linux target, exported through the kernel NVMe target
+# (nvmet, configured through configfs) over NVMe/TCP and consumed with native
+# NVMe multipath.
+#
+# The ZFS dataset is the source of truth for a namespace: its owner NQN, NSID
+# and UUID are ZFS user properties. Everything else on the target is derived
+# state that the plugin rebuilds after a target restart.
+#
+# Storage parsing, validation and planning happen locally. SSH carries quoted
+# commands, sometimes joined with `&&`, to read or change ZFS/configfs state.
+# A pmxcfs domain lock per target serializes the changes within the cluster,
+# held like the storage lock of the other shared storage types.
+
+# ---------------------------------------------------------------------------
+# Constants and regular expressions
+# ---------------------------------------------------------------------------
+
+# LogLevel=ERROR keeps a login banner out of the first stderr line, which is
+# the error text of a failed call.
+my @ssh_cmd = (
+ '/usr/bin/ssh', '-o', 'BatchMode=yes', '-o', 'ConnectTimeout=10', '-o', 'LogLevel=ERROR',
+);
+my $id_rsa_path = '/etc/pve/priv/zfs';
+my $secret_dir = '/etc/pve/priv/storage';
+my $max_paths = 16;
+my $max_hosts = 64;
+my $slow_path_backoff = 60;
+
+my $nvmet_root = '/sys/kernel/config/nvmet';
+my $nvmet_model = 'Proxmox ZFS NVMe';
+my $nvmet_marker = 'ZFSNVME-CONFIGFS';
+my $nvmet_max_nsid = 0xfffffffe;
+my $nvmet_max_command = 65536;
+# Seconds to wait for the target lock while another node changes the target;
+# a destroy or rollback can hold it for several seconds.
+my $nvmet_lock_wait = 30;
+# On the pool dataset: the highest NSID handed out for a volume of the pool.
+my $nvmet_last_nsid = 'proxmox:nvme-last-nsid';
+
+my @nvmet_zfs_props =
+ ('type', 'proxmox:nvme-subsys', 'proxmox:nvme-nsid', 'proxmox:nvme-uuid', $nvmet_last_nsid);
+my %nvmet_zfs_keys = (
+ type => 'type',
+ 'proxmox:nvme-subsys' => 'nqn',
+ 'proxmox:nvme-nsid' => 'nsid',
+ 'proxmox:nvme-uuid' => 'uuid',
+ $nvmet_last_nsid => 'last_nsid',
+);
+my @nvmet_cfs_attrs = qw(
+ enable device_path device_uuid buffered_io
+ attr_model attr_serial attr_allow_any_host
+ addr_trtype addr_adrfam addr_traddr addr_trsvcid
+);
+# Every command the plugin runs on the target, besides the `dd | tee` pipe of
+# a key write. The configfs read runs find through env to set its locale.
+my %nvmet_commands = map { $_ => 1 } qw(
+ cat chmod env grep ln mkdir modprobe mount printf rm rmdir test zfs
+);
+
+my $RE_NQN = qr{
+ \A
+ nqn \.
+ [A-Za-z0-9] [A-Za-z0-9.-]*
+ :
+ [A-Za-z0-9] [A-Za-z0-9._:-]*
+ \z
+}nxx;
+my $RE_IPV4_PORTAL = qr{
+ \A
+ (?<address> [^:]+)
+ (?: : (?<port> [0-9]+))?
+ \z
+}nxx;
+my $RE_IPV6_PORTAL = qr{
+ \A
+ \[ (?<address> [^\]]+) \]
+ (?: : (?<port> [0-9]+))?
+ \z
+}nxx;
+# A canonical IPv4-mapped IPv6 address.
+my $RE_IPV4_MAPPED = qr{\A ::ffff: (?<address> [0-9.]+) \z}nxx;
+my $RE_HOST_IFACE = qr{\A [A-Za-z0-9_.-]+ \z}nxx;
+my $RE_DHCHAP_KEY = qr{
+ \A DHHC-1 : (?<hash> 0[0-3]) : (?<secret> [A-Za-z0-9+/]+ ={0,2}) : \z
+}nxx;
+# A stopped PVE task: the worker's signal handler dies with the first text,
+# run_command adds the command context, and run_fork_with_timeout reports
+# the signal in its own words.
+my $RE_TASK_INTERRUPT = qr{
+ \A (?: (?: command \x20 ' .* ' \x20 failed: \x20 )? received \x20 interrupt
+ | interrupted \x20 by \x20 unexpected \x20 signal ) \n \z
+}nsxx;
+my $RE_COMMAND_TIMEOUT = qr{
+ \A (?: command \x20 ' .* ' \x20 failed: \x20 )? got \x20 timeout \n \z
+}nsxx;
+my $RE_COMMAND_EXIT =
+ qr{\A command \x20 ' .* ' \x20 failed: \x20 exit \x20 code \x20 (?<code> [0-9]+) \n \z}nsxx;
+my $RE_FABRICS_VALUE = qr{\A [^\s,\x00]+ \z}nxx;
+my $RE_FABRICS_RESULT = qr{\A instance=(?<instance> [0-9]+) ,cntlid= [0-9]+ \n \z}nxx;
+my $RE_ZERO_UUID = qr{\A 0{8} (?: -0{4}){3} -0{12} \z}nxx;
+my $RE_CONFIG_INT = qr{\A -?[0-9]+ \z}nxx;
+my $RE_NVME_CONTROLLER = qr{\A nvme [0-9]+ \z}nxx;
+my $RE_NVME_SUBSYSTEM = qr{\A nvme-subsys [0-9]+ \z}nxx;
+my $RE_NVME_NAMESPACE = qr{\A nvme [0-9]+ n [0-9]+ \z}nxx;
+my $RE_NVME_PARTITION = qr{\A nvme [0-9]+ n [0-9]+ p [0-9]+ \z}nxx;
+my $RE_TRADDR = qr{(?: \A | ,) traddr=(?<value>[^,]+)}nxx;
+my $RE_TRSVCID = qr{(?: \A | ,) trsvcid=(?<value>[^,]+)}nxx;
+my $RE_HOST_IFACE_ADDRESS = qr{(?: \A | ,) host_iface=(?<value>[^,]+)}nxx;
+my $RE_ZVOL_OWNER = qr{^ (?:vm|base|subvol|basevol)- (?<owner>\d+) - \S+ $}nxx;
+my $RE_VOLNAME = qr{
+ ^
+ (?: (?<base> (?:base|basevol)- (?<base_vmid>\d+) - \S+) /)?
+ (?<name> (?<type>base|basevol|vm|subvol)- (?<vmid>\d+) - \S+)
+ $
+}nxx;
+my $RE_BASE_SNAPSHOT = qr{^ (?<base>\S+) \@__base__ $}nxx;
+my $RE_UNSIGNED_INTEGER = qr{^ (?<value>\d+) $}nxx;
+
+my $RE_NVMET_UUID = qr{\A [0-9a-fA-F]{8} (?: - [0-9a-fA-F]{4}){3} - [0-9a-fA-F]{12} \z}nxx;
+my $RE_NVMET_NSID = qr{\A [1-9] [0-9]* \z}nxx;
+my $RE_NVMET_POOL = qr{\A [A-Za-z0-9] [A-Za-z0-9_.:/-]* \z}nxx;
+my $RE_NVMET_DATASET_NAME = qr{\A [A-Za-z0-9] [A-Za-z0-9_.:-]* \z}nxx;
+my $RE_NVMET_SNAPSHOT = qr{\A [A-Za-z0-9_.:-]+ \z}nxx;
+my $RE_NVMET_TEMPLATE_ZVOL = qr{/ (?<type> vm | base) - (?<suffix> [^/]+) \z}nxx;
+my $RE_NVMET_INHERITED = qr{\A inherited \x20 from \x20 (?<source> .+) \z}nxx;
+# Errors after which the target state is not known: an interrupted task, or a
+# change whose command may still be running on the target.
+my $RE_NVMET_ABANDONED = qr{
+ (?: \A received \x20 interrupt | ; \x20 the \x20 target \x20 state \x20 is \x20 unknown) \n \z
+}nxx;
+
+# Lines of the configfs read (nvmet_configfs_step below): `D <dir>`,
+# `L <link>`, `<file>:<value>` from grep, and `<sha256> <file>` from
+# sha256sum. Attribute names anchor the split, so an NQN containing ':' is
+# never ambiguous.
+my $cfs_root = quotemeta($nvmet_root);
+my $RE_CFS_TOP = qr{\A D \x20 $cfs_root (?: / (?<top> hosts | ports | subsystems))? \z}nxx;
+my $RE_CFS_SUBSYS_DIR = qr{\A D \x20 $cfs_root /subsystems/ (?<nqn> [^/]+) \z}nxx;
+my $RE_CFS_SUBSYS_GROUP = qr{
+ \A D \x20 $cfs_root /subsystems/ [^/]+ / (?: namespaces | allowed_hosts) \z
+}nxx;
+my $RE_CFS_NS_DIR = qr{
+ \A D \x20 $cfs_root /subsystems/ (?<nqn> [^/]+) /namespaces/ (?<nsid> [0-9]+) \z
+}nxx;
+my $RE_CFS_PORT_DIR = qr{\A D \x20 $cfs_root /ports/ (?<port> [0-9]+) \z}nxx;
+my $RE_CFS_PORT_GROUP = qr{\A D \x20 $cfs_root /ports/ [0-9]+ /subsystems \z}nxx;
+my $RE_CFS_HOST_DIR = qr{\A D \x20 $cfs_root /hosts/ (?<host> [^/]+) \z}nxx;
+my $RE_CFS_PORT_LINK = qr{
+ \A L \x20 $cfs_root /ports/ (?<port> [0-9]+) /subsystems/ (?<nqn> [^/]+) \z
+}nxx;
+my $RE_CFS_ACL_LINK = qr{
+ \A L \x20 $cfs_root /subsystems/ (?<nqn> [^/]+) /allowed_hosts/ (?<host> [^/]+) \z
+}nxx;
+my $RE_CFS_OTHER_ENTRY = qr{\A [DL] \x20 $cfs_root /}nxx;
+my $RE_CFS_NS_ATTR = qr{
+ \A $cfs_root /subsystems/ (?<nqn> [^/]+) /namespaces/ (?<nsid> [0-9]+)
+ / (?<attr> enable | device_path | device_uuid | buffered_io) : (?<value> .*) \z
+}nxx;
+my $RE_CFS_SUBSYS_ATTR = qr{
+ \A $cfs_root /subsystems/ (?<nqn> [^/]+)
+ / (?<attr> attr_model | attr_serial | attr_allow_any_host) : (?<value> .*) \z
+}nxx;
+my $RE_CFS_PORT_ATTR = qr{
+ \A $cfs_root /ports/ (?<port> [0-9]+)
+ / (?<attr> addr_trtype | addr_adrfam | addr_traddr | addr_trsvcid) : (?<value> .*) \z
+}nxx;
+my $RE_CFS_KEY_DIGEST = qr{
+ \A (?<sha> [0-9a-f]{64}) \x20 [\x20*] $cfs_root /hosts/ (?<host> [^/]+) /dhchap_key \z
+}nxx;
+my $RE_CFS_OTHER_ATTR = qr{\A $cfs_root /}nxx;
+
+# ---------------------------------------------------------------------------
+# Target addressing and validated names
+# ---------------------------------------------------------------------------
+
+my sub nvmet_server($scfg) {
+ return $scfg->{server};
+}
+
+my sub nvmet_ssh_key($scfg) {
+ return "$id_rsa_path/" . nvmet_server($scfg) . '_id_rsa';
+}
+
+my sub nvmet_lock_id($scfg) {
+ my $server = nvmet_server($scfg) // '';
+ return 'zfsnvme-' . ($server =~ s/[^A-Za-z0-9.-]/_/gr);
+}
+
+my sub secret_path($storeid) {
+ return "$secret_dir/$storeid.nvme-dhchap";
+}
+
+my sub nvmet_nqn($scfg) {
+ my $nqn = $scfg->{subsysnqn};
+ die "invalid NVMe subsystem NQN\n" if !defined($nqn) || !verify_nvme_nqn($nqn, 1);
+ return $nqn;
+}
+
+my sub nvmet_pool($scfg) {
+ my $pool = $scfg->{pool};
+ die "invalid ZFS pool name\n" if !defined($pool) || $pool !~ $RE_NVMET_POOL;
+ return $pool;
+}
+
+my sub nvmet_dataset($scfg, $name) {
+ die "invalid ZFS volume name\n" if !defined($name) || $name !~ $RE_NVMET_DATASET_NAME;
+ return nvmet_pool($scfg) . "/$name";
+}
+
+# zfs destroy reads '%' and ',' in a snapshot name as a range and a list.
+my sub nvmet_snapshot($scfg, $name, $snap) {
+ die "invalid snapshot name\n" if !defined($snap) || $snap !~ $RE_NVMET_SNAPSHOT;
+ return nvmet_dataset($scfg, $name) . "\@$snap";
+}
+
+my sub nvmet_valid_nsid($value) {
+ return defined($value) && $value =~ $RE_NVMET_NSID && $value <= $nvmet_max_nsid;
+}
+
+# IPv6 addresses have many spellings; compare and store only the canonical one.
+my sub nvmet_canonical_address($address) {
+ return $address if !defined($address) || index($address, ':') < 0;
+ my $packed = inet_pton(AF_INET6, $address) // return $address;
+ return inet_ntop(AF_INET6, $packed);
+}
+
+# The address family and address that a portal of parse_nvme_portals reaches:
+# an IPv4-mapped IPv6 address reaches the IPv4 address.
+my sub nvmet_listener_address($portal) {
+ return ('ipv4', $+{address})
+ if $portal->{family} eq 'ipv6' && $portal->{address} =~ $RE_IPV4_MAPPED;
+ return $portal->@{qw(family address)};
+}
+
+# Target-provided names only appear in messages when they are well-formed.
+my sub nvmet_label($name) {
+ return $name =~ $RE_NQN ? " '$name'" : '';
+}
+
+# The other name of a zvol that a template conversion renames: vm-* and base-*.
+my sub nvmet_template_twin($name) {
+ return '' if $name !~ $RE_NVMET_TEMPLATE_ZVOL;
+ my $twin = ($+{type} eq 'vm' ? 'base' : 'vm') . "-$+{suffix}";
+ return substr($name, 0, $-[0]) . "/$twin";
+}
+
+# The timeout of a call that changes ZFS: ZFS waits for transaction groups,
+# which a busy pool can take minutes to sync.
+my sub nvmet_long_timeout() {
+ return PVE::RPCEnvironment->is_worker() ? 60 * 60 : 60;
+}
+
+# ---------------------------------------------------------------------------
+# Remote steps
+#
+# A step is an argv array, a write `{ write => $path, value => $value }` or a
+# key write `{ key => [$path, ...] }` fed from stdin. A unit is a list of steps
+# that must stay in one call.
+# ---------------------------------------------------------------------------
+
+my sub nvmet_subsys_path($nqn) {
+ return "$nvmet_root/subsystems/$nqn";
+}
+
+my sub nvmet_ns_path($nqn, $nsid) {
+ return "$nvmet_root/subsystems/$nqn/namespaces/$nsid";
+}
+
+my sub nvmet_port_path($id) {
+ return "$nvmet_root/ports/$id";
+}
+
+my sub nvmet_host_path($hostnqn) {
+ return "$nvmet_root/hosts/$hostnqn";
+}
+
+my sub nvmet_write($path, $value) {
+ return { write => $path, value => $value };
+}
+
+# Parts of a parsed configfs state; missing parts read as empty.
+my sub nvmet_namespaces($cfs, $nqn) {
+ my $subsys = $cfs->{subsystems}->{$nqn};
+ return ($subsys ? $subsys->{namespaces} : undef) // {};
+}
+
+my sub nvmet_port_linked($cfs, $id, $nqn) {
+ my $port = $cfs->{ports}->{$id};
+ return $port && ($port->{links} // {})->{$nqn};
+}
+
+my sub nvmet_linked_hosts($cfs) {
+ my %linked;
+ for my $subsys (values $cfs->{subsystems}->%*) {
+ $linked{$_} = 1 for keys(($subsys->{acl} // {})->%*);
+ }
+ return \%linked;
+}
+
+# A subsystem whose creation by this plugin stopped after its model was
+# written, before its serial: nobody else writes our model, and it exports,
+# allows and publishes nothing, so completing it takes nothing from anybody.
+# A subsystem with the kernel's default model may belong to another tool.
+my sub nvmet_unfinished_subsystem($cfs, $nqn) {
+ my $subsys = $cfs->{subsystems}->{$nqn};
+ return 0 if !$subsys;
+ return
+ $subsys->{attr_model} eq $nvmet_model
+ && !($subsys->{namespaces} // {})->%*
+ && !($subsys->{acl} // {})->%*
+ && !grep { nvmet_port_linked($cfs, $_, $nqn) } keys $cfs->{ports}->%*;
+}
+
+# Same write order as the kernel requires: the device before enabling it.
+my sub nvmet_build_unit($nqn, $nsid, $uuid, $dev) {
+ my $ns = nvmet_ns_path($nqn, $nsid);
+ return [
+ ['test', '-b', $dev],
+ ['mkdir', $ns],
+ nvmet_write("$ns/device_path", $dev),
+ nvmet_write("$ns/device_uuid", $uuid),
+ nvmet_write("$ns/buffered_io", 0),
+ nvmet_write("$ns/enable", 1),
+ ];
+}
+
+my sub nvmet_enable_unit($nqn, $nsid, $dev) {
+ return [['test', '-b', $dev], nvmet_write(nvmet_ns_path($nqn, $nsid) . '/enable', 1)];
+}
+
+my sub nvmet_unexport_unit($nqn, $nsid, $enabled) {
+ my $ns = nvmet_ns_path($nqn, $nsid);
+ return [($enabled ? (nvmet_write("$ns/enable", 0)) : ()), ['rmdir', $ns]];
+}
+
+my sub nvmet_identity($nqn, $nsid, $uuid) {
+ return ("proxmox:nvme-subsys=$nqn", "proxmox:nvme-nsid=$nsid", "proxmox:nvme-uuid=$uuid");
+}
+
+# The pool and its direct children with their identity properties.
+my sub nvmet_zfs_inventory_step($pool) {
+ return [
+ 'zfs',
+ 'get',
+ '-H',
+ '-p',
+ '-d',
+ '1',
+ '-t',
+ 'filesystem,volume',
+ '-o',
+ 'name,property,value,source',
+ join(',', @nvmet_zfs_props),
+ $pool,
+ ];
+}
+
+# Every configfs object the plugin uses, in one traversal. Keys are never
+# read back; with host NQNs, only their sha256 digests are returned. Only the
+# kernel's own groups of objects the plugin does not use are skipped, so an
+# object of another tool with one of their names is still listed.
+my sub nvmet_configfs_step($hostnqns) {
+ my @names = map { ('-o', '-name', $_) } @nvmet_cfs_attrs;
+ shift @names;
+ my @keys = map { ('-o', '-path', nvmet_host_path($_) . '/dhchap_key') } $hostnqns->@*;
+ shift @keys;
+ return [
+ 'env',
+ 'LC_ALL=C',
+ 'find',
+ $nvmet_root,
+ '(',
+ '-path',
+ "$nvmet_root/subsystems/*/passthru",
+ '-o',
+ '-path',
+ "$nvmet_root/ports/*/ana_groups",
+ '-o',
+ '-path',
+ "$nvmet_root/ports/*/referrals",
+ ')',
+ '-type',
+ 'd',
+ '-prune',
+ '-o',
+ '-type',
+ 'd',
+ '-exec',
+ 'printf',
+ 'D %s\n',
+ '{}',
+ '+',
+ '-o',
+ '-type',
+ 'l',
+ '-exec',
+ 'printf',
+ 'L %s\n',
+ '{}',
+ '+',
+ '-o',
+ '-type',
+ 'f',
+ '(',
+ @names,
+ ')',
+ '-exec',
+ 'grep',
+ '',
+ '/dev/null',
+ '{}',
+ '+',
+ (@keys ? ('-o', '-type', 'f', '(', @keys, ')', '-exec', 'sha256sum', '{}', '+') : ()),
+ ];
+}
+
+# The whole target state: the ZFS inventory, a marker and configfs. Only a
+# locked read loads nvmet first (modprobe).
+my sub nvmet_state_steps($pool, $hostnqns, $modprobe) {
+ return [
+ ($modprobe ? (['modprobe', 'nvmet_tcp']) : ()),
+ nvmet_zfs_inventory_step($pool),
+ ['printf', '%s\n', $nvmet_marker],
+ nvmet_configfs_step($hostnqns),
+ ];
+}
+
+# ---------------------------------------------------------------------------
+# Renderer and runner
+# ---------------------------------------------------------------------------
+
+my sub nvmet_valid_word($word) {
+ return defined($word) && !ref($word) && $word !~ /[\0\r\n]/;
+}
+
+my sub nvmet_valid_path($path) {
+ return nvmet_valid_word($path) && $path =~ m{\A/};
+}
+
+my sub nvmet_quote($word) {
+ return PVE::Tools::shellquote($word);
+}
+
+# Renders steps into one POSIX shell command line: simple commands joined with
+# `&&`, plus `>` for configfs writes and one `dd | tee` pipe for a key read from
+# stdin. There are no loops, variables, conditionals or substitutions. Every
+# word is quoted; all operands are built from validated configuration.
+sub _nvmet_render($steps) {
+ my $invalid = "internal error: invalid NVMe target step\n";
+ die $invalid if ref($steps) ne 'ARRAY' || !$steps->@*;
+
+ my $keys = 0;
+ my @commands;
+ for my $step ($steps->@*) {
+ if (ref($step) eq 'ARRAY') {
+ die $invalid if !$step->@* || grep { !nvmet_valid_word($_) } $step->@*;
+ die $invalid if !$nvmet_commands{ $step->[0] };
+ push @commands, join(' ', map { nvmet_quote($_) } $step->@*);
+ } elsif (ref($step) eq 'HASH' && exists($step->{write})) {
+ die $invalid
+ if keys($step->%*) != 2
+ || !nvmet_valid_path($step->{write})
+ || !nvmet_valid_word($step->{value});
+ push @commands,
+ q{printf '%s\n' }
+ . nvmet_quote($step->{value}) . ' > '
+ . nvmet_quote($step->{write});
+ } elsif (ref($step) eq 'HASH' && exists($step->{key})) {
+ my $paths = $step->{key};
+ die $invalid
+ if keys($step->%*) != 1
+ || ref($paths) ne 'ARRAY'
+ || !$paths->@*
+ || grep { !nvmet_valid_path($_) } $paths->@*;
+ $keys++;
+ die $invalid if $keys > 1;
+ # configfs needs the whole key in one write(2). dd collects its
+ # input into one output block: GNU dd without bs=, busybox dd only
+ # when ibs and obs differ. tee writes it to every host object.
+ push @commands,
+ 'dd ibs=4096 obs=8192 2>/dev/null | tee '
+ . join(' ', map { nvmet_quote($_) } $paths->@*)
+ . ' >/dev/null';
+ } else {
+ die $invalid;
+ }
+ }
+ return join(' && ', @commands);
+}
+
+# Packs whole units into calls of at most $max bytes of rendered command, the
+# single argument sshd runs, below its limit. A rendered command is ASCII
+# (every operand is validated), so its length is its size in bytes.
+sub _nvmet_chunk($units, $max = $nvmet_max_command) {
+ my (@chunks, @current);
+ for my $unit ($units->@*) {
+ my @next = (@current, $unit->@*);
+ if (length(_nvmet_render(\@next)) > $max) {
+ die "internal error: NVMe target command too long\n"
+ if !@current || length(_nvmet_render($unit)) > $max;
+ push @chunks, [@current];
+ @next = $unit->@*;
+ }
+ @current = @next;
+ }
+ push @chunks, [@current] if @current;
+ return \@chunks;
+}
+
+# Never propagate the context run_command adds to the task marker: it quotes
+# the command line.
+my sub rethrow_task_interrupt($error) {
+ die "received interrupt\n" if $error =~ $RE_TASK_INTERRUPT;
+}
+
+# The only code that runs anything on the target. Returns
+# { rc => exit code, or -1 when ssh did not run to its end, out => [lines],
+# err => text } and never dies on a failing command; a stopped task dies with
+# the task marker. %opts: op (label, required), timeout and input (stdin,
+# used for keys).
+sub _nvmet_run($scfg, $steps, %opts) {
+ die "internal error: NVMe target call without label\n" if !defined($opts{op});
+ my $command = _nvmet_render($steps);
+ die "internal error: NVMe target command too long\n" if length($command) > $nvmet_max_command;
+ my $cmd = [@ssh_cmd, '-i', nvmet_ssh_key($scfg), 'root@' . nvmet_server($scfg), $command];
+ my (@out, $err);
+ my $rc = 0;
+ # Not noerr: run_command would then also swallow the interrupt of a stopped
+ # task, which must end the call. The exit code is taken from the exception.
+ eval {
+ run_command(
+ $cmd,
+ timeout => $opts{timeout} // 15,
+ outfunc => sub($line) { push @out, $line },
+ errfunc => sub($line) { $err = $line if !defined($err) && $line ne '' },
+ (defined($opts{input}) ? (input => $opts{input}) : ()),
+ );
+ };
+ if (my $error = $@) {
+ rethrow_task_interrupt($error);
+ return { rc => -1, out => [], err => 'timeout' } if $error =~ $RE_COMMAND_TIMEOUT;
+ return { rc => -1, out => [], err => 'ssh failed' } if $error !~ $RE_COMMAND_EXIT;
+ $rc = $+{code};
+ }
+ return { rc => $rc, out => \@out, err => $err // ($rc ? "exit code $rc" : '') };
+}
+
+# Monotonic clock for the activation backoff and local waits (a test seam).
+sub _now() {
+ return Time::HiRes::clock_gettime(Time::HiRes::CLOCK_MONOTONIC());
+}
+
+# Every local wait of the plugin (a test seam).
+sub _sleep($seconds) {
+ Time::HiRes::sleep($seconds);
+ return;
+}
+
+# ---------------------------------------------------------------------------
+# Parsers (pure)
+# ---------------------------------------------------------------------------
+
+sub _nvmet_split_state($text) {
+ my @parts = split /^\Q$nvmet_marker\E\n/m, $text, -1;
+ die "malformed NVMe target state\n" if @parts != 2;
+ return @parts;
+}
+
+sub _nvmet_property_value($value, $source) {
+ die "malformed ZFS property value\n"
+ if !defined($value)
+ || !defined($source)
+ || $value eq ''
+ || $source eq ''
+ || $value =~ /[\r\n\t]/
+ || $source =~ /[\r\n\t]/;
+ return $value if $source eq 'local' || $source eq 'received';
+ return '-' if $value eq '-' && $source eq '-';
+ if ($source =~ $RE_NVMET_INHERITED) {
+ die "invalid ZFS pool name\n" if $+{source} !~ $RE_NVMET_POOL;
+ return '-';
+ }
+ die "invalid ZFS property source\n";
+}
+
+# Parses the ZFS inventory into { types => { dataset => type }, volumes =>
+# { dataset => { nqn, nsid, uuid } }, last_nsid => highest NSID handed out }.
+# The inventory must be complete: the pool rows and all rows of every
+# dataset, or the read fails. Children whose names PVE never uses (for
+# example with a space) are kept as datasets but never count as owned.
+sub _nvmet_parse_zfs_inventory($text, $pool) {
+ die "invalid ZFS pool name\n" if $pool !~ $RE_NVMET_POOL;
+
+ my (%rows, %foreign);
+ for my $line (split /\n/, $text) {
+ my @fields = split /\t/, $line, -1;
+ die "malformed ZFS inventory row\n" if @fields != 4 || $line =~ /\r/;
+ my ($dataset, $property, $value, $source) = @fields;
+ if ($dataset ne $pool) {
+ die "ZFS dataset is outside configured pool\n" if index($dataset, "$pool/") != 0;
+ my $child = substr($dataset, length($pool) + 1);
+ die "ZFS dataset is outside configured pool\n"
+ if $child eq '' || index($child, '/') >= 0;
+ $foreign{$dataset} = 1 if $child !~ $RE_NVMET_DATASET_NAME;
+ }
+ my $on = $foreign{$dataset} ? '' : " on '$dataset'";
+ my $key = $nvmet_zfs_keys{$property} // die "unexpected ZFS property$on\n";
+ die "duplicate ZFS property '$property'$on\n" if exists($rows{$dataset}->{$key});
+ if ($key eq 'type') {
+ die "invalid ZFS dataset type$on\n"
+ if ($value ne 'filesystem' && $value ne 'volume') || $source ne '-';
+ $rows{$dataset}->{type} = $value;
+ } else {
+ $rows{$dataset}->{$key} = _nvmet_property_value($value, $source);
+ }
+ }
+ die "ZFS dataset inventory is missing the configured pool\n" if !$rows{$pool};
+
+ my (%types, %volumes);
+ for my $dataset (keys %rows) {
+ my $row = $rows{$dataset};
+ die "incomplete ZFS properties" . ($foreign{$dataset} ? '' : " on '$dataset'") . "\n"
+ if keys($row->%*) != scalar(@nvmet_zfs_props);
+ $types{$dataset} = $row->{type};
+ next if $row->{type} ne 'volume';
+ # nvmet shows device_uuid and udev names the by-id link in lower case,
+ # so identities are compared in lower case.
+ $volumes{$dataset} =
+ $foreign{$dataset}
+ ? { nqn => '-', nsid => '-', uuid => '-' }
+ : { nqn => $row->{nqn}, nsid => $row->{nsid}, uuid => lc($row->{uuid}) };
+ }
+ my $last_nsid = $rows{$pool}->{last_nsid};
+ return {
+ types => \%types,
+ volumes => \%volumes,
+ last_nsid => nvmet_valid_nsid($last_nsid) ? $last_nsid : 0,
+ };
+}
+
+# Parses the configfs read into
+# { subsystems => { nqn => { attr_model, attr_serial, attr_allow_any_host,
+# acl => { hostnqn => 1 }, namespaces => { nsid => { enable, device_path,
+# device_uuid, buffered_io } } } },
+# ports => { id => { addr_trtype, addr_adrfam, addr_traddr, addr_trsvcid,
+# links => { nqn => 1 } } },
+# hosts => { hostnqn => { key_sha256 => hex or undef } } }
+# Every object must be complete; unknown objects of newer kernels are ignored.
+sub _nvmet_parse_configfs($text) {
+ my $state = { subsystems => {}, ports => {}, hosts => {} };
+ my (%top, %dir);
+
+ my $subsystem = sub($nqn) {
+ return $state->{subsystems}->{$nqn} //= { acl => {}, namespaces => {} };
+ };
+ my $namespace = sub($nqn, $nsid) {
+ return $subsystem->($nqn)->{namespaces}->{$nsid} //= {};
+ };
+ my $port = sub($id) { return $state->{ports}->{$id} //= { links => {} } };
+ my $host = sub($hostnqn) { return $state->{hosts}->{$hostnqn} //= { key_sha256 => undef } };
+ my $set = sub($object, $attr, $value) {
+ die "duplicate NVMe target configfs attribute\n" if exists($object->{$attr});
+ $object->{$attr} = $value;
+ };
+
+ for my $line (split /\n/, $text) {
+ if ($line =~ $RE_CFS_TOP) {
+ $top{ $+{top} // 'root' } = 1;
+ } elsif ($line =~ $RE_CFS_SUBSYS_DIR) {
+ my $nqn = $+{nqn};
+ $subsystem->($nqn);
+ $dir{"subsystem $nqn"} = 1;
+ } elsif ($line =~ $RE_CFS_SUBSYS_GROUP || $line =~ $RE_CFS_PORT_GROUP) {
+ # default groups, always present with their parent
+ } elsif ($line =~ $RE_CFS_NS_DIR) {
+ my ($nqn, $nsid) = @+{qw(nqn nsid)};
+ $namespace->($nqn, $nsid);
+ $dir{"namespace $nqn/$nsid"} = 1;
+ } elsif ($line =~ $RE_CFS_PORT_DIR) {
+ my $id = $+{port};
+ $port->($id);
+ $dir{"port $id"} = 1;
+ } elsif ($line =~ $RE_CFS_HOST_DIR) {
+ my $hostnqn = $+{host};
+ $host->($hostnqn);
+ $dir{"host $hostnqn"} = 1;
+ } elsif ($line =~ $RE_CFS_PORT_LINK) {
+ my ($id, $nqn) = @+{qw(port nqn)};
+ $port->($id)->{links}->{$nqn} = 1;
+ } elsif ($line =~ $RE_CFS_ACL_LINK) {
+ my ($nqn, $hostnqn) = @+{qw(nqn host)};
+ $subsystem->($nqn)->{acl}->{$hostnqn} = 1;
+ } elsif ($line =~ $RE_CFS_OTHER_ENTRY) {
+ # objects of a shape the plugin does not use
+ } elsif ($line =~ $RE_CFS_NS_ATTR) {
+ my ($nqn, $nsid, $attr, $value) = @+{qw(nqn nsid attr value)};
+ $set->($namespace->($nqn, $nsid), $attr, $value);
+ } elsif ($line =~ $RE_CFS_SUBSYS_ATTR) {
+ my ($nqn, $attr, $value) = @+{qw(nqn attr value)};
+ $set->($subsystem->($nqn), $attr, $value);
+ } elsif ($line =~ $RE_CFS_PORT_ATTR) {
+ my ($id, $attr, $value) = @+{qw(port attr value)};
+ $set->($port->($id), $attr, $value);
+ } elsif ($line =~ $RE_CFS_KEY_DIGEST) {
+ my ($sha, $hostnqn) = @+{qw(sha host)};
+ my $object = $host->($hostnqn);
+ die "duplicate DH-HMAC-CHAP key digest" . nvmet_label($hostnqn) . "\n"
+ if defined($object->{key_sha256});
+ $object->{key_sha256} = $sha;
+ } elsif ($line =~ $RE_CFS_OTHER_ATTR) {
+ # attribute of an object shape the plugin does not use
+ } else {
+ die "unexpected NVMe target configfs line\n";
+ }
+ }
+
+ for my $name (qw(root hosts ports subsystems)) {
+ die "NVMe target configfs is unavailable\n" if !$top{$name};
+ }
+ for my $nqn (keys $state->{subsystems}->%*) {
+ my $subsys = $state->{subsystems}->{$nqn};
+ die "incomplete NVMe subsystem" . nvmet_label($nqn) . "\n"
+ if !$dir{"subsystem $nqn"}
+ || grep { !defined($subsys->{$_}) } qw(attr_model attr_serial attr_allow_any_host);
+ for my $nsid (keys $subsys->{namespaces}->%*) {
+ my $ns = $subsys->{namespaces}->{$nsid};
+ die "incomplete NVMe namespace '$nsid' of subsystem" . nvmet_label($nqn) . "\n"
+ if !$dir{"namespace $nqn/$nsid"}
+ || grep { !defined($ns->{$_}) } qw(enable device_path device_uuid buffered_io);
+ }
+ }
+ for my $id (keys $state->{ports}->%*) {
+ my $object = $state->{ports}->{$id};
+ my @attrs = qw(addr_trtype addr_adrfam addr_traddr addr_trsvcid);
+ die "incomplete NVMe port '$id'\n"
+ if !$dir{"port $id"} || grep { !defined($object->{$_}) } @attrs;
+ }
+ for my $hostnqn (keys $state->{hosts}->%*) {
+ die "DH-HMAC-CHAP key digest for unknown NVMe host" . nvmet_label($hostnqn) . "\n"
+ if !$dir{"host $hostnqn"};
+ }
+ return $state;
+}
+
+sub _nvmet_parse_mounts($text) {
+ for my $line (split /\n/, $text) {
+ my (undef, $mountpoint, $type) = split / /, $line;
+ return 1 if ($mountpoint // '') eq '/sys/kernel/config' && ($type // '') eq 'configfs';
+ }
+ return 0;
+}
+
+# ---------------------------------------------------------------------------
+# Planners (pure)
+# ---------------------------------------------------------------------------
+
+sub _nvmet_serial($nqn) {
+ return 'PVEZFS' . substr(sha256_hex($nqn), 0, 14);
+}
+
+sub _nvmet_template_name($dataset) {
+ die "only VM zvols can become templates\n"
+ if $dataset !~ $RE_NVMET_TEMPLATE_ZVOL || $+{type} ne 'vm';
+ return nvmet_template_twin($dataset);
+}
+
+# The durable identity (NSID, UUID) of a volume owned by $nqn. Refuses a
+# volume that another owned volume duplicates anywhere in the pool.
+sub _nvmet_owned_identity($inv, $nqn, $dataset) {
+ my $type = $inv->{types}->{$dataset} // die "ZFS volume '$dataset' does not exist\n";
+ die "'$dataset' is not a ZFS volume\n" if $type ne 'volume';
+ my $row = $inv->{volumes}->{$dataset};
+ die "ZFS volume '$dataset' is not owned by NVMe subsystem '$nqn'\n" if $row->{nqn} ne $nqn;
+ die "invalid NSID on '$dataset'\n" if !nvmet_valid_nsid($row->{nsid});
+ die "invalid namespace UUID on '$dataset'\n" if $row->{uuid} !~ $RE_NVMET_UUID;
+ for my $other (keys $inv->{volumes}->%*) {
+ next if $other eq $dataset;
+ my $identity = $inv->{volumes}->{$other};
+ next if $identity->{nqn} ne $nqn;
+ die "duplicate NVMe identity on '$dataset'\n"
+ if $identity->{nsid} eq $row->{nsid} || $identity->{uuid} eq $row->{uuid};
+ }
+ return ($row->{nsid}, $row->{uuid});
+}
+
+# { nsid => [uuid, device] } for every volume owned by $nqn, after validating
+# all of them, so a bad identity refuses the whole plan before any change.
+sub _nvmet_desired_namespaces($inv, $nqn) {
+ my (%desired, %uuids);
+ for my $dataset (sort keys $inv->{volumes}->%*) {
+ my $row = $inv->{volumes}->{$dataset};
+ next if $row->{nqn} ne $nqn;
+ die "invalid NSID on '$dataset'\n" if !nvmet_valid_nsid($row->{nsid});
+ die "invalid namespace UUID on '$dataset'\n" if $row->{uuid} !~ $RE_NVMET_UUID;
+ die "duplicate NSID '$row->{nsid}'\n" if $desired{ $row->{nsid} };
+ die "duplicate namespace UUID '$row->{uuid}'\n" if $uuids{ lc($row->{uuid}) }++;
+ $desired{ $row->{nsid} } = [$row->{uuid}, "/dev/zvol/$dataset"];
+ }
+ return \%desired;
+}
+
+# NSIDs grow monotonically, also past the last NSID handed out in the pool,
+# so a host that missed a namespace removal never sees a new volume at an old
+# NSID. Only after the last NSID is used does the allocation fall back to the
+# lowest free one.
+sub _nvmet_allocate_nsid($inv, $cfs, $nqn) {
+ my %used;
+ for my $row (values $inv->{volumes}->%*) {
+ $used{ $row->{nsid} } = 1 if $row->{nqn} eq $nqn && nvmet_valid_nsid($row->{nsid});
+ }
+ $used{$_} = 1 for keys nvmet_namespaces($cfs, $nqn)->%*;
+
+ my $highest = max(0, $inv->{last_nsid} // 0, keys %used);
+ return $highest + 1 if $highest < $nvmet_max_nsid;
+ my $nsid = 1;
+ $nsid++ while $used{$nsid};
+ die "no free namespace ID\n" if $nsid > $nvmet_max_nsid;
+ return $nsid;
+}
+
+# Export state of (NSID, UUID, device) in the subsystem: present, disabled,
+# absent, incomplete (a disabled namespace with another identity, for example
+# one being built), stale (enabled on the other template name of the zvol) or
+# nosubsys. Dies on a conflict.
+sub _nvmet_export_state($cfs, $nqn, $nsid, $uuid, $dev) {
+ return 'nosubsys' if !$cfs->{subsystems}->{$nqn};
+ my $namespaces = nvmet_namespaces($cfs, $nqn);
+ for my $id (keys $namespaces->%*) {
+ next if $id eq $nsid;
+ die "namespace UUID '$uuid' is already in use\n"
+ if $namespaces->{$id}->{device_uuid} eq $uuid;
+ }
+ my $ns = $namespaces->{$nsid} // return 'absent';
+ if ($ns->{device_uuid} eq $uuid && $ns->{device_path} eq $dev) {
+ return $ns->{enable} eq '1' ? 'present' : 'disabled';
+ }
+ # A template conversion or its undo that ssh gave up on can rename the
+ # zvol on the target after an activation exported it under its old name.
+ my $twin = nvmet_template_twin($dev);
+ return 'stale'
+ if $ns->{enable} eq '1'
+ && $ns->{device_uuid} eq $uuid
+ && $twin ne ''
+ && $ns->{device_path} eq $twin;
+ die "NVMe namespace ID '$nsid' has a different identity; refusing to replace it\n"
+ if $ns->{enable} ne '0';
+ return 'incomplete';
+}
+
+# Units that export (NSID, UUID, device) and the state they start from:
+# absent, present, disabled, reclaim or stale. A disabled namespace with
+# another identity, or a stale one, is rebuilt, which is only correct under
+# the target lock. A stale namespace belongs to a template, or to a volume
+# whose template conversion failed, so no guest uses it.
+sub _nvmet_plan_export($cfs, $nqn, $nsid, $uuid, $dev) {
+ die "NVMe subsystem does not exist\n" if !$cfs->{subsystems}->{$nqn};
+ my $state = _nvmet_export_state($cfs, $nqn, $nsid, $uuid, $dev);
+ return ([nvmet_build_unit($nqn, $nsid, $uuid, $dev)], 'absent') if $state eq 'absent';
+ return ([], 'present') if $state eq 'present';
+ return ([nvmet_enable_unit($nqn, $nsid, $dev)], 'disabled') if $state eq 'disabled';
+ my $unexport = nvmet_unexport_unit($nqn, $nsid, $state eq 'stale');
+ my $build = nvmet_build_unit($nqn, $nsid, $uuid, $dev);
+ return ([[$unexport->@*, $build->@*]], $state eq 'stale' ? 'stale' : 'reclaim');
+}
+
+# Units that remove the namespace exporting $uuid, and its previous state.
+sub _nvmet_plan_unexport($cfs, $nqn, $uuid) {
+ my $namespaces = nvmet_namespaces($cfs, $nqn);
+ my @matches = grep { $namespaces->{$_}->{device_uuid} eq $uuid }
+ sort { $a <=> $b } keys $namespaces->%*;
+ die "duplicate namespace UUID '$uuid'\n" if @matches > 1;
+ return ([], undef) if !@matches;
+ my $enabled = $namespaces->{ $matches[0] }->{enable} eq '1';
+ return (
+ [nvmet_unexport_unit($nqn, $matches[0], $enabled)],
+ { nsid => $matches[0], enabled => $enabled },
+ );
+}
+
+# The port serving a portal ({ family, address, port } of parse_nvme_portals).
+# Among complete, matching ports, the one already linked to $nqn wins, then
+# the lowest id. Incomplete ports never match.
+sub _nvmet_find_port($cfs, $nqn, $portal) {
+ my $wanted = nvmet_canonical_address($portal->{address});
+ my @matches = grep {
+ my $port = $cfs->{ports}->{$_};
+ $port->{addr_trtype} eq 'tcp'
+ && $port->{addr_adrfam} eq $portal->{family}
+ && $port->{addr_traddr} ne ''
+ && $port->{addr_trsvcid} eq "$portal->{port}"
+ && nvmet_canonical_address($port->{addr_traddr}) eq $wanted
+ } sort {
+ $a <=> $b
+ } keys $cfs->{ports}->%*;
+ my @linked = grep { nvmet_port_linked($cfs, $_, $nqn) } @matches;
+ return $linked[0] // $matches[0];
+}
+
+# Units that bring the target to the configured state, except publishing:
+# objects (subsystem, ports, hosts), keys, ACLs, namespaces in NSID order,
+# then removal of undesired namespaces.
+# $conf: { nqn, portals (of parse_nvme_portals), hostnqns, keysha }
+sub _nvmet_plan_activation($inv, $cfs, $conf) {
+ my ($nqn, $keysha) = $conf->@{qw(nqn keysha)};
+ my $desired = _nvmet_desired_namespaces($inv, $nqn); # refuse before any unit
+ my $subsys_path = nvmet_subsys_path($nqn);
+ my $subsys = $cfs->{subsystems}->{$nqn};
+ my (@objects, @key_hosts, @acls, @namespaces, %port_ids);
+
+ if (!$subsys) {
+ push @objects,
+ [
+ ['mkdir', $subsys_path],
+ nvmet_write("$subsys_path/attr_model", $nvmet_model),
+ nvmet_write("$subsys_path/attr_serial", _nvmet_serial($nqn)),
+ nvmet_write("$subsys_path/attr_allow_any_host", 0),
+ ];
+ } else {
+ my $serial = _nvmet_serial($nqn);
+ my @attrs;
+ if ($subsys->{attr_model} ne $nvmet_model || $subsys->{attr_serial} ne $serial) {
+ die "refusing to take over existing NVMe subsystem '$nqn': it was not created by"
+ . " Proxmox VE; remove it on the target or use another NQN\n"
+ if !nvmet_unfinished_subsystem($cfs, $nqn);
+ push @attrs, nvmet_write("$subsys_path/attr_serial", $serial);
+ }
+ push @attrs, nvmet_write("$subsys_path/attr_allow_any_host", 0)
+ if $subsys->{attr_allow_any_host} ne '0';
+ push @objects, [@attrs] if @attrs;
+ }
+
+ my %taken = map { $_ => 1 } keys $cfs->{ports}->%*;
+ for my $portal ($conf->{portals}->@*) {
+ my $id = _nvmet_find_port($cfs, $nqn, $portal);
+ if (!defined($id)) {
+ $id = 1;
+ $id++ while $taken{$id};
+ $taken{$id} = 1;
+ my $port_path = nvmet_port_path($id);
+ push @objects,
+ [
+ ['mkdir', $port_path],
+ nvmet_write("$port_path/addr_trtype", 'tcp'),
+ nvmet_write("$port_path/addr_adrfam", $portal->{family}),
+ nvmet_write("$port_path/addr_traddr", $portal->{address}),
+ nvmet_write("$port_path/addr_trsvcid", $portal->{port}),
+ ];
+ }
+ $port_ids{$id} = 1;
+ }
+
+ # nvmet keeps the key on the global host object. A key that differs from
+ # ours is only replaced while no subsystem links the host.
+ my $linked = nvmet_linked_hosts($cfs);
+ for my $hostnqn ($conf->{hostnqns}->@*) {
+ my $host_path = nvmet_host_path($hostnqn);
+ if (my $host = $cfs->{hosts}->{$hostnqn}) {
+ my $sha = $host->{key_sha256}
+ // die "NVMe target does not expose a DH-HMAC-CHAP key for host '$hostnqn'\n";
+ if ($sha ne $keysha) {
+ die "refusing to replace an in-use DH-HMAC-CHAP key for '$hostnqn'\n"
+ if $linked->{$hostnqn};
+ push @key_hosts, $host_path;
+ }
+ } else {
+ push @objects, [['mkdir', $host_path]];
+ push @key_hosts, $host_path;
+ }
+ push @acls, [['ln', '-s', $host_path, "$subsys_path/allowed_hosts/$hostnqn"]]
+ if !$subsys || !($subsys->{acl} // {})->{$hostnqn};
+ }
+ # nvmet creates the key attributes world-readable. The chmod before the
+ # write keeps every later open from reading the key; a descriptor opened
+ # while an attribute was still world-readable is not affected by it.
+ my @keys;
+ if (@key_hosts) {
+ my @attrs = map { ("$_/dhchap_key", "$_/dhchap_ctrl_key") } @key_hosts;
+ @keys = ([['chmod', '0600', @attrs], { key => [map { "$_/dhchap_key" } @key_hosts] }]);
+ }
+
+ my $current = { subsystems => { $nqn => $subsys // { acl => {}, namespaces => {} } } };
+ for my $nsid (sort { $a <=> $b } keys $desired->%*) {
+ my ($units) = _nvmet_plan_export($current, $nqn, $nsid, $desired->{$nsid}->@*);
+ push @namespaces, $units->@*;
+ }
+ my $existing = nvmet_namespaces($cfs, $nqn);
+ for my $nsid (sort { $a <=> $b } keys $existing->%*) {
+ next if $desired->{$nsid};
+ push @namespaces, nvmet_unexport_unit($nqn, $nsid, $existing->{$nsid}->{enable} eq '1');
+ }
+
+ return {
+ prepublish => [@objects, @keys, @acls, @namespaces],
+ port_ids => [sort { $a <=> $b } keys %port_ids],
+ };
+}
+
+# One call per port: each link repeats the guards, so it only succeeds while
+# the subsystem denies unknown hosts and every configured host has its ACL.
+# Under the target lock, only a manager of the target outside this cluster
+# can change that after the verify read; the guards stop the publish then.
+# Links to ports that are no longer configured are removed.
+# Returns [{ port => $id, action => 'link' | 'unlink', steps => [...] }, ...].
+sub _nvmet_plan_publish($cfs, $nqn, $port_ids, $hostnqns) {
+ my $subsys_path = nvmet_subsys_path($nqn);
+ my %wanted = map { $_ => 1 } $port_ids->@*;
+ my @guards = (
+ ['grep', '-qx', '0', "$subsys_path/attr_allow_any_host"],
+ map { ['test', '-L', "$subsys_path/allowed_hosts/$_"] } $hostnqns->@*,
+ );
+
+ my @units;
+ for my $id ($port_ids->@*) {
+ next if nvmet_port_linked($cfs, $id, $nqn);
+ my $link = ['ln', '-s', $subsys_path, nvmet_port_path($id) . "/subsystems/$nqn"];
+ push @units, { port => $id, action => 'link', steps => [@guards, $link] };
+ }
+ for my $id (sort { $a <=> $b } keys $cfs->{ports}->%*) {
+ next if $wanted{$id} || !nvmet_port_linked($cfs, $id, $nqn);
+ my $unlink = ['rm', nvmet_port_path($id) . "/subsystems/$nqn"];
+ push @units, { port => $id, action => 'unlink', steps => [$unlink] };
+ }
+ return \@units;
+}
+
+# Teardown units for the subsystem and the hosts that may become orphans.
+sub _nvmet_plan_delete_target($inv, $cfs, $nqn, $hostnqns) {
+ for my $dataset (sort keys $inv->{volumes}->%*) {
+ die "refusing to delete NVMe subsystem '$nqn': owned ZFS volume '$dataset' exists\n"
+ if $inv->{volumes}->{$dataset}->{nqn} eq $nqn;
+ }
+ my $subsys = $cfs->{subsystems}->{$nqn} // return ([], [$hostnqns->@*]);
+ die "refusing to delete foreign NVMe subsystem '$nqn'\n"
+ if $subsys->{attr_model} ne $nvmet_model || $subsys->{attr_serial} ne _nvmet_serial($nqn);
+ die "refusing to delete NVMe subsystem '$nqn': namespaces remain\n"
+ if nvmet_namespaces($cfs, $nqn)->%*;
+
+ my $subsys_path = nvmet_subsys_path($nqn);
+ my @acl = sort keys(($subsys->{acl} // {})->%*);
+ die "refusing to delete NVMe subsystem '$nqn': it allows a malformed host name\n"
+ if grep { $_ !~ $RE_NQN } @acl;
+ my @links = (
+ (
+ map { nvmet_port_path($_) . "/subsystems/$nqn" }
+ grep { nvmet_port_linked($cfs, $_, $nqn) }
+ sort { $a <=> $b } keys $cfs->{ports}->%*
+ ),
+ (map { "$subsys_path/allowed_hosts/$_" } @acl),
+ );
+ # A retry after a teardown that failed at its rmdir finds no ACL left, so
+ # the configured hosts are candidates as well.
+ my %candidates = map { $_ => 1 } @acl, $hostnqns->@*;
+ return (
+ [[(@links ? (['rm', @links]) : ()), ['rmdir', $subsys_path]]], [sort keys %candidates],
+ );
+}
+
+# Candidate hosts that exist and that no subsystem links any more.
+sub _nvmet_plan_orphan_hosts($cfs, $candidates) {
+ my $linked = nvmet_linked_hosts($cfs);
+ my %seen;
+ my @orphans = map { nvmet_host_path($_) }
+ grep { !$seen{$_}++ && $_ =~ $RE_NQN && $cfs->{hosts}->{$_} && !$linked->{$_} }
+ sort $candidates->@*;
+ return @orphans ? [[['rmdir', @orphans]]] : [];
+}
+
+# ---------------------------------------------------------------------------
+# Target lock and remote execution
+# ---------------------------------------------------------------------------
+
+my %nvmet_lock_owner; # lock id => pid of the process holding the domain lock
+
+# Runs $code under the pmxcfs domain lock of the target. Like the pmxcfs
+# lock itself, it is not re-entrant, and a child forked inside is not the
+# owner.
+my sub nvmet_locked($scfg, $code) {
+ my $id = nvmet_lock_id($scfg);
+ die "cluster not quorate - refusing NVMe target changes\n"
+ if !PVE::Cluster::check_cfs_quorum(1);
+ my $res = PVE::Cluster::cfs_lock_domain(
+ $id,
+ $nvmet_lock_wait,
+ sub {
+ local $nvmet_lock_owner{$id} = $$;
+ return $code->();
+ },
+ );
+ die $@ if $@;
+ return $res;
+}
+
+# Runs steps on the target. %opts: op, timeout, input, change (the steps
+# change the target, which needs the lock).
+#
+# ssh does not stop a command on the target when it gives up on it: a change
+# whose connection broke (exit code 255), any call under the lock that ssh
+# did not run to its end (-1: its timeout, or ssh was killed), and a call
+# during which the task was stopped (it dies with the task marker) may still
+# be running there. Its outcome is unknown, so it is never compensated from a
+# read that could come before the rest of it. The next activation repairs
+# from the target state.
+my sub nvmet_exec($scfg, $steps, %opts) {
+ my $locked = ($nvmet_lock_owner{ nvmet_lock_id($scfg) } // 0) == $$;
+ die "internal error: NVMe target change without target lock\n"
+ if $opts{change} && !$locked;
+ my $res = _nvmet_run(
+ $scfg, $steps,
+ op => $opts{op},
+ timeout => $opts{timeout} // 15,
+ (defined($opts{input}) ? (input => $opts{input}) : ()),
+ );
+ die "NVMe target operation '$opts{op}' did not complete ($res->{err});"
+ . " the target state is unknown\n"
+ if ($locked && $res->{rc} == -1)
+ || ($opts{change} && $res->{rc} == 255);
+ return $res;
+}
+
+# An abandoned operation is passed on instead of being compensated.
+my sub nvmet_rethrow_abandoned($error) {
+ die $error if $error =~ $RE_NVMET_ABANDONED;
+}
+
+my sub nvmet_unreachable($res) {
+ return $res->{rc} == 255 || $res->{rc} == -1;
+}
+
+my sub nvmet_check($res, $op) {
+ die "NVMe target operation '$op' failed: $res->{err}\n" if $res->{rc};
+}
+
+my sub nvmet_read_failed($scfg, $res) {
+ die "NVMe target '" . nvmet_server($scfg) . "' is unreachable: $res->{err}\n"
+ if nvmet_unreachable($res);
+ die "cannot read NVMe target state: $res->{err}\n";
+}
+
+my sub nvmet_output($res) {
+ return join('', map { "$_\n" } $res->{out}->@*);
+}
+
+# A read, retried twice after 200 ms unless the target is unreachable: a read
+# without the lock can meet a configfs object that another node removes
+# while find lists it, which makes find fail once.
+my sub nvmet_read_call($scfg, $steps, %opts) {
+ my $res;
+ for my $attempt (0 .. 2) {
+ _sleep(0.2) if $attempt;
+ $res = nvmet_exec($scfg, $steps, %opts, timeout => 15);
+ last if !$res->{rc} || nvmet_unreachable($res);
+ }
+ return $res;
+}
+
+# Reads the target state and returns (ZFS inventory, configfs state).
+# keys => [hostnqns] adds the digests of their keys. modprobe => 1, for locked
+# reads only, loads nvmet first and mounts configfs if that is what failed.
+# With tolerant => 1, a failed or unparsable read returns an empty list,
+# unless the target is unreachable.
+my sub nvmet_read($scfg, %opts) {
+ my $pool = nvmet_pool($scfg);
+ my $hostnqns = $opts{keys} // [];
+ # loading a module that may still be loading is harmless
+ my @modprobe = $opts{modprobe} ? (change => 1) : ();
+ my $read = sub () {
+ return nvmet_read_call(
+ $scfg,
+ nvmet_state_steps($pool, $hostnqns, $opts{modprobe}),
+ op => 'read target state',
+ @modprobe,
+ );
+ };
+ my $res = $read->();
+ if ($res->{rc} && $opts{modprobe} && !nvmet_unreachable($res)) {
+ my $mounts = nvmet_exec($scfg, [['cat', '/proc/mounts']], op => 'read mounts');
+ nvmet_check($mounts, 'read mounts');
+ if (!_nvmet_parse_mounts(nvmet_output($mounts))) {
+ my $mount = ['mount', '-t', 'configfs', 'none', '/sys/kernel/config'];
+ my $mounted = nvmet_exec($scfg, [$mount], op => 'mount configfs', change => 1);
+ nvmet_check($mounted, 'mount configfs');
+ $res = $read->();
+ }
+ }
+ if ($res->{rc}) {
+ return () if $opts{tolerant} && !nvmet_unreachable($res);
+ nvmet_read_failed($scfg, $res);
+ }
+ my @state = eval {
+ my ($zfs, $configfs) = _nvmet_split_state(nvmet_output($res));
+ (_nvmet_parse_zfs_inventory($zfs, $pool), _nvmet_parse_configfs($configfs));
+ };
+ if (my $error = $@) {
+ return () if $opts{tolerant};
+ die $error;
+ }
+ return @state;
+}
+
+my sub nvmet_read_zfs($scfg) {
+ my $pool = nvmet_pool($scfg);
+ my $res = nvmet_read_call($scfg, [nvmet_zfs_inventory_step($pool)], op => 'read ZFS inventory');
+ nvmet_read_failed($scfg, $res) if $res->{rc};
+ return _nvmet_parse_zfs_inventory(nvmet_output($res), $pool);
+}
+
+# Applies units in as few calls as possible and stops at the first failing
+# call, whose result it returns. A chain only saves SSH round trips: every
+# decision was taken locally before, and `&&` only stops at the first
+# failing command. Only the call with the key step gets stdin. The default
+# timeout suits configfs chains; ZFS changes pass their own.
+my sub nvmet_apply($scfg, $units, %opts) {
+ my $input = delete $opts{input};
+ $opts{timeout} //= 30;
+ for my $chunk (_nvmet_chunk($units)->@*) {
+ my $with_key = grep { ref($_) eq 'HASH' && exists($_->{key}) } $chunk->@*;
+ my @input = $with_key ? (input => $input) : ();
+ my $res = nvmet_exec($scfg, $chunk, %opts, @input, change => 1);
+ return $res if $res->{rc};
+ }
+ return { rc => 0, out => [], err => '' };
+}
+
+my sub nvmet_change($scfg, $units, %opts) {
+ nvmet_check(nvmet_apply($scfg, $units, %opts), $opts{op});
+}
+
+my sub nvmet_wait_device($scfg, $dev) {
+ my $deadline = _now() + 10;
+ while (1) {
+ my $res = nvmet_exec($scfg, [['test', '-b', $dev]], op => 'wait for zvol');
+ return if !$res->{rc};
+ nvmet_read_failed($scfg, $res) if nvmet_unreachable($res);
+ last if _now() >= $deadline;
+ _sleep(0.25);
+ }
+ die "zvol '$dev' is not a block device after 10 seconds\n";
+}
+
+my sub nvmet_new_uuid($inv, $cfs) {
+ my %used = map { lc($_->{uuid}) => 1 } values $inv->{volumes}->%*;
+ for my $subsys (values $cfs->{subsystems}->%*) {
+ $used{ lc($_->{device_uuid}) } = 1 for values(($subsys->{namespaces} // {})->%*);
+ }
+ for (1 .. 10) {
+ my $uuid = lc(file_read_firstline('/proc/sys/kernel/random/uuid') // '');
+ return $uuid if $uuid =~ $RE_NVMET_UUID && !$used{$uuid};
+ }
+ die "cannot generate an unused namespace UUID\n";
+}
+
+my sub nvmet_has_identity($row, $nqn, $nsid, $uuid) {
+ return $row && $row->{nqn} eq $nqn && $row->{nsid} eq "$nsid" && $row->{uuid} eq $uuid;
+}
+
+# The configfs state after a successful unexport of namespace $pre->{nsid}.
+my sub nvmet_without_namespace($cfs, $nqn, $pre) {
+ return $cfs if !$pre;
+ my $subsys = $cfs->{subsystems}->{$nqn};
+ my %namespaces = $subsys->{namespaces}->%*;
+ delete $namespaces{ $pre->{nsid} };
+ my $changed = { $subsys->%*, namespaces => \%namespaces };
+ return { $cfs->%*, subsystems => { $cfs->{subsystems}->%*, $nqn => $changed } };
+}
+
+# Links the subsystem to each desired port separately. One port that cannot
+# be bound (for example, its address is missing on the target) only warns
+# while another desired port serves the subsystem.
+my sub nvmet_publish($scfg, $cfs, $nqn, $port_ids, $hostnqns) {
+ my @failed;
+ for my $unit (_nvmet_plan_publish($cfs, $nqn, $port_ids, $hostnqns)->@*) {
+ my $op = "$unit->{action} NVMe subsystem on port $unit->{port}";
+ my $res = nvmet_apply($scfg, [$unit->{steps}], op => $op);
+ push @failed, { $unit->%*, err => $res->{err} } if $res->{rc};
+ }
+ return if !@failed;
+
+ my (undef, $fresh) = eval { nvmet_read($scfg) };
+ nvmet_rethrow_abandoned($@) if $@;
+ my $published = $fresh && grep { nvmet_port_linked($fresh, $_, $nqn) } $port_ids->@*;
+ my ($first) = grep { $_->{action} eq 'link' } @failed;
+ $first //= $failed[0];
+ die "cannot publish NVMe subsystem '$nqn': $first->{err}\n" if !$published;
+ for my $failure (@failed) {
+ my $what = $failure->{action} eq 'link' ? 'publish' : 'unpublish';
+ log_warn("cannot $what NVMe subsystem '$nqn' on port $failure->{port}: $failure->{err}");
+ }
+}
+
+# ---------------------------------------------------------------------------
+# Target flows
+# ---------------------------------------------------------------------------
+
+# Brings the target to the configured state and publishes the subsystem.
+# A converged target costs one read and no lock. Otherwise, under the lock:
+# read, apply, verify with a fresh read, and publish only once the verify
+# read plans nothing more. $portals and $hostnqns are the parsed and validated
+# storage configuration.
+sub _nvmet_activate_target($storeid, $scfg, $portals, $hostnqns, $key) {
+ my $nqn = nvmet_nqn($scfg);
+ my $conf = {
+ nqn => $nqn,
+ portals => $portals,
+ hostnqns => $hostnqns,
+ keysha => sha256_hex("$key\n"),
+ };
+ my $converged = sub($inv, $cfs) {
+ my $plan = _nvmet_plan_activation($inv, $cfs, $conf);
+ return !$plan->{prepublish}->@*
+ && !_nvmet_plan_publish($cfs, $nqn, $plan->{port_ids}, $hostnqns)->@*;
+ };
+
+ # Without the lock, a refusal or failed read only means "take the lock".
+ my @state = nvmet_read($scfg, keys => $hostnqns, tolerant => 1);
+ return if @state && eval { $converged->(@state) };
+
+ nvmet_locked(
+ $scfg,
+ sub {
+ # Removing the storage deletes its key before it removes the target
+ # configuration, so an activation that waited for the lock meanwhile
+ # does not restore it.
+ die "storage '$storeid' is being removed or its key changed\n"
+ if (file_read_firstline(secret_path($storeid)) // '') ne $key;
+ my ($inv, $cfs) = nvmet_read($scfg, keys => $hostnqns, modprobe => 1);
+ my $plan = _nvmet_plan_activation($inv, $cfs, $conf);
+ my $error;
+ for my $round (1 .. 2) {
+ last if !$plan->{prepublish}->@*;
+ my $apply = nvmet_apply(
+ $scfg, $plan->{prepublish},
+ op => 'configure NVMe target',
+ input => "$key\n",
+ );
+ $error //= $apply->{err} if $apply->{rc};
+ ($inv, $cfs) = nvmet_read($scfg, keys => $hostnqns);
+ $plan = _nvmet_plan_activation($inv, $cfs, $conf);
+ }
+ die "NVMe target did not converge: " . ($error // 'state mismatch') . "\n"
+ if $plan->{prepublish}->@*;
+ nvmet_publish($scfg, $cfs, $nqn, $plan->{port_ids}, $hostnqns);
+ },
+ );
+ return;
+}
+
+# The namespace UUID of an owned volume, from one read-only ZFS query.
+sub _nvmet_volume_uuid($scfg, $name) {
+ my $dataset = nvmet_dataset($scfg, $name);
+ my $inv = nvmet_read_zfs($scfg);
+ my (undef, $uuid) = _nvmet_owned_identity($inv, nvmet_nqn($scfg), $dataset);
+ return $uuid;
+}
+
+# Creates (size => KiB) or clones (origin => pool/base@snap) a zvol with its
+# identity set at creation, then exports it. A failure removes only what this
+# call created, decided from a fresh read.
+my sub nvmet_create_volume($scfg, $name, %opts) {
+ my $nqn = nvmet_nqn($scfg);
+ my $dataset = nvmet_dataset($scfg, $name);
+ my $dev = "/dev/zvol/$dataset";
+ die "internal error: invalid zvol creation request\n"
+ if defined($opts{size}) == defined($opts{origin})
+ || (defined($opts{size}) && $opts{size} !~ $RE_UNSIGNED_INTEGER);
+
+ return nvmet_locked(
+ $scfg,
+ sub {
+ my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1);
+ die "zvol '$dataset' already exists\n" if exists($inv->{types}->{$dataset});
+ die "NVMe subsystem does not exist; activate the storage first\n"
+ if !$cfs->{subsystems}->{$nqn};
+ my $nsid = _nvmet_allocate_nsid($inv, $cfs, $nqn);
+ my $uuid = nvmet_new_uuid($inv, $cfs);
+ my (undef, $state) = _nvmet_plan_export($cfs, $nqn, $nsid, $uuid, $dev);
+ die "internal error: namespace ID '$nsid' is in use\n" if $state ne 'absent';
+
+ # Record the NSID on the pool before the zvol exists: a creation that
+ # completes on the target only after this call gave up on it must
+ # never share its NSID with a later volume.
+ my $timeout = nvmet_long_timeout();
+ if ($nsid > $inv->{last_nsid}) {
+ my $reserve = ['zfs', 'set', "$nvmet_last_nsid=$nsid", nvmet_pool($scfg)];
+ nvmet_change(
+ $scfg, [[$reserve]],
+ op => 'reserve namespace ID',
+ timeout => $timeout,
+ );
+ }
+
+ my @identity = map { ('-o', $_) } nvmet_identity($nqn, $nsid, $uuid);
+ my @create_opts = (
+ ($scfg->{sparse} ? ('-s') : ()),
+ ($scfg->{blocksize} ? ('-b', $scfg->{blocksize}) : ()),
+ );
+ my $create =
+ defined($opts{origin})
+ ? ['zfs', 'clone', @identity, $opts{origin}, $dataset]
+ : ['zfs', 'create', @create_opts, @identity, '-V', "$opts{size}k", $dataset];
+ my $res = nvmet_apply(
+ $scfg, [[$create]],
+ op => 'create zvol',
+ timeout => $timeout,
+ );
+ if ($res->{rc}) {
+ # A definite command failure may still have created the zvol.
+ # Only a zvol with this call's identity counts. Ambiguous SSH
+ # outcomes have already aborted the transaction in nvmet_exec.
+ my $check = eval { nvmet_read_zfs($scfg) };
+ nvmet_rethrow_abandoned($@) if $@;
+ my $row = $check ? $check->{volumes}->{$dataset} : undef;
+ die "NVMe target operation 'create zvol' failed: $res->{err}\n"
+ if !nvmet_has_identity($row, $nqn, $nsid, $uuid);
+ }
+
+ my $verified;
+ eval {
+ nvmet_wait_device($scfg, $dev);
+ my $build_unit = nvmet_build_unit($nqn, $nsid, $uuid, $dev);
+ my $build = nvmet_apply($scfg, [$build_unit], op => 'export namespace');
+ $verified = [nvmet_read($scfg)];
+ my ($vinv, $vcfs) = $verified->@*;
+ my ($vnsid, $vuuid) = _nvmet_owned_identity($vinv, $nqn, $dataset);
+ die "NVMe identity of zvol '$dataset' changed\n"
+ if $vnsid ne $nsid || $vuuid ne $uuid;
+ if (_nvmet_export_state($vcfs, $nqn, $nsid, $uuid, $dev) ne 'present') {
+ nvmet_check($build, 'export namespace');
+ die "NVMe namespace '$nsid' was not exported\n";
+ }
+ };
+ my $error = $@ or return $uuid;
+ nvmet_rethrow_abandoned($error);
+
+ eval {
+ my ($cinv, $ccfs) = $verified ? $verified->@* : nvmet_read($scfg);
+ my $namespaces = nvmet_namespaces($ccfs, $nqn);
+ # The build writes the UUID before it enables the namespace, so an
+ # enabled namespace without our UUID was not built by this call.
+ my @unexport =
+ map { nvmet_unexport_unit($nqn, $_, $namespaces->{$_}->{enable} eq '1') }
+ grep {
+ my $ns = $namespaces->{$_};
+ $ns->{device_uuid} eq $uuid
+ || ($ns->{enable} ne '1' && $ns->{device_path} eq $dev)
+ } sort {
+ $a <=> $b
+ } keys $namespaces->%*;
+ nvmet_change($scfg, \@unexport, op => 'remove namespace') if @unexport;
+ # -o at creation proves that this call created the zvol
+ if (nvmet_has_identity($cinv->{volumes}->{$dataset}, $nqn, $nsid, $uuid)) {
+ my $destroy = ['zfs', 'destroy', '-r', $dataset];
+ nvmet_change(
+ $scfg, [[$destroy]],
+ op => 'destroy zvol',
+ timeout => $timeout,
+ );
+ }
+ };
+ log_warn("failed to clean up zvol '$name': $@") if $@;
+ die $error;
+ },
+ );
+}
+
+# Unexports and destroys an owned zvol. A missing zvol is already destroyed.
+my sub nvmet_destroy_volume($scfg, $name) {
+ my $nqn = nvmet_nqn($scfg);
+ my $dataset = nvmet_dataset($scfg, $name);
+ my $dev = "/dev/zvol/$dataset";
+ my $destroy = ['zfs', 'destroy', '-r', $dataset];
+ my $timeout = nvmet_long_timeout();
+
+ nvmet_locked(
+ $scfg,
+ sub {
+ my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1);
+ return if !exists($inv->{types}->{$dataset});
+ my ($nsid, $uuid) = _nvmet_owned_identity($inv, $nqn, $dataset);
+ my ($unexport, $pre) = ([], undef);
+ if ($cfs->{subsystems}->{$nqn}) {
+ _nvmet_plan_export($cfs, $nqn, $nsid, $uuid, $dev); # refusals only
+ ($unexport, $pre) = _nvmet_plan_unexport($cfs, $nqn, $uuid);
+ }
+ my $units = [$unexport->@*, [$destroy]];
+ my $res = nvmet_apply($scfg, $units, op => 'destroy zvol', timeout => $timeout);
+ return if !$res->{rc};
+ my $error = $res->{err};
+
+ # Decide from a fresh read what the failed call changed.
+ my ($finv, $fcfs) = eval { nvmet_read($scfg) };
+ if (my $read_error = $@) {
+ nvmet_rethrow_abandoned($read_error);
+ die "failed to destroy '$dataset': $error\n";
+ }
+ return if !exists($finv->{types}->{$dataset});
+ my $namespaces = nvmet_namespaces($fcfs, $nqn);
+ my ($current) =
+ grep { $namespaces->{$_}->{device_uuid} eq $uuid } keys $namespaces->%*;
+ if (defined($current)) {
+ die "cannot disable namespace '$current': $error\n"
+ if $namespaces->{$current}->{enable} eq '1';
+ die "cannot remove namespace '$current': $error\n" if !$pre || !$pre->{enabled};
+ my $enable = [nvmet_enable_unit($nqn, $current, $dev)];
+ my $restore = nvmet_apply($scfg, $enable, op => 'enable namespace');
+ die "cannot remove namespace '$current': $error"
+ . ($restore->{rc} ? "; cannot restore namespace: $restore->{err}" : '') . "\n";
+ }
+
+ # The namespace is gone and the zvol is still there, typically busy.
+ for my $attempt (1 .. 5) {
+ _sleep(1);
+ my $retry =
+ nvmet_apply($scfg, [[$destroy]], op => 'destroy zvol', timeout => $timeout);
+ return if !$retry->{rc};
+ $error = $retry->{err};
+ }
+ die "failed to destroy '$dataset': $error\n" if !$pre;
+ eval {
+ my (undef, $rcfs) = nvmet_read($scfg);
+ my ($units) = _nvmet_plan_export($rcfs, $nqn, $nsid, $uuid, $dev);
+ nvmet_change($scfg, $units, op => 'restore namespace');
+ };
+ die "failed to destroy '$dataset': $error; namespace restoration also failed: $@"
+ if $@;
+ die "failed to destroy '$dataset': $error; namespace restored\n";
+ },
+ );
+ return;
+}
+
+# Rolls an owned zvol back and always exports it again with the same identity.
+my sub nvmet_rollback_volume($scfg, $name, $snap) {
+ my $snapshot = nvmet_snapshot($scfg, $name, $snap);
+ my $nqn = nvmet_nqn($scfg);
+ my $dataset = nvmet_dataset($scfg, $name);
+ my $dev = "/dev/zvol/$dataset";
+
+ my $timeout = nvmet_long_timeout();
+
+ nvmet_locked(
+ $scfg,
+ sub {
+ my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1);
+ my ($nsid, $uuid) = _nvmet_owned_identity($inv, $nqn, $dataset);
+ die "NVMe subsystem does not exist\n" if !$cfs->{subsystems}->{$nqn};
+ _nvmet_plan_export($cfs, $nqn, $nsid, $uuid, $dev); # refusals only
+ my ($unexport, $pre) = _nvmet_plan_unexport($cfs, $nqn, $uuid);
+ my $set_identity = ['zfs', 'set', nvmet_identity($nqn, $nsid, $uuid), $dataset];
+ my $rollback = ['zfs', 'rollback', $snapshot];
+ my $res = nvmet_apply(
+ $scfg, [$unexport->@*, [$rollback], [$set_identity]],
+ op => 'rollback zvol',
+ timeout => $timeout,
+ );
+ my $error;
+ $error = "NVMe target operation 'rollback zvol' failed: $res->{err}" if $res->{rc};
+
+ # Always export the volume again, from a fresh read if anything failed.
+ eval {
+ my ($rinv, $rcfs) =
+ $error ? nvmet_read($scfg) : ($inv, nvmet_without_namespace($cfs, $nqn, $pre));
+ nvmet_wait_device($scfg, $dev);
+ my @units;
+ push @units, [$set_identity]
+ if !nvmet_has_identity($rinv->{volumes}->{$dataset}, $nqn, $nsid, $uuid);
+ my ($export) = _nvmet_plan_export($rcfs, $nqn, $nsid, $uuid, $dev);
+ push @units, $export->@*;
+ nvmet_change($scfg, \@units, op => 'restore namespace', timeout => $timeout);
+ };
+ chomp(my $restore_error = $@);
+
+ die "rollback failed: $error; namespace restoration failed: $restore_error\n"
+ if defined($error) && $restore_error;
+ die "namespace restoration failed after rollback: $restore_error\n"
+ if $restore_error;
+ die "$error\n" if defined($error);
+ },
+ );
+ return;
+}
+
+# Renames vm-* to base-*, exports it under the same identity and takes the
+# __base__ snapshot. A failed conversion is undone from a fresh read.
+my sub nvmet_template_volume($scfg, $name) {
+ my $nqn = nvmet_nqn($scfg);
+ my $old = nvmet_dataset($scfg, $name);
+ my $new = _nvmet_template_name($old);
+ my ($old_dev, $new_dev) = ("/dev/zvol/$old", "/dev/zvol/$new");
+ my $timeout = nvmet_long_timeout();
+
+ nvmet_locked(
+ $scfg,
+ sub {
+ my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1);
+ my ($nsid, $uuid) = _nvmet_owned_identity($inv, $nqn, $old);
+ die "template zvol '$new' already exists\n" if exists($inv->{types}->{$new});
+ die "NVMe subsystem does not exist\n" if !$cfs->{subsystems}->{$nqn};
+ _nvmet_plan_export($cfs, $nqn, $nsid, $uuid, $old_dev); # refusals only
+ my ($unexport) = _nvmet_plan_unexport($cfs, $nqn, $uuid);
+
+ eval {
+ my $rename = [$unexport->@*, [['zfs', 'rename', $old, $new]]];
+ nvmet_change($scfg, $rename, op => 'rename zvol', timeout => $timeout);
+ nvmet_wait_device($scfg, $new_dev);
+ my $build = nvmet_build_unit($nqn, $nsid, $uuid, $new_dev);
+ my $template = [$build, [['zfs', 'snapshot', "$new\@__base__"]]];
+ nvmet_change($scfg, $template, op => 'create template', timeout => $timeout);
+ };
+ my $error = $@ or return;
+ nvmet_rethrow_abandoned($error);
+ chomp($error);
+
+ eval {
+ my ($rinv, $rcfs) = nvmet_read($scfg);
+ if (exists($rinv->{types}->{$new})) {
+ my ($units) = _nvmet_plan_unexport($rcfs, $nqn, $uuid);
+ my $back = [$units->@*, [['zfs', 'rename', $new, $old]]];
+ nvmet_change($scfg, $back, op => 'rename zvol back', timeout => $timeout);
+ (undef, $rcfs) = nvmet_read($scfg);
+ } elsif (!exists($rinv->{types}->{$old})) {
+ die "zvol '$old' does not exist\n";
+ }
+ nvmet_wait_device($scfg, $old_dev);
+ my ($units) = _nvmet_plan_export($rcfs, $nqn, $nsid, $uuid, $old_dev);
+ nvmet_change($scfg, $units, op => 'restore namespace');
+ };
+ die "template conversion failed: $error; restoration failed: $@" if $@;
+ die "template conversion failed: $error\n";
+ },
+ );
+ return;
+}
+
+# Grows an owned zvol and lets an exported namespace pick up the new size.
+my sub nvmet_resize_volume($scfg, $name, $size) {
+ my $nqn = nvmet_nqn($scfg);
+ my $dataset = nvmet_dataset($scfg, $name);
+ die "internal error: invalid zvol size\n" if $size !~ $RE_UNSIGNED_INTEGER;
+
+ # The volume is checked before it changes; under the lock, the namespace
+ # read here is still the one to revalidate after the resize. It is
+ # revalidated in the same call, so whenever the zvol grows, also after a
+ # lost reply, the hosts see the new size. A namespace that is not enabled
+ # reads the new size when it is enabled.
+ nvmet_locked(
+ $scfg,
+ sub {
+ my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1);
+ my (undef, $uuid) = _nvmet_owned_identity($inv, $nqn, $dataset);
+ my $namespaces = nvmet_namespaces($cfs, $nqn);
+ my @matches =
+ grep { $namespaces->{$_}->{device_uuid} eq $uuid } keys $namespaces->%*;
+ die "duplicate namespace UUID '$uuid'\n" if @matches > 1;
+ my @resize = (['zfs', 'set', "volsize=${size}k", $dataset]);
+ push @resize, nvmet_write(nvmet_ns_path($nqn, $matches[0]) . '/revalidate_size', 1)
+ if @matches && $namespaces->{ $matches[0] }->{enable} eq '1';
+ nvmet_change(
+ $scfg, [\@resize],
+ op => 'resize zvol',
+ timeout => nvmet_long_timeout(),
+ );
+ },
+ );
+ return;
+}
+
+# Removes the subsystem once it has no volumes and namespaces, then the hosts
+# that no subsystem links any more. Ports are never removed.
+sub _nvmet_delete_target($scfg) {
+ my $nqn = nvmet_nqn($scfg);
+ my $hostnqns = parse_nvme_host_nqns($scfg->{'nvme-host-nqns'}, 1) // [];
+
+ nvmet_locked(
+ $scfg,
+ sub {
+ my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1);
+ my ($units, $candidates) = _nvmet_plan_delete_target($inv, $cfs, $nqn, $hostnqns);
+ my $fresh = $cfs;
+ if ($units->@*) {
+ my $res = nvmet_apply($scfg, $units, op => 'remove NVMe subsystem');
+ (undef, $fresh) = eval { nvmet_read($scfg) };
+ my $read_error = $@;
+ nvmet_rethrow_abandoned($read_error) if $read_error;
+ die "cannot remove NVMe subsystem '$nqn': $res->{err}\n"
+ if $res->{rc} && (!$fresh || $fresh->{subsystems}->{$nqn});
+ die $read_error if !$fresh;
+ }
+ my $orphans = _nvmet_plan_orphan_hosts($fresh, $candidates);
+ return if !$orphans->@*;
+ my $res = nvmet_apply($scfg, $orphans, op => 'remove orphan NVMe hosts');
+ log_warn("could not remove orphan NVMe host(s): $res->{err}") if $res->{rc};
+ },
+ );
+ return;
+}
+
+# ---------------------------------------------------------------------------
+# Configuration
+# ---------------------------------------------------------------------------
+
+sub verify_nvme_nqn($value, $noerr = undef) {
+
+ if (
+ length($value) > 223
+ || $value !~ $RE_NQN
+ ) {
+ return undef if $noerr;
+ die "value is not a valid NVMe qualified name\n";
+ }
+
+ return $value;
+}
+
+sub parse_nvme_portals($value, $noerr = undef) {
+ my $result = [];
+ my $seen = {};
+
+ for my $entry (split(/,/, $value // '')) {
+ $entry = trim($entry);
+ my ($address, $port, $family);
+ if ($entry =~ $RE_IPV6_PORTAL) {
+ ($address, $port, $family) = ($+{address}, $+{port} // 4420, 'ipv6');
+ } elsif ($entry =~ $RE_IPV4_PORTAL) {
+ ($address, $port, $family) = ($+{address}, $+{port} // 4420, 'ipv4');
+ } else {
+ return undef if $noerr;
+ die "invalid NVMe/TCP portal '$entry'\n";
+ }
+
+ if (
+ !PVE::JSONSchema::pve_verify_ip($address, 1)
+ || ($family eq 'ipv4' && index($address, ':') >= 0)
+ || ($family eq 'ipv6' && index($address, ':') < 0)
+ || $port < 1
+ || $port > 65535
+ ) {
+ return undef if $noerr;
+ die "invalid NVMe/TCP portal '$entry'\n";
+ }
+ $address = nvmet_canonical_address($address) if $family eq 'ipv6';
+ # one listener: compare like _assert_unique_target
+ my $id = join(
+ "\0",
+ nvmet_listener_address({ family => $family, address => $address }),
+ int($port),
+ );
+ if ($seen->{$id}++) {
+ return undef if $noerr;
+ die "duplicate NVMe/TCP portal '$entry'\n";
+ }
+ push $result->@*,
+ {
+ address => $address,
+ port => int($port),
+ family => $family,
+ };
+ if (scalar($result->@*) > $max_paths) {
+ return undef if $noerr;
+ die "at most $max_paths NVMe/TCP portals are supported\n";
+ }
+ }
+
+ if (!$result->@*) {
+ return undef if $noerr;
+ die "at least one NVMe/TCP portal is required\n";
+ }
+
+ return $result;
+}
+
+my sub verify_nvme_portals($value, $noerr = undef) {
+ return undef if !parse_nvme_portals($value, $noerr);
+ return $value;
+}
+
+sub parse_nvme_host_ifaces($value, $noerr = undef) {
+ my $result = [];
+
+ for my $iface (split(/,/, $value // '')) {
+ $iface = trim($iface);
+ # the kernel never names an interface '.' or '..'
+ if (
+ length($iface) < 1
+ || length($iface) > 15
+ || $iface !~ $RE_HOST_IFACE
+ || $iface eq '.'
+ || $iface eq '..'
+ ) {
+ return undef if $noerr;
+ die "invalid NVMe/TCP host interface '$iface'\n";
+ }
+ push $result->@*, $iface;
+ }
+
+ if (!$result->@*) {
+ return undef if $noerr;
+ die "at least one NVMe/TCP host interface is required\n";
+ }
+
+ return $result;
+}
+
+sub parse_nvme_host_nqns($value, $noerr = undef) {
+ my $result = [];
+ my $seen = {};
+
+ for my $hostnqn (split(/,/, $value // '')) {
+ $hostnqn = trim($hostnqn);
+ if (!verify_nvme_nqn($hostnqn, 1) || $seen->{$hostnqn}++) {
+ return undef if $noerr;
+ die "invalid or duplicate NVMe host NQN '$hostnqn'\n";
+ }
+ push $result->@*, $hostnqn;
+ if (scalar($result->@*) > $max_hosts) {
+ return undef if $noerr;
+ die "at most $max_hosts NVMe host NQNs are supported\n";
+ }
+ }
+
+ if (!$result->@*) {
+ return undef if $noerr;
+ die "at least one NVMe host NQN is required\n";
+ }
+
+ return $result;
+}
+
+my sub verify_nvme_host_ifaces($value, $noerr = undef) {
+ return undef if !parse_nvme_host_ifaces($value, $noerr);
+ return $value;
+}
+
+my sub verify_nvme_host_nqns($value, $noerr = undef) {
+ return undef if !parse_nvme_host_nqns($value, $noerr);
+ return $value;
+}
+
+sub _configured_portals($scfg) {
+ my $portals = parse_nvme_portals($scfg->{'nvme-portals'});
+ my $ifaces = parse_nvme_host_ifaces($scfg->{'nvme-host-ifaces'});
+ die "nvme-host-ifaces must contain one interface for each nvme-portals entry\n"
+ if scalar($ifaces->@*) != scalar($portals->@*);
+
+ for (my $i = 0; $i < scalar($portals->@*); $i++) {
+ $portals->[$i]->{host_iface} = $ifaces->[$i];
+ }
+ return $portals;
+}
+
+sub _local_iface_exists($iface) {
+ return -d "/sys/class/net/$iface";
+}
+
+sub _validate_local_ifaces($portals) {
+ for my $portal ($portals->@*) {
+ my $iface = $portal->{host_iface};
+ die "NVMe/TCP host interface '$iface' does not exist on this node\n"
+ if !_local_iface_exists($iface);
+ }
+}
+
+PVE::JSONSchema::register_format('pve-storage-nvme-nqn', \&verify_nvme_nqn);
+PVE::JSONSchema::register_format('pve-storage-nvme-portals', \&verify_nvme_portals);
+PVE::JSONSchema::register_format('pve-storage-nvme-host-ifaces', \&verify_nvme_host_ifaces);
+PVE::JSONSchema::register_format('pve-storage-nvme-host-nqns', \&verify_nvme_host_nqns);
+
+sub type($class) {
+ return 'zfsnvme';
+}
+
+sub plugindata($class) {
+ return {
+ content => [{ images => 1 }, { images => 1 }],
+ 'sensitive-properties' => { 'dhchap-key' => 1 },
+ };
+}
+
+sub properties($class) {
+ return {
+ subsysnqn => {
+ description => "NVMe subsystem qualified name.",
+ type => 'string',
+ format => 'pve-storage-nvme-nqn',
+ },
+ 'nvme-portals' => {
+ description =>
+ "Comma-separated NVMe/TCP target IP addresses, at most 16. Put IPv6 addresses in"
+ . " brackets. The default TCP port is 4420.",
+ type => 'string',
+ format => 'pve-storage-nvme-portals',
+ maxLength => 2048,
+ },
+ 'nvme-host-ifaces' => {
+ description =>
+ "Comma-separated local interfaces, in portal order. Interface names must be"
+ . " identical on every cluster node.",
+ type => 'string',
+ format => 'pve-storage-nvme-host-ifaces',
+ maxLength => 512,
+ },
+ 'nvme-host-nqns' => {
+ description =>
+ "Comma-separated /etc/nvme/hostnqn values for every cluster node allowed to use"
+ . " this storage, at most 64. Host NQNs can only be added.",
+ type => 'string',
+ format => 'pve-storage-nvme-host-nqns',
+ maxLength => 8192,
+ },
+ 'dhchap-key' => {
+ description =>
+ "NVMe DH-HMAC-CHAP key in secret representation format (DHHC-1). Required when the"
+ . " storage is created; it cannot be changed.",
+ type => 'string',
+ maxLength => 256,
+ },
+ 'nvme-iopolicy' => {
+ description => "Native NVMe multipath I/O policy.",
+ type => 'string',
+ enum => ['numa', 'round-robin', 'queue-depth'],
+ default => 'round-robin',
+ },
+ 'nvme-keep-alive-tmo' => {
+ description => "NVMe keep-alive timeout in seconds.",
+ type => 'integer',
+ minimum => 1,
+ maximum => 120,
+ default => 5,
+ },
+ 'nvme-reconnect-delay' => {
+ description => "Delay between NVMe reconnect attempts in seconds.",
+ type => 'integer',
+ minimum => 1,
+ maximum => 120,
+ default => 2,
+ },
+ 'nvme-ctrl-loss-tmo' => {
+ description => "Time to keep retrying a lost NVMe controller in seconds.",
+ type => 'integer',
+ minimum => -1,
+ maximum => 86400,
+ default => 600,
+ },
+ 'nvme-fast-io-fail-tmo' => {
+ description =>
+ "Optional time before failing I/O on a reconnecting NVMe controller. Unset queues"
+ . " I/O until controller loss timeout.",
+ type => 'integer',
+ minimum => 0,
+ maximum => 86400,
+ optional => 1,
+ },
+ 'nvme-nr-io-queues' => {
+ description => "Number of NVMe/TCP I/O queues per controller.",
+ type => 'integer',
+ minimum => 1,
+ maximum => 1024,
+ optional => 1,
+ },
+ };
+}
+
+sub options($class) {
+ return {
+ server => { fixed => 1 },
+ subsysnqn => { fixed => 1 },
+ 'nvme-portals' => { fixed => 1 },
+ 'nvme-host-ifaces' => {},
+ 'nvme-host-nqns' => {},
+ pool => { fixed => 1 },
+ blocksize => { fixed => 1 },
+ sparse => { optional => 1 },
+ 'dhchap-key' => { optional => 1 },
+ 'nvme-iopolicy' => { optional => 1 },
+ 'nvme-keep-alive-tmo' => { optional => 1 },
+ 'nvme-reconnect-delay' => { optional => 1 },
+ 'nvme-ctrl-loss-tmo' => { optional => 1 },
+ 'nvme-fast-io-fail-tmo' => { optional => 1 },
+ 'nvme-nr-io-queues' => { optional => 1 },
+ nodes => { optional => 1 },
+ disable => { optional => 1 },
+ content => { optional => 1 },
+ bwlimit => { optional => 1 },
+ };
+}
+
+sub _validate_fail_fast_timeout($config, $default_ctrl_loss_tmo = undef) {
+ my $fast = $config->{'nvme-fast-io-fail-tmo'};
+ return if !defined($fast);
+
+ my $ctrl = $config->{'nvme-ctrl-loss-tmo'};
+ $ctrl = $default_ctrl_loss_tmo if !defined($ctrl);
+ return if !defined($ctrl) || $ctrl < 0;
+
+ die "nvme-fast-io-fail-tmo must not exceed nvme-ctrl-loss-tmo\n"
+ if $fast > $ctrl;
+}
+
+# Defaults are not injected here: this also runs when storage.cfg is parsed.
+# Every use site applies the schema default itself.
+sub check_config($class, $section_id, $config, $create, $skip_schema_check = undef) {
+ nvmet_pool($config) if defined($config->{pool});
+ verify_nvme_nqn($config->{subsysnqn}) if defined($config->{subsysnqn});
+ parse_nvme_portals($config->{'nvme-portals'}) if defined($config->{'nvme-portals'});
+ if (defined($config->{'nvme-host-ifaces'}) && defined($config->{'nvme-portals'})) {
+ _configured_portals($config);
+ } elsif (defined($config->{'nvme-host-ifaces'})) {
+ parse_nvme_host_ifaces($config->{'nvme-host-ifaces'});
+ }
+ parse_nvme_host_nqns($config->{'nvme-host-nqns'})
+ if defined($config->{'nvme-host-nqns'});
+ _validate_fail_fast_timeout($config, $create ? 600 : undef);
+ return $class->SUPER::check_config($section_id, $config, $create, $skip_schema_check);
+}
+
+# ---------------------------------------------------------------------------
+# ZFS helpers (same formats as the other ZFS plugins)
+# ---------------------------------------------------------------------------
+
+my sub zfs_request($scfg, $timeout, $method, @args) {
+ $timeout = PVE::RPCEnvironment->is_worker() ? 60 * 60 : 10 if !$timeout;
+ my $change = $method ne 'get' && $method ne 'list';
+ my $res = nvmet_exec(
+ $scfg, [['zfs', $method, @args]],
+ op => "zfs $method",
+ timeout => $timeout,
+ change => $change,
+ );
+ die "zfs error on '" . nvmet_server($scfg) . "': $res->{err}\n" if $res->{rc};
+ return nvmet_output($res);
+}
+
+# The direct children of the pool with PVE volume names. The plugin creates
+# and owns zvols only; a filesystem child holds its name.
+my sub zfs_parse_zvol_list($text, $pool) {
+ my $list = [];
+ for my $line (split /\n/, $text // '') {
+ my ($dataset, $size, $origin, $type) = split(/\s+/, $line);
+ next if !defined($type) || ($type ne 'volume' && $type ne 'filesystem');
+ my @parts = split /\//, $dataset;
+ next if @parts < 2;
+ my $name = pop @parts;
+ next if join('/', @parts) ne $pool;
+ next if $name !~ $RE_ZVOL_OWNER;
+ push $list->@*,
+ {
+ name => $name,
+ owner => $+{owner},
+ type => $type,
+ ($type eq 'volume' ? (size => $size + 0) : ()),
+ ($origin ne '-' ? (origin => $origin) : ()),
+ };
+ }
+ return $list;
+}
+
+# The pool's direct children with PVE volume names, so that a name is never
+# handed out twice; with $owned_only, the zvols owned by this storage's NVMe
+# subsystem.
+my sub zfs_list_zvol($scfg, $owned_only) {
+ my $pool = nvmet_pool($scfg);
+ my $text = zfs_request(
+ $scfg,
+ 10,
+ 'list',
+ '-o',
+ 'name,volsize,origin,type',
+ '-t',
+ 'volume,filesystem',
+ '-d1',
+ '-Hp',
+ $pool,
+ );
+ my $list = {};
+ my $prefix = "$pool/";
+ for my $zvol (zfs_parse_zvol_list($text, $pool)->@*) {
+ next if $owned_only && $zvol->{type} ne 'volume';
+ my $parent = $zvol->{origin};
+ # an origin in the pool is named relative to it, like the ZFS plugins do
+ $parent = substr($parent, length($prefix)) if $parent && index($parent, $prefix) == 0;
+ $list->{ $zvol->{name} } = {
+ name => $zvol->{name},
+ size => $zvol->{size},
+ parent => $parent,
+ format => 'raw',
+ vmid => $zvol->{owner},
+ };
+ }
+ return $list if !$owned_only;
+
+ my $properties = zfs_request(
+ $scfg,
+ 10,
+ 'get',
+ '-H',
+ '-d',
+ '1',
+ '-o',
+ 'name,value,source',
+ 'proxmox:nvme-subsys',
+ $pool,
+ );
+ my $owned = {};
+ for my $line (split(/\n/, $properties)) {
+ my ($dataset, $nqn, $source) = split(/\t/, $line, 3);
+ next if !defined($source) || ($source ne 'local' && $source ne 'received');
+ next if !defined($nqn) || $nqn ne $scfg->{subsysnqn};
+ next if index($dataset, $prefix) != 0;
+ $owned->{ substr($dataset, length($prefix)) } = 1;
+ }
+ for my $name (keys $list->%*) {
+ delete $list->{$name} if !$owned->{$name};
+ }
+ return $list;
+}
+
+my sub zfs_get_properties($scfg, $properties, $dataset, $timeout = undef) {
+ my $text = zfs_request($scfg, $timeout, 'get', '-o', 'value', '-Hp', $properties, $dataset);
+ my @values = split /\n/, $text;
+ return wantarray ? @values : $values[0];
+}
+
+my sub zfs_get_sorted_snapshot_list($scfg, $name, $sort_params) {
+ return [
+ map { s/^.*\@//r } split /\n/,
+ zfs_request(
+ $scfg,
+ undef,
+ 'list',
+ '-H',
+ '-r',
+ '-t',
+ 'snapshot',
+ '-o',
+ 'name',
+ $sort_params->@*,
+ nvmet_dataset($scfg, $name),
+ ),
+ ];
+}
+
+# ---------------------------------------------------------------------------
+# Storage API: volumes
+# ---------------------------------------------------------------------------
+
+sub parse_volname($class, $volname) {
+ if ($volname =~ $RE_VOLNAME) {
+ my $format = ($+{type} eq 'subvol' || $+{type} eq 'basevol') ? 'subvol' : 'raw';
+ my $is_base = $+{type} eq 'base' || $+{type} eq 'basevol';
+ return ('images', $+{name}, $+{vmid}, $+{base}, $+{base_vmid}, $is_base, $format);
+ }
+ die "unable to parse zfs volume name '$volname'\n";
+}
+
+sub list_images($class, $storeid, $scfg, $vmid = undef, $vollist = undef, $cache = undef) {
+ my $res = [];
+ for my $info (values zfs_list_zvol($scfg, 1)->%*) {
+ my $volname =
+ $info->{parent} && $info->{parent} =~ $RE_BASE_SNAPSHOT
+ ? "$storeid:$+{base}/$info->{name}"
+ : "$storeid:$info->{name}";
+ next
+ if $vollist
+ ? !grep { $_ eq $volname } $vollist->@*
+ : defined($vmid) && $info->{vmid} ne $vmid;
+ $info->{volid} = $volname;
+ push $res->@*, $info;
+ }
+ return $res;
+}
+
+# Names are allocated against every direct child of the pool, owned or not, so
+# a leftover zvol that list_images hides is never handed out again.
+sub find_free_diskname($class, $storeid, $scfg, $vmid, $fmt = undef, $add_fmt_suffix = undef) {
+ my $disk_list = [map { "$storeid:$_" } sort keys zfs_list_zvol($scfg, 0)->%*];
+ return PVE::Storage::Plugin::get_next_vm_diskname(
+ $disk_list, $storeid, $vmid, $fmt, $scfg, $add_fmt_suffix,
+ );
+}
+
+sub status($class, $storeid, $scfg, $cache = undef) {
+ my ($available, $used) =
+ eval { zfs_get_properties($scfg, 'available,used', nvmet_pool($scfg)) };
+ if (my $err = $@) {
+ warn "storage '$storeid': $err";
+ return (0, 0, 0, 0);
+ }
+ if (
+ !defined($available)
+ || !defined($used)
+ || $available !~ $RE_UNSIGNED_INTEGER
+ || $used !~ $RE_UNSIGNED_INTEGER
+ ) {
+ warn "unexpected ZFS pool usage for storage '$storeid'\n";
+ return (0, 0, 0, 0);
+ }
+ return ($available + $used, $available, $used, 1);
+}
+
+sub volume_size_info($class, $scfg, $storeid, $volname, $timeout = undef) {
+ my (undef, $name, undef, $parent, undef, undef, $format) = $class->parse_volname($volname);
+ die "volume_size_info requires a ZFS volume\n" if $format ne 'raw';
+ my ($size, $used) = zfs_get_properties(
+ $scfg, 'volsize,usedbydataset', nvmet_dataset($scfg, $name), $timeout,
+ );
+ die "Could not get zfs volume size\n" if !defined($size) || $size !~ $RE_UNSIGNED_INTEGER;
+ $used = defined($used) && $used =~ $RE_UNSIGNED_INTEGER ? $used + 0 : 0;
+ return wantarray ? ($size + 0, 'raw', $used, $parent) : $size + 0;
+}
+
+sub volume_snapshot($class, $scfg, $storeid, $volname, $snap) {
+ my (undef, $name, undef, undef, undef, undef, $format) = $class->parse_volname($volname);
+ die "volume_snapshot requires a ZFS volume\n" if $format ne 'raw';
+ my $snapshot = nvmet_snapshot($scfg, $name, $snap);
+ nvmet_locked($scfg, sub { zfs_request($scfg, undef, 'snapshot', $snapshot) });
+ return;
+}
+
+sub volume_snapshot_delete($class, $scfg, $storeid, $volname, $snap, $running = undef) {
+ my $name = ($class->parse_volname($volname))[1];
+ my $snapshot = nvmet_snapshot($scfg, $name, $snap);
+ nvmet_locked($scfg, sub { zfs_request($scfg, undef, 'destroy', $snapshot) });
+ return;
+}
+
+sub volume_rollback_is_possible($class, $scfg, $storeid, $volname, $snap, $blockers = undef) {
+ my $name = ($class->parse_volname($volname))[1];
+ nvmet_snapshot($scfg, $name, $snap); # validation only
+ my $found;
+ $blockers //= [];
+ for my $snapshot (zfs_get_sorted_snapshot_list($scfg, $name, ['-s', 'creation'])->@*) {
+ $found = 1 if $snapshot eq $snap;
+ push $blockers->@*, $snapshot if $found && $snapshot ne $snap;
+ }
+ die "can't rollback, snapshot '$snap' does not exist on '${storeid}:${volname}'\n" if !$found;
+ die "can't rollback, '$snap' is not most recent snapshot on '${storeid}:${volname}'\n"
+ if $blockers->@*;
+ return 1;
+}
+
+sub volume_snapshot_rollback($class, $scfg, $storeid, $volname, $snap) {
+ my (undef, $name, undef, undef, undef, undef, $format) = $class->parse_volname($volname);
+ die "snapshot rollback requires a ZFS volume\n" if $format ne 'raw';
+ nvmet_rollback_volume($scfg, $name, $snap);
+}
+
+sub volume_snapshot_info($class, $scfg, $storeid, $volname) {
+ my $name = ($class->parse_volname($volname))[1];
+ my $info = {};
+ my $text = zfs_request(
+ $scfg,
+ undef,
+ 'list',
+ '-Hp',
+ '-r',
+ '-t',
+ 'snapshot',
+ '-o',
+ 'name,guid,creation',
+ nvmet_dataset($scfg, $name),
+ );
+ for my $line (split /\n/, $text) {
+ my ($snapshot, $guid, $creation) = split /\s+/, $line;
+ $snapshot =~ s/^.*\@//;
+ $info->{$snapshot} = { id => $guid, timestamp => $creation };
+ }
+ return $info;
+}
+
+sub free_image($class, $storeid, $scfg, $volname, $is_base = undef, $format = undef) {
+ my (undef, $name, undef, undef, undef, undef, $volformat) = $class->parse_volname($volname);
+ die "free_image requires a ZFS volume\n" if $volformat ne 'raw';
+ nvmet_destroy_volume($scfg, $name);
+ return undef;
+}
+
+sub create_base($class, $storeid, $scfg, $volname) {
+ my (undef, $name, undef, $basename, undef, $is_base, $format) = $class->parse_volname($volname);
+ die "create_base not possible with base image\n" if $is_base;
+ die "create_base requires a ZFS volume\n" if $format ne 'raw';
+ my $newname = $name =~ s/^vm-/base-/r;
+ nvmet_template_volume($scfg, $name);
+ return $basename ? "$basename/$newname" : $newname;
+}
+
+sub clone_image($class, $scfg, $storeid, $volname, $vmid, $snap = undef) {
+ $snap ||= '__base__';
+ my (undef, $basename, undef, undef, undef, $is_base, $format) = $class->parse_volname($volname);
+ die "clone_image only works on base images\n" if !$is_base;
+ die "clone_image requires a ZFS volume\n" if $format ne 'raw';
+ my $origin = nvmet_snapshot($scfg, $basename, $snap);
+ my $name = $class->find_free_diskname($storeid, $scfg, $vmid, $format);
+ nvmet_create_volume($scfg, $name, origin => $origin);
+ return "$basename/$name";
+}
+
+sub alloc_image($class, $storeid, $scfg, $vmid, $fmt, $name, $size) {
+ die "unsupported format '$fmt'" if $fmt ne 'raw';
+ die "illegal name '$name' - should be 'vm-$vmid-*'\n"
+ if $name && index($name, "vm-$vmid-") != 0;
+ my $volname = $name || $class->find_free_diskname($storeid, $scfg, $vmid, $fmt);
+ $size += (1024 - $size % 1024) % 1024;
+ nvmet_create_volume($scfg, $volname, size => $size);
+ return $volname;
+}
+
+sub volume_resize($class, $scfg, $storeid, $volname, $size, $running = undef, $snapname = undef) {
+ # QEMU can resize a host_device once the device itself has grown, but the
+ # plugin cannot yet wait for the local NVMe namespace to report the new
+ # size. Until it can, refuse online resize before changing the zvol, so a
+ # running VM never sees a size different from its configuration.
+ die "online resize is not supported for NVMe/TCP block devices; stop the VM first\n"
+ if $running;
+ die "resizing a snapshot is not supported for $class\n" if $snapname;
+ my $new_size = int($size / 1024);
+ $new_size += (1024 - $new_size % 1024) % 1024;
+ my $name = ($class->parse_volname($volname))[1];
+ nvmet_resize_volume($scfg, $name, $new_size);
+ return $new_size;
+}
+
+sub volume_has_feature(
+ $class, $scfg, $feature, $storeid, $volname,
+ $snapname = undef,
+ $running = undef,
+ $opts = undef,
+) {
+ my $features = {
+ snapshot => { current => 1, snap => 1 },
+ clone => { base => 1 },
+ template => { current => 1 },
+ copy => { base => 1, current => 1 },
+ };
+ my $is_base = ($class->parse_volname($volname))[5];
+ my $key = $snapname ? 'snap' : $is_base ? 'base' : 'current';
+ return $features->{$feature}->{$key};
+}
+
+# No stream transfers or renames; the storage has no path, so the base class
+# offers no export or import format either.
+sub volume_export($class, @args) {
+ die "ZFS stream export is not supported for NVMe/TCP storage\n";
+}
+
+sub volume_import($class, @args) {
+ die "ZFS stream import is not supported for NVMe/TCP storage\n";
+}
+
+sub rename_volume($class, @args) {
+ die "renaming volumes is not supported for NVMe/TCP storage\n";
+}
+
+sub rename_snapshot($class, @args) {
+ die "rename_snapshot is not supported for $class";
+}
+
+# ---------------------------------------------------------------------------
+# Secrets and configuration hooks
+# ---------------------------------------------------------------------------
+
+# Each storage needs its own NVMe/TCP listener (address family, address and
+# port), and a disabled storage keeps its listener. nvmet answers a connect to
+# a listener that does not publish the subsystem yet with DNR, and the host
+# then deletes the controller whatever its loss timeout: after a target
+# restart, the storage published second would lose its paths. Storages that
+# reach the same data address must spell the server alike, as it names the
+# target lock.
+sub _assert_unique_target($storeid, $scfg, $cfg = undef) {
+ $cfg //= PVE::Storage::config();
+ my (%listeners, %addresses);
+ for my $portal (parse_nvme_portals($scfg->{'nvme-portals'})->@*) {
+ my $address = join("\0", nvmet_listener_address($portal));
+ $listeners{"$address\0$portal->{port}"} = 1;
+ $addresses{$address} = 1;
+ }
+ my $ids = $cfg->{ids} // {};
+ for my $other_id (sort keys $ids->%*) {
+ next if $other_id eq $storeid;
+ my $other = $ids->{$other_id};
+ next if ($other->{type} // '') ne 'zfsnvme';
+
+ die "NVMe subsystem NQN is already used by storage '$other_id'\n"
+ if ($other->{subsysnqn} // '') eq $scfg->{subsysnqn};
+ my $other_server = $other->{server} // '';
+ # a storage whose portals do not parse cannot be activated
+ for my $portal ((parse_nvme_portals($other->{'nvme-portals'}, 1) // [])->@*) {
+ my $address = join("\0", nvmet_listener_address($portal));
+ die "NVMe/TCP portal '$portal->{address}' port $portal->{port} is already used by"
+ . " storage '$other_id'; give each storage its own address or port\n"
+ if $listeners{"$address\0$portal->{port}"};
+ die "storage '$other_id' reaches target address '$portal->{address}' through server"
+ . " '$other_server'; use the same server value for both storages\n"
+ if $addresses{$address} && $other_server ne $scfg->{server};
+ }
+ next if $other_server ne $scfg->{server};
+ my ($pool, $other_pool) = ($scfg->{pool}, $other->{pool} // '');
+ die "ZFS pool '$pool' on '$scfg->{server}' is already used by storage '$other_id'\n"
+ if $other_pool eq $pool;
+ die "ZFS pool '$pool' on '$scfg->{server}' overlaps the pool of storage '$other_id'\n"
+ if index("$other_pool/", "$pool/") == 0 || index("$pool/", "$other_pool/") == 0;
+ }
+}
+
+# nvmet keeps one DH-HMAC-CHAP key per host NQN, so storages on the same target
+# whose host NQNs overlap must use the same key.
+sub _assert_shared_host_key($storeid, $scfg, $key, $cfg = undef) {
+ $cfg //= PVE::Storage::config();
+ my %hosts = map { $_ => 1 } (parse_nvme_host_nqns($scfg->{'nvme-host-nqns'}, 1) // [])->@*;
+ my $ids = $cfg->{ids} // {};
+ for my $other_id (sort keys $ids->%*) {
+ next if $other_id eq $storeid;
+ my $other = $ids->{$other_id};
+ next if ($other->{type} // '') ne 'zfsnvme';
+ next if ($other->{server} // '') ne ($scfg->{server} // '');
+ my $other_hosts = parse_nvme_host_nqns($other->{'nvme-host-nqns'}, 1) // [];
+ next if !grep { $hosts{$_} } $other_hosts->@*;
+ my $other_key = file_read_firstline(secret_path($other_id));
+ die "storage '$other_id' uses a different DH-HMAC-CHAP key for the same NVMe host NQNs"
+ . " on '$scfg->{server}'\n"
+ if defined($other_key) && $other_key ne $key;
+ }
+}
+
+# DHHC-1:<hash>:<base64 of key and its little-endian CRC-32>: as generated by
+# `nvme gen-dhchap-key`. Messages never contain the key.
+sub _validate_secret($key) {
+ die "missing NVMe DH-HMAC-CHAP key\n" if !defined($key) || $key eq '';
+ my $invalid = "invalid NVMe DH-HMAC-CHAP key representation\n";
+ die $invalid if $key !~ $RE_DHCHAP_KEY;
+ my ($hash, $encoded) = @+{qw(hash secret)};
+ die $invalid if length($encoded) % 4;
+ my $decoded = decode_base64($encoded);
+ my $length = length($decoded) - 4;
+ my %lengths = ('00' => [32, 48, 64], '01' => [32], '02' => [48], '03' => [64]);
+ die $invalid if !grep { $_ == $length } $lengths{$hash}->@*;
+ die $invalid if pack('V', crc32(substr($decoded, 0, $length))) ne substr($decoded, $length);
+ return $key;
+}
+
+my sub set_secret($storeid, $key) {
+ _validate_secret($key);
+ make_path($secret_dir, { mode => 0700 });
+ file_set_contents(secret_path($storeid), "$key\n", 0600);
+}
+
+my sub get_secret($storeid) {
+ my $key = file_read_firstline(secret_path($storeid));
+ return _validate_secret($key);
+}
+
+# Removes the key file of a storage (a test seam).
+sub _unlink_file($path) {
+ return unlink($path);
+}
+
+my sub delete_secret($storeid) {
+ _unlink_file(secret_path($storeid));
+}
+
+sub on_add_hook($class, $storeid, $scfg, %sensitive) {
+ _configured_portals($scfg);
+ parse_nvme_host_nqns($scfg->{'nvme-host-nqns'});
+ _assert_unique_target($storeid, $scfg);
+ my $key = _validate_secret($sensitive{'dhchap-key'});
+ _assert_shared_host_key($storeid, $scfg, $key);
+ set_secret($storeid, $key);
+ return;
+}
+
+sub on_update_hook_full($class, $storeid, $scfg, $update, $delete = undef, $sensitive = undef) {
+ $sensitive //= {};
+ my %prospective = ($scfg->%*, $update->%*);
+ delete @prospective{ $delete->@* } if $delete;
+ verify_nvme_nqn($prospective{subsysnqn});
+ _configured_portals(\%prospective);
+ my $new_hostnqns = parse_nvme_host_nqns($prospective{'nvme-host-nqns'});
+ my %new_hosts = map { $_ => 1 } $new_hostnqns->@*;
+ for my $old_hostnqn (parse_nvme_host_nqns($scfg->{'nvme-host-nqns'})->@*) {
+ die "removing NVMe host NQN '$old_hostnqn' is not supported; move every volume off"
+ . " the storage and recreate it with a new subsystem NQN and a new DH-HMAC-CHAP"
+ . " key\n"
+ if !$new_hosts{$old_hostnqn};
+ }
+ _validate_fail_fast_timeout(\%prospective, 600);
+ _assert_unique_target($storeid, \%prospective);
+
+ my $old_key = file_read_firstline(secret_path($storeid));
+ my $key = exists($sensitive->{'dhchap-key'}) ? $sensitive->{'dhchap-key'} : $old_key;
+ _validate_secret($key);
+
+ # nvmet stores authentication on the global Host NQN object, which can be
+ # shared by multiple subsystems. Replacing an active key in place can make
+ # unrelated storages unrecoverable on their next reconnect. Until PVE can
+ # coordinate a rolling rotation on every node, fail before mutating the
+ # cluster-wide secret.
+ die "NVMe DH-HMAC-CHAP key rotation is not supported; create a new storage with a new"
+ . " subsystem NQN\n"
+ if defined($old_key) && $key ne $old_key;
+ _assert_shared_host_key($storeid, \%prospective, $key);
+
+ set_secret($storeid, $key) if exists($sensitive->{'dhchap-key'});
+ return;
+}
+
+# ---------------------------------------------------------------------------
+# Local NVMe host side
+# ---------------------------------------------------------------------------
+
+my %slow_path_last; # storeid => time of the last slow-path activation
+
+# Probes of the local host (test seams).
+sub _block_device($path) {
+ return -b $path;
+}
+
+sub _link_target($path) {
+ return readlink($path);
+}
+
+sub _fabrics_device() {
+ return '/dev/nvme-fabrics';
+}
+
+# Runs $code in a child and gives up on it after $timeout seconds. Returns the
+# child's result; the child's errors pass through as exceptions, a timeout
+# becomes one. The bound is not hard: as with PVE::Tools::run_command, a
+# child stuck in the kernel (a write or delete in D state) is only reaped
+# once the kernel returns, and a timeout does not mean that the operation did
+# not happen.
+#
+# A stopped task (TERM, also HUP, INT and QUIT) stays pending in the parent
+# while the child runs, and the child inherits the mask: a connect that ends
+# before PVE escalates the stop to KILL, 5 seconds after the TERM, creates
+# the controller and restricts its secret attributes without interruption,
+# and run_fork_with_timeout, whose own handlers would consume the signal when
+# the child delivers a result, never sees it. Once the child is reaped and the
+# mask is restored, the handler of the worker dies with the task marker, also
+# over a result or an error of the child. ALRM, the timeout, is not blocked.
+# The KILL ends the parent and the child together; a controller that the
+# child had created but not restricted yet keeps its secret attributes
+# world-readable until the next activation restricts them.
+#
+# run_fork_with_timeout is called in list context on purpose: only there does
+# run_with_timeout report a timeout as its second return value; in scalar
+# context the timeout would be a warning and an undef result.
+my sub run_bounded($timeout, $what, $code) {
+ my $previous = POSIX::SigSet->new();
+ sigprocmask(SIG_BLOCK, POSIX::SigSet->new(SIGHUP, SIGINT, SIGQUIT, SIGTERM), $previous)
+ or die "cannot block signals: $!\n";
+ my $outer_warn = $SIG{__WARN__};
+ my ($res, $timed_out) = eval {
+ # Only a child that leaves no result makes run_fork_with_timeout warn
+ # in the parent; that case is an exception below. The child keeps the
+ # handler of the caller, so its own warnings reach the task log.
+ local $SIG{__WARN__} = sub($message) { };
+ PVE::Tools::run_fork_with_timeout(
+ $timeout,
+ sub {
+ local $SIG{__WARN__} = $outer_warn;
+ return $code->();
+ },
+ );
+ };
+ my $error = $@;
+ sigprocmask(SIG_SETMASK, $previous) or die "cannot restore signals: $!\n";
+ if ($error) {
+ rethrow_task_interrupt($error);
+ die $error;
+ }
+ die "$what did not complete within $timeout seconds\n" if $timed_out;
+ die "$what left no result\n" if !defined($res);
+ return $res;
+}
+
+# The connect options of one controller, as the fabrics device takes them,
+# and their names. The values are validated configuration and host identities
+# checked by the caller; the kernel reads each one up to the next ','.
+sub _fabrics_options($scfg, $portal, $hostnqn, $hostid, $key) {
+ my @options = (
+ [transport => 'tcp'],
+ [traddr => $portal->{address}],
+ [trsvcid => $portal->{port}],
+ [host_iface => $portal->{host_iface}],
+ [nqn => $scfg->{subsysnqn}],
+ [hostnqn => $hostnqn],
+ [hostid => $hostid],
+ [dhchap_secret => $key],
+ [keep_alive_tmo => $scfg->{'nvme-keep-alive-tmo'} // 5],
+ [reconnect_delay => $scfg->{'nvme-reconnect-delay'} // 2],
+ [ctrl_loss_tmo => $scfg->{'nvme-ctrl-loss-tmo'} // 600],
+ );
+ # unlike unset, which leaves fast I/O failure off, 0 is a real timeout
+ push @options, [fast_io_fail_tmo => $scfg->{'nvme-fast-io-fail-tmo'}]
+ if defined($scfg->{'nvme-fast-io-fail-tmo'});
+ push @options, [nr_io_queues => $scfg->{'nvme-nr-io-queues'}]
+ if defined($scfg->{'nvme-nr-io-queues'});
+ for my $option (@options) {
+ die "internal error: invalid NVMe connect option '$option->[0]'\n"
+ if !defined($option->[1]) || $option->[1] !~ $RE_FABRICS_VALUE;
+ }
+ return (join(',', map { "$_->[0]=$_->[1]" } @options), [map { $_->[0] } @options]);
+}
+
+# The child of a connect (a test seam): creates exactly one controller with
+# one write of $options to the fabrics device and returns its instance. The
+# write is never repeated, because the kernel may have created the controller
+# even when it reports an error. Errors carry the errno text of the failed
+# operation and never the options, which hold the key.
+sub _fabrics_connect($options, $names) {
+ my $device = _fabrics_device();
+ my $probe;
+ if (!sysopen($probe, $device, O_RDONLY | O_NOFOLLOW)) {
+ die "'$device' does not exist; load the nvme-tcp kernel module\n" if $! == ENOENT;
+ die "cannot open '$device': $!\n";
+ }
+ my @stat = stat($probe);
+ die "'$device' is not a character device\n" if !@stat || !S_ISCHR($stat[2]);
+ # without a controller, a read lists the supported options as patterns
+ my $supported = '';
+ sysread($probe, $supported, 4096) // die "cannot read '$device': $!\n";
+ close($probe);
+ my %supported = map { s/=.*//r => 1 } split(/,/, $supported =~ s/\n\z//r);
+ for my $name ($names->@*) {
+ die "the running kernel does not support the NVMe connect option '$name'\n"
+ if !$supported{$name};
+ }
+
+ # Not PVE::SysFSTools::file_write: the kernel returns the instance of the
+ # new controller on a read of the file descriptor that wrote the options.
+ sysopen(my $fh, $device, O_RDWR | O_NOFOLLOW) or die "cannot open '$device': $!\n";
+ my $written = _fabrics_write($fh, $options);
+ die "connect failed: $!\n" if !defined($written);
+ die "connect failed: short write\n" if $written != length($options);
+ my $result = '';
+ sysread($fh, $result, 128) // die "cannot read the connect result: $!\n";
+ die "unexpected connect result\n" if $result !~ $RE_FABRICS_RESULT;
+ my $instance = int($+{instance});
+ # The kernel creates the secret attributes of a controller world-readable;
+ # the caller keeps a stopped task from interrupting before this, short of
+ # a KILL.
+ _restrict_attr("/sys/class/nvme/nvme$instance/$_") for qw(dhchap_secret dhchap_ctrl_secret);
+ close($fh);
+ return $instance;
+}
+
+# The write that creates a controller (a test seam).
+sub _fabrics_write($fh, $options) {
+ return syswrite($fh, $options);
+}
+
+# Writes a sysfs attribute and returns the errno text of a failure, or undef.
+# PVE::SysFSTools::file_write returns undef without a warning when the open
+# fails, leaving its errno in $!, and 0 after warning "error writing ...:
+# <errno text>" when the write fails; that warning is not passed on. With
+# $missing_ok, an attribute that does not exist, because its controller went
+# away, is not a failure.
+my sub write_attribute($path, $value, $missing_ok = 0) {
+ my $warning;
+ my $written = do {
+ local $SIG{__WARN__} = sub($message) { $warning = $message };
+ PVE::SysFSTools::file_write($path, $value);
+ };
+ return if $written;
+ return if !defined($warning) && $missing_ok && $! == ENOENT;
+ return "$!" if !defined($warning);
+ return $warning =~ s/\A.*: //sr =~ s/\n\z//r;
+}
+
+# Deletes controllers (the body of a child, a test seam): the deletion can
+# block in the kernel. The error names the controller and the errno text; a
+# controller that is already gone counts as deleted.
+sub _delete_controllers($controllers) {
+ for my $controller ($controllers->@*) {
+ my $path = "/sys/class/nvme/$controller->{name}/delete_controller";
+ my $error = write_attribute($path, '1', 1) // next;
+ die "cannot delete NVMe controller '$controller->{name}': $error\n";
+ }
+ return scalar($controllers->@*);
+}
+
+# Every local controller of the subsystem.
+my sub nvme_controllers($nqn) {
+ my $controllers = [];
+ PVE::File::dir_glob_foreach(
+ '/sys/class/nvme',
+ $RE_NVME_CONTROLLER,
+ sub($entry) {
+ my $base = "/sys/class/nvme/$entry";
+ my $subsys = file_read_firstline("$base/subsysnqn");
+ return if !defined($subsys) || $subsys ne $nqn;
+ my $address = file_read_firstline("$base/address") // '';
+ my $traddr = $address =~ $RE_TRADDR ? $+{value} : undef;
+ my $trsvcid = $address =~ $RE_TRSVCID ? $+{value} : undef;
+ my $host_iface = $address =~ $RE_HOST_IFACE_ADDRESS ? $+{value} : undef;
+ push $controllers->@*,
+ {
+ name => $entry,
+ state => file_read_firstline("$base/state") // 'unknown',
+ traddr => nvmet_canonical_address($traddr),
+ trsvcid => $trsvcid,
+ host_iface => $host_iface,
+ };
+ },
+ );
+ return $controllers;
+}
+
+# Controllers by "address:port", with canonical IPv6 addresses. A portal has
+# more than one controller while its path moves to another interface.
+my sub controller_states($nqn) {
+ my $states = {};
+ for my $controller (nvme_controllers($nqn)->@*) {
+ next if !defined($controller->{traddr}) || !defined($controller->{trsvcid});
+ push $states->{"$controller->{traddr}:$controller->{trsvcid}"}->@*, $controller;
+ }
+ return $states;
+}
+
+# The multipath namespace devices of the subsystem and their partitions. One
+# dir_glob_foreach per directory level, because PVE::File::dir_glob_regex
+# returns only the first match.
+my sub namespace_devices($nqn) {
+ my $devices = {};
+
+ PVE::File::dir_glob_foreach(
+ '/sys/class/nvme-subsystem',
+ $RE_NVME_SUBSYSTEM,
+ sub($entry) {
+ my $base = "/sys/class/nvme-subsystem/$entry";
+ my $subsys = file_read_firstline("$base/subsysnqn");
+ return if !defined($subsys) || $subsys ne $nqn;
+
+ PVE::File::dir_glob_foreach(
+ $base,
+ $RE_NVME_NAMESPACE,
+ sub($device) {
+ $devices->{"/dev/$device"} = $device if _block_device("/dev/$device");
+ PVE::File::dir_glob_foreach(
+ "/sys/class/block/$device",
+ $RE_NVME_PARTITION,
+ sub($partition) {
+ $devices->{"/dev/$partition"} = $partition
+ if _block_device("/dev/$partition");
+ },
+ );
+ },
+ );
+ },
+ );
+ return $devices;
+}
+
+sub _namespace_openers($nqn) {
+ my $devices = namespace_devices($nqn);
+ return [] if !$devices->%*;
+
+ my %openers;
+ PVE::File::dir_glob_foreach(
+ '/proc',
+ '[0-9]+',
+ sub($pid) {
+ PVE::File::dir_glob_foreach(
+ "/proc/$pid/fd",
+ '[0-9]+',
+ sub($fd) {
+ my $target = _link_target("/proc/$pid/fd/$fd");
+ return if !defined($target) || !exists($devices->{$target});
+ my $comm = eval { file_read_firstline("/proc/$pid/comm") } // 'unknown';
+ $openers{"$pid:$target"} = "$comm (PID $pid, $target)";
+ },
+ );
+ },
+ );
+
+ # Kernel consumers such as device-mapper do not necessarily keep a userspace
+ # file descriptor open, but expose their dependency in the holders directory.
+ for my $path (keys $devices->%*) {
+ my $device = $devices->{$path};
+ my $holders = "/sys/class/block/$device/holders";
+ PVE::File::dir_glob_foreach(
+ $holders,
+ '[^\\.].*',
+ sub($holder) {
+ $openers{"holder:$device:$holder"} = "$path held by $holder";
+ },
+ );
+ }
+
+ return [sort values %openers];
+}
+
+my sub portal_reachable($portal) {
+ return PVE::Network::tcp_ping($portal->{address}, $portal->{port}, 2) // 0;
+}
+
+# The number of portals with a live controller on their configured interface,
+# or, with $any_iface, on any interface: such a path carries I/O, even while
+# it waits to be moved to its configured interface.
+sub _live_portal_count($states, $portals, $any_iface = 0) {
+ my $live = 0;
+ for my $portal ($portals->@*) {
+ my $id = "$portal->{address}:$portal->{port}";
+ $live++ if grep {
+ $_->{state} eq 'live'
+ && ($any_iface || ($_->{host_iface} // '') eq $portal->{host_iface})
+ } ($states->{$id} // [])->@*;
+ }
+ return $live;
+}
+
+# Creates the controller of a portal in a child, giving up after 10 seconds.
+sub _connect_portal($scfg, $portal, $hostnqn, $hostid, $key) {
+ my ($options, $names) = _fabrics_options($scfg, $portal, $hostnqn, $hostid, $key);
+ return run_bounded(10, 'NVMe/TCP connect', sub { return _fabrics_connect($options, $names) });
+}
+
+my sub set_iopolicy($nqn, $policy) {
+ PVE::File::dir_glob_foreach(
+ '/sys/class/nvme-subsystem',
+ $RE_NVME_SUBSYSTEM,
+ sub($entry) {
+ my $base = "/sys/class/nvme-subsystem/$entry";
+ my $subsys = file_read_firstline("$base/subsysnqn");
+ return if !defined($subsys) || $subsys ne $nqn;
+ my $error = write_attribute("$base/iopolicy", "$policy\n");
+ die "unable to set NVMe multipath policy: $error\n" if defined($error);
+ },
+ );
+}
+
+# Applies the reconnect timeouts to connected controllers, which otherwise
+# keep the values of their connect. sysfs shows "off" for -1, and shows the
+# loss timeout rounded up to a multiple of the reconnect delay.
+my sub set_tunables($scfg) {
+ my $delay = $scfg->{'nvme-reconnect-delay'} // 2;
+ my $loss = $scfg->{'nvme-ctrl-loss-tmo'} // 600;
+ my $fast = $scfg->{'nvme-fast-io-fail-tmo'};
+ my $shown_loss = $loss < 0 ? 'off' : int(($loss + $delay - 1) / $delay) * $delay;
+ my $shown_fast = defined($fast) && $fast >= 0 ? $fast : 'off';
+
+ for my $controller (nvme_controllers($scfg->{subsysnqn})->@*) {
+ my $base = "/sys/class/nvme/$controller->{name}";
+ my $current_delay = file_read_firstline("$base/reconnect_delay") // next;
+ my $write = sub($attr, $value) {
+ my $error = write_attribute("$base/$attr", "$value\n") // return;
+ log_warn("cannot set $attr of NVMe controller '$controller->{name}': $error");
+ };
+ my $delay_changed = $current_delay ne "$delay";
+ $write->('reconnect_delay', $delay) if $delay_changed;
+ # the loss timeout is stored as a number of reconnects of the delay
+ $write->('ctrl_loss_tmo', $loss < 0 ? -1 : $loss)
+ if $delay_changed
+ || (file_read_firstline("$base/ctrl_loss_tmo") // '') ne "$shown_loss";
+ $write->('fast_io_fail_tmo', $shown_fast eq 'off' ? -1 : $fast)
+ if (file_read_firstline("$base/fast_io_fail_tmo") // '') ne "$shown_fast";
+ }
+}
+
+# Attempts to make a controller attribute readable by root only (a test seam).
+sub _restrict_attr($path) {
+ my @stat = stat($path);
+ if (!@stat) {
+ log_warn("cannot stat NVMe controller attribute '$path': $!") if $! != ENOENT;
+ return;
+ }
+ return if !($stat[2] & 077);
+ chmod(0600, $path) or log_warn("cannot restrict permissions of '$path': $!");
+}
+
+# The kernel creates the DH-HMAC-CHAP secret attributes world-readable.
+my sub restrict_secret_attrs($nqn) {
+ for my $controller (nvme_controllers($nqn)->@*) {
+ _restrict_attr("/sys/class/nvme/$controller->{name}/$_")
+ for qw(dhchap_secret dhchap_ctrl_secret);
+ }
+}
+
+# Best effort: a controller that went away since it was listed needs no scan,
+# and one that cannot be rescanned is a warning.
+my sub rescan_namespaces($nqn) {
+ for my $controller (nvme_controllers($nqn)->@*) {
+ next if $controller->{state} ne 'live';
+ my $path = "/sys/class/nvme/$controller->{name}/rescan_controller";
+ my $error = write_attribute($path, "1\n", 1) // next;
+ log_warn("cannot rescan NVMe controller '$controller->{name}': $error");
+ }
+}
+
+# Deletes a controller that is dead or that moved to another interface. The
+# activation reconnects or keeps the path either way, so this only warns.
+my sub disconnect_controller($controller) {
+ my $name = $controller->{name};
+ eval {
+ run_bounded(
+ 5,
+ "deleting NVMe controller '$name'",
+ sub { return _delete_controllers([$controller]) },
+ );
+ };
+ if (my $error = $@) {
+ rethrow_task_interrupt($error);
+ log_warn($error);
+ }
+}
+
+# Waits up to 10 seconds for a live controller of the portal on its
+# configured interface.
+my sub wait_for_live_path($nqn, $portal) {
+ for (my $attempt = 0; $attempt < 40; $attempt++) {
+ return 1 if _live_portal_count(controller_states($nqn), [$portal]);
+ _sleep(0.25);
+ }
+ return 0;
+}
+
+# Connects every missing or dead path on its configured interface. A path
+# that is connected on another interface moves make-before-break: its old
+# controller is only disconnected once the new one is live, so changing an
+# interface mapping never takes away a path that carries I/O.
+my sub connect_portals($scfg, $portals, $hostnqn, $hostid, $key) {
+ my $nqn = $scfg->{subsysnqn};
+ for my $portal ($portals->@*) {
+ my $id = "$portal->{address}:$portal->{port}";
+ my $iface = $portal->{host_iface};
+ my @controllers = (controller_states($nqn)->{$id} // [])->@*;
+ my ($configured) = grep { ($_->{host_iface} // '') eq $iface } @controllers;
+ my @moved = grep { ($_->{host_iface} // '') ne $iface } @controllers;
+
+ if (!$configured || $configured->{state} eq 'dead') {
+ disconnect_controller($configured) if $configured;
+ if (!portal_reachable($portal)) {
+ log_warn("NVMe/TCP portal '$id' is unreachable");
+ next;
+ }
+ eval { _connect_portal($scfg, $portal, $hostnqn, $hostid, $key) };
+ my $connect_error = $@;
+ # A failed or timed-out connect can still have created a controller.
+ eval { restrict_secret_attrs($nqn); 1 };
+ if (my $error = $@) {
+ rethrow_task_interrupt($error);
+ log_warn("cannot restrict NVMe controller attributes for '$nqn'");
+ }
+ rethrow_task_interrupt($connect_error) if $connect_error;
+ if ($connect_error) {
+ log_warn("$id: $connect_error");
+ next;
+ }
+ }
+ next if !@moved;
+ if (!wait_for_live_path($nqn, $portal)) {
+ log_warn("NVMe/TCP portal '$id' is not live on '$iface' yet; keeping its path on"
+ . " another interface");
+ next;
+ }
+ disconnect_controller($_) for @moved;
+ }
+}
+
+sub activate_storage($class, $storeid, $scfg, $cache = undef) {
+ $cache //= {};
+ my $nqn = $scfg->{subsysnqn};
+ # First, whatever the rest of the activation does, also for controllers
+ # connected by hand. Also try after each connect attempt below, including
+ # failures that may have created a controller.
+ restrict_secret_attrs($nqn);
+
+ my $multipath = file_read_firstline('/sys/module/nvme_core/parameters/multipath')
+ // die "the NVMe kernel modules are not loaded; load the nvme-tcp kernel module\n";
+ die "native NVMe multipath is disabled in the running kernel\n" if $multipath ne 'Y';
+
+ my $portals = _configured_portals($scfg);
+ verify_nvme_nqn($nqn);
+ _validate_local_ifaces($portals);
+ _assert_unique_target($storeid, $scfg);
+ my $hostnqn = file_read_firstline('/etc/nvme/hostnqn')
+ // die "missing /etc/nvme/hostnqn; the nvme-cli package generates it\n";
+ verify_nvme_nqn($hostnqn);
+ my $hostnqns = parse_nvme_host_nqns($scfg->{'nvme-host-nqns'});
+ die "local NVMe host NQN '$hostnqn' is missing from nvme-host-nqns\n"
+ if !grep { $_ eq $hostnqn } $hostnqns->@*;
+ my $policy = $scfg->{'nvme-iopolicy'} // 'round-robin';
+ my $force_reconcile = delete($cache->{'zfsnvme-force-reconcile'}->{$storeid}) // 0;
+ my $states = controller_states($nqn);
+ my $healthy = _live_portal_count($states, $portals);
+ my $usable = _live_portal_count($states, $portals, 1);
+ my $last = $slow_path_last{$storeid};
+ my $recent = defined($last) && _now() - $last < $slow_path_backoff;
+
+ # This method is called by the periodic storage status loop. Once every
+ # configured path is live, lifecycle operations keep the target converged,
+ # so no remote call is needed. A degraded storage retries the target at
+ # most once a minute: in between, one with a live path (also one waiting
+ # to move to its configured interface) is usable, and one without fails
+ # at once instead of waiting for a path again. Missing namespaces
+ # explicitly force the slow path from activate_volume().
+ if (!$force_reconcile && ($healthy == scalar($portals->@*) || $recent)) {
+ die "no live NVMe/TCP path for storage '$storeid'\n" if !$usable;
+ die "missing NVMe DH-HMAC-CHAP key\n"
+ if (file_read_firstline(secret_path($storeid)) // '') eq '';
+ set_iopolicy($nqn, $policy);
+ set_tunables($scfg);
+ return 1;
+ }
+
+ $slow_path_last{$storeid} = _now();
+ my $hostid = file_read_firstline('/etc/nvme/hostid')
+ // die "missing /etc/nvme/hostid; the nvme-cli package generates it\n";
+ die "invalid NVMe host ID\n"
+ if $hostid !~ $RE_NVMET_UUID || $hostid =~ $RE_ZERO_UUID;
+ my $key = get_secret($storeid);
+ for my $name (qw(keep-alive-tmo reconnect-delay ctrl-loss-tmo fast-io-fail-tmo nr-io-queues)) {
+ my $value = $scfg->{"nvme-$name"};
+ die "invalid NVMe connection parameter '$name'\n"
+ if defined($value) && $value !~ $RE_CONFIG_INT;
+ }
+ _nvmet_activate_target($storeid, $scfg, $portals, $hostnqns, $key);
+ connect_portals($scfg, $portals, $hostnqn, $hostid, $key);
+
+ # Existing controllers can be in the middle of their kernel reconnect delay
+ # after a target restart. Do not create duplicates, but give that recovery
+ # cycle enough time to complete before declaring the storage unavailable.
+ for (my $attempt = 0; $attempt < 60; $attempt++) {
+ $states = controller_states($nqn);
+ last if _live_portal_count($states, $portals, 1);
+ _sleep(0.25);
+ }
+ die "no live NVMe/TCP path for storage '$storeid'\n"
+ if !_live_portal_count($states, $portals, 1);
+ $healthy = _live_portal_count($states, $portals);
+ log_warn("storage '$storeid' is degraded: $healthy/" . scalar($portals->@*) . " paths live")
+ if $healthy < scalar($portals->@*);
+
+ set_iopolicy($nqn, $policy);
+ set_tunables($scfg);
+ return 1;
+}
+
+sub deactivate_storage($class, $storeid, $scfg, $cache = undef) {
+ my $nqn = $scfg->{subsysnqn};
+ my $openers = _namespace_openers($nqn);
+ die "refusing to disconnect NVMe storage '$storeid': namespace in use by "
+ . join(', ', $openers->@*) . "\n"
+ if $openers->@*;
+
+ # One child deletes every controller, bounded by 15 seconds in all. When it
+ # gives up, the controllers that are left stay connected and the
+ # deactivation fails.
+ my $controllers = nvme_controllers($nqn);
+ run_bounded(
+ 15,
+ "disconnecting NVMe subsystem '$nqn'",
+ sub { return _delete_controllers($controllers) },
+ ) if $controllers->@*;
+ return 1;
+}
+
+# Removing a storage definition never destroys data and never depends on the
+# target, like for every other storage type. Target cleanup is best effort
+# and only happens once the storage owns no volume. Other nodes keep their
+# connections until they are disconnected there.
+sub on_delete_hook($class, $storeid, $scfg) {
+ # First: an activation elsewhere that waits for the target lock then
+ # refuses instead of restoring what is removed below.
+ delete_secret($storeid);
+ eval { $class->deactivate_storage($storeid, $scfg) };
+ if (my $err = $@) {
+ rethrow_task_interrupt($err);
+ chomp($err);
+ log_warn("not disconnecting NVMe storage '$storeid': $err");
+ }
+ eval { _nvmet_delete_target($scfg) };
+ if (my $err = $@) {
+ rethrow_task_interrupt($err);
+ chomp($err);
+ log_warn("keeping the NVMe target configuration of storage '$storeid': $err");
+ }
+ return;
+}
+
+sub path($class, $scfg, $volname, $storeid, $snapname = undef) {
+ die "direct access to snapshots not implemented\n" if defined($snapname);
+ my ($vtype, $name, $vmid) = $class->parse_volname($volname);
+ my $uuid = _nvmet_volume_uuid($scfg, $name);
+ my $path = "/dev/disk/by-id/nvme-uuid.$uuid";
+ return ($path, $vmid, $vtype);
+}
+
+sub qemu_blockdev_options(
+ $class, $scfg, $storeid, $volname,
+ $machine_version = undef,
+ $options = undef,
+) {
+ die "direct access to snapshots not implemented\n" if $options->{'snapshot-name'};
+ my ($path) = $class->path($scfg, $volname, $storeid);
+ return { driver => 'host_device', filename => $path };
+}
+
+# Whether the local device behind a namespace link is the namespace (NSID,
+# UUID) of this subsystem, so a UUID claimed by another subsystem is never
+# handed to a guest.
+sub _nvmet_local_namespace_ok($path, $nqn, $nsid, $uuid) {
+ my $link = _link_target($path) // return 0;
+ my $device = (split m{/}, $link)[-1];
+ return 0 if !defined($device) || $device !~ $RE_NVME_NAMESPACE;
+ my $sys = "/sys/block/$device";
+ return 0 if lc(file_read_firstline("$sys/uuid") // '') ne lc($uuid);
+ return 0 if (file_read_firstline("$sys/nsid") // '') ne "$nsid";
+ return 0 if (file_read_firstline("$sys/device/subsysnqn") // '') ne $nqn;
+ return 1;
+}
+
+sub activate_volume(
+ $class, $storeid, $scfg, $volname,
+ $snapname = undef,
+ $cache = undef,
+ $hints = undef,
+) {
+ die "unable to activate snapshot from remote zfs storage\n" if $snapname;
+ my (undef, $name) = $class->parse_volname($volname);
+ my $nqn = nvmet_nqn($scfg);
+ my $dataset = nvmet_dataset($scfg, $name);
+
+ my ($inv, $cfs) = nvmet_read($scfg);
+ my ($nsid, $uuid) = _nvmet_owned_identity($inv, $nqn, $dataset);
+ my $state = _nvmet_export_state($cfs, $nqn, $nsid, $uuid, "/dev/zvol/$dataset");
+ my $path = "/dev/disk/by-id/nvme-uuid.$uuid";
+ my $ok = sub {
+ return _block_device($path) && _nvmet_local_namespace_ok($path, $nqn, $nsid, $uuid);
+ };
+ return 1 if $state eq 'present' && $ok->();
+
+ $cache //= {};
+ $cache->{'zfsnvme-force-reconcile'}->{$storeid} = 1;
+ $class->activate_storage($storeid, $scfg, $cache);
+ for (my $attempt = 0; $attempt < 40 && !_block_device($path); $attempt++) {
+ # a host can miss the change notice; rescanning is harmless
+ rescan_namespaces($nqn) if $attempt % 8 == 0;
+ _sleep(0.25);
+ }
+ die "NVMe namespace for '$volname' did not appear\n" if !_block_device($path);
+ die "NVMe namespace for '$volname' has an unexpected identity\n" if !$ok->();
+ return 1;
+}
+
+sub deactivate_volume(
+ $class, $storeid, $scfg, $volname, $snapname = undef, $cache = undef,
+) {
+ die "unable to deactivate snapshot from remote zfs storage\n" if $snapname;
+ return 1;
+}
+
+1;
^ permalink raw reply related [flat|nested] 7+ messages in thread
* [PATCH storage v3 2/4] test: add zfsnvme plugin tests
2026-10-05 0:26 [PATCH storage v3 0/4] add ZFS over NVMe/TCP storage plugin Joaquin Varela
2026-10-05 0:26 ` [PATCH storage v3 1/4] zfsnvme: " Joaquin Varela
@ 2026-10-05 0:26 ` Joaquin Varela
2026-10-05 0:26 ` [PATCH storage v3 3/4] zfsnvme: fence target commands of abandoned transactions Joaquin Varela
` (3 subsequent siblings)
5 siblings, 0 replies; 7+ messages in thread
From: Joaquin Varela @ 2026-10-05 0:26 UTC (permalink / raw)
To: pve-devel
Add two test files to the plugin test harness and register them in
run_plugin_tests.pl.
zfsnvme_test.pm covers the host side: option and portal parsing, the
schema and the config checks, the key handling (the add and update
hooks with /etc/pve/priv/storage, and the 'dhchap-key' file mapping of
'pvesm add' and 'pvesm set'), the SSH runner and its error
classification, the /dev/nvme-fabrics option builder and the connect
child (against a pseudoterminal that answers like the fabrics device;
skipped without /dev/ptmx), the sysfs-based rescan and controller
deletion, path selection and activation, vdisk_alloc under the core
storage lock, and the volume path and QEMU blockdev options. Nothing in
it opens a connection or touches a local NVMe device.
zfsnvme_target_test.pm covers the target side: the parsers and planners
on fixture reads of an nvmet target (Linux 6.8, ZFS 2.2; addresses and
NQNs replaced by documentation values), the command renderer and
chunking, and every target flow against an in-memory target
(FakeTarget) that interprets exactly the commands the plugin sends and
records anything else as a protocol violation. A fault sweep fails
every mutating call of every flow in each way an SSH call can fail
(before or after it ran, a connection cut or lost in the middle of a
chain, or a call the target runs to its end after ssh gave up on it)
and checks that identities stay unique, nothing of other storages
changes, every change ran under one domain lock, and the next
activation converges. The rendered chains also run with /bin/sh on a
configfs stand-in in a temporary directory, which checks the quoting
and the POSIX command shapes; these tests bail out unless /bin/sh,
find, grep, sha256sum, dd and tee are available.
Signed-off-by: Joaquin Varela <joaquinvarela@neatech.ar>
---
src/test/run_plugin_tests.pl | 2 +
.../zfsnvme_fixtures/configfs_snapshot.txt | 38 +
src/test/zfsnvme_fixtures/zfs_inventory.txt | 15 +
src/test/zfsnvme_target_test.pm | 4849 +++++++++++++++++
src/test/zfsnvme_test.pm | 2353 ++++++++
5 files changed, 7257 insertions(+)
create mode 100644 src/test/zfsnvme_fixtures/configfs_snapshot.txt
create mode 100644 src/test/zfsnvme_fixtures/zfs_inventory.txt
create mode 100644 src/test/zfsnvme_target_test.pm
create mode 100644 src/test/zfsnvme_test.pm
diff --git a/src/test/run_plugin_tests.pl b/src/test/run_plugin_tests.pl
index 8bce9d3b..d063a3d2 100755
--- a/src/test/run_plugin_tests.pl
+++ b/src/test/run_plugin_tests.pl
@@ -17,6 +17,8 @@ my $res = $harness->runtests(
"get_subdir_test.pm",
"filesystem_path_test.pm",
"prune_backups_test.pm",
+ "zfsnvme_test.pm",
+ "zfsnvme_target_test.pm",
);
exit -1 if !$res || $res->{failed} || $res->{parse_errors};
diff --git a/src/test/zfsnvme_fixtures/configfs_snapshot.txt b/src/test/zfsnvme_fixtures/configfs_snapshot.txt
new file mode 100644
index 00000000..1b7672a0
--- /dev/null
+++ b/src/test/zfsnvme_fixtures/configfs_snapshot.txt
@@ -0,0 +1,38 @@
+D /sys/kernel/config/nvmet
+D /sys/kernel/config/nvmet/hosts
+D /sys/kernel/config/nvmet/hosts/nqn.2014-08.org.nvmexpress:uuid:0a0a0a0a-0a0a-4a0a-8a0a-0a0a0a0a0a0a
+D /sys/kernel/config/nvmet/hosts/nqn.2014-08.org.nvmexpress:uuid:0b0b0b0b-0b0b-4b0b-8b0b-0b0b0b0b0b0b
+D /sys/kernel/config/nvmet/ports
+D /sys/kernel/config/nvmet/ports/2
+D /sys/kernel/config/nvmet/ports/2/subsystems
+D /sys/kernel/config/nvmet/ports/1
+D /sys/kernel/config/nvmet/ports/1/subsystems
+D /sys/kernel/config/nvmet/subsystems
+D /sys/kernel/config/nvmet/subsystems/nqn.2026-01.com.example:zfsnvme
+D /sys/kernel/config/nvmet/subsystems/nqn.2026-01.com.example:zfsnvme/allowed_hosts
+D /sys/kernel/config/nvmet/subsystems/nqn.2026-01.com.example:zfsnvme/namespaces
+D /sys/kernel/config/nvmet/subsystems/nqn.2026-01.com.example:zfsnvme/namespaces/6
+D /sys/kernel/config/nvmet/subsystems/nqn.2026-01.com.example:zfsnvme/namespaces/5
+L /sys/kernel/config/nvmet/ports/2/subsystems/nqn.2026-01.com.example:zfsnvme
+L /sys/kernel/config/nvmet/ports/1/subsystems/nqn.2026-01.com.example:zfsnvme
+L /sys/kernel/config/nvmet/subsystems/nqn.2026-01.com.example:zfsnvme/allowed_hosts/nqn.2014-08.org.nvmexpress:uuid:0a0a0a0a-0a0a-4a0a-8a0a-0a0a0a0a0a0a
+L /sys/kernel/config/nvmet/subsystems/nqn.2026-01.com.example:zfsnvme/allowed_hosts/nqn.2014-08.org.nvmexpress:uuid:0b0b0b0b-0b0b-4b0b-8b0b-0b0b0b0b0b0b
+/sys/kernel/config/nvmet/ports/2/addr_trtype:tcp
+/sys/kernel/config/nvmet/ports/2/addr_trsvcid:4420
+/sys/kernel/config/nvmet/ports/2/addr_traddr:192.0.2.22
+/sys/kernel/config/nvmet/ports/2/addr_adrfam:ipv4
+/sys/kernel/config/nvmet/ports/1/addr_trtype:tcp
+/sys/kernel/config/nvmet/ports/1/addr_trsvcid:4420
+/sys/kernel/config/nvmet/ports/1/addr_traddr:192.0.2.21
+/sys/kernel/config/nvmet/ports/1/addr_adrfam:ipv4
+/sys/kernel/config/nvmet/subsystems/nqn.2026-01.com.example:zfsnvme/namespaces/6/buffered_io:0
+/sys/kernel/config/nvmet/subsystems/nqn.2026-01.com.example:zfsnvme/namespaces/6/enable:1
+/sys/kernel/config/nvmet/subsystems/nqn.2026-01.com.example:zfsnvme/namespaces/6/device_uuid:0d0d0d0d-0d0d-4d0d-8d0d-0d0d0d0d0d0d
+/sys/kernel/config/nvmet/subsystems/nqn.2026-01.com.example:zfsnvme/namespaces/6/device_path:/dev/zvol/tank/vm-100-cloudinit
+/sys/kernel/config/nvmet/subsystems/nqn.2026-01.com.example:zfsnvme/namespaces/5/buffered_io:0
+/sys/kernel/config/nvmet/subsystems/nqn.2026-01.com.example:zfsnvme/namespaces/5/enable:1
+/sys/kernel/config/nvmet/subsystems/nqn.2026-01.com.example:zfsnvme/namespaces/5/device_uuid:0c0c0c0c-0c0c-4c0c-8c0c-0c0c0c0c0c0c
+/sys/kernel/config/nvmet/subsystems/nqn.2026-01.com.example:zfsnvme/namespaces/5/device_path:/dev/zvol/tank/vm-100-disk-0
+/sys/kernel/config/nvmet/subsystems/nqn.2026-01.com.example:zfsnvme/attr_model:Proxmox ZFS NVMe
+/sys/kernel/config/nvmet/subsystems/nqn.2026-01.com.example:zfsnvme/attr_serial:PVEZFS51b2c3c90ea254
+/sys/kernel/config/nvmet/subsystems/nqn.2026-01.com.example:zfsnvme/attr_allow_any_host:0
diff --git a/src/test/zfsnvme_fixtures/zfs_inventory.txt b/src/test/zfsnvme_fixtures/zfs_inventory.txt
new file mode 100644
index 00000000..69ba9238
--- /dev/null
+++ b/src/test/zfsnvme_fixtures/zfs_inventory.txt
@@ -0,0 +1,15 @@
+tank type filesystem -
+tank proxmox:nvme-subsys - -
+tank proxmox:nvme-nsid - -
+tank proxmox:nvme-uuid - -
+tank proxmox:nvme-last-nsid - -
+tank/vm-100-cloudinit type volume -
+tank/vm-100-cloudinit proxmox:nvme-subsys nqn.2026-01.com.example:zfsnvme local
+tank/vm-100-cloudinit proxmox:nvme-nsid 6 local
+tank/vm-100-cloudinit proxmox:nvme-uuid 0d0d0d0d-0d0d-4d0d-8d0d-0d0d0d0d0d0d local
+tank/vm-100-cloudinit proxmox:nvme-last-nsid - -
+tank/vm-100-disk-0 type volume -
+tank/vm-100-disk-0 proxmox:nvme-subsys nqn.2026-01.com.example:zfsnvme local
+tank/vm-100-disk-0 proxmox:nvme-nsid 5 local
+tank/vm-100-disk-0 proxmox:nvme-uuid 0c0c0c0c-0c0c-4c0c-8c0c-0c0c0c0c0c0c local
+tank/vm-100-disk-0 proxmox:nvme-last-nsid - -
diff --git a/src/test/zfsnvme_target_test.pm b/src/test/zfsnvme_target_test.pm
new file mode 100644
index 00000000..991d2084
--- /dev/null
+++ b/src/test/zfsnvme_target_test.pm
@@ -0,0 +1,4849 @@
+# Target side of the zfsnvme storage plugin: the parsers and planners on the
+# fixture reads, the command renderer, and every target flow against an
+# in-memory NVMe target (FakeTarget) that interprets exactly the commands the
+# plugin sends. Nothing here opens a connection or touches local NVMe devices.
+
+use v5.36;
+
+use lib qw(..);
+
+use Compress::Zlib qw(crc32);
+use Digest::SHA qw(sha256_hex);
+use File::Temp qw(tempdir);
+use FindBin;
+use IPC::Open3;
+use MIME::Base64 qw(encode_base64);
+use POSIX ();
+use Storable qw(dclone freeze);
+use Symbol qw(gensym);
+use Test::MockModule;
+use Test::More;
+
+use PVE::Cluster;
+use PVE::Storage;
+use PVE::Storage::ZFSNVMePlugin;
+
+my $PLUGIN = 'PVE::Storage::ZFSNVMePlugin';
+my $ROOT = '/sys/kernel/config/nvmet';
+my $MARKER = 'ZFSNVME-CONFIGFS';
+my $MODEL = 'Proxmox ZFS NVMe';
+my $FIXTURES = "$FindBin::Bin/zfsnvme_fixtures";
+my $UUID_RE = qr/\A[0-9a-f]{8}(?:-[0-9a-f]{4}){3}-[0-9a-f]{12}\z/i;
+my $EMPTY_KEY_SHA = sha256_hex("\n");
+
+# The fixtures are the ZFS inventory and the configfs read of a target (Linux
+# 6.8 nvmet, ZFS 2.2) that serves one subsystem with two namespaces to two
+# cluster nodes, as the plugin reads them.
+my $FIXTURE_NQN = 'nqn.2026-01.com.example:zfsnvme';
+my @FIXTURE_HOSTS = (
+ 'nqn.2014-08.org.nvmexpress:uuid:0a0a0a0a-0a0a-4a0a-8a0a-0a0a0a0a0a0a',
+ 'nqn.2014-08.org.nvmexpress:uuid:0b0b0b0b-0b0b-4b0b-8b0b-0b0b0b0b0b0b',
+);
+my $FIXTURE_UUID5 = '0c0c0c0c-0c0c-4c0c-8c0c-0c0c0c0c0c0c';
+my $FIXTURE_UUID6 = '0d0d0d0d-0d0d-4d0d-8d0d-0d0d0d0d0d0d';
+
+my $NQN = 'nqn.2026-01.com.example:zfsnvme-test';
+my $FOREIGN_NQN = 'nqn.2026-01.com.example:foreign';
+my @HOSTS =
+ map { sprintf('nqn.2014-08.org.nvmexpress:uuid:00000000-0000-4000-8000-%012d', $_) } 1, 2;
+
+# A portal as parse_nvme_portals() returns it.
+sub portal($family, $address, $port) {
+ return { family => $family, address => $address, port => $port };
+}
+my $PORTALS = [portal('ipv4', '192.0.2.21', 4420), portal('ipv4', '192.0.2.22', 4420)];
+my %U = map {
+ my $d = $_;
+ ($d => join('-', $d x 8, $d x 4, '4' . $d x 3, '8' . $d x 3, $d x 12))
+} 1 .. 9;
+
+sub make_key($hash, $length, $seed = 1) {
+ my $raw = join('', map { chr(($_ * 131 + $seed * 17) % 256) } 1 .. $length);
+ return "DHHC-1:$hash:" . encode_base64($raw . pack('V', crc32($raw)), '') . ':';
+}
+my $KEY = make_key('00', 32);
+my $KEY_B = make_key('00', 32, 2);
+my $KEY_SHA = sha256_hex("$KEY\n");
+
+sub scfg(%override) {
+ return {
+ type => 'zfsnvme',
+ server => '192.0.2.10',
+ pool => 'tank',
+ subsysnqn => $NQN,
+ blocksize => '16k',
+ sparse => 1,
+ 'nvme-portals' => '192.0.2.21,192.0.2.22',
+ 'nvme-host-ifaces' => 'ens19,ens20',
+ 'nvme-host-nqns' => join(',', @HOSTS),
+ %override,
+ };
+}
+
+sub slurp($path) {
+ open(my $fh, '<', $path) or die "cannot open '$path': $!\n";
+ local $/;
+ return scalar(<$fh>);
+}
+
+# Calls a plugin package sub by name, so a mocked sub is the one called.
+sub nv($name, @args) {
+ my $code = $PLUGIN->can($name) // die "plugin has no sub '$name'\n";
+ return $code->(@args);
+}
+
+sub same($x, $y) {
+ local $Storable::canonical = 1;
+ return freeze([$x]) eq freeze([$y]);
+}
+
+# ---------------------------------------------------------------------------
+# Mocks: the only way to the target is _nvmet_run, and it reaches $FAKE.
+# The pmxcfs lock, quorum, clock, waits, UUID source, local files and local
+# devices are mocked. Waits advance $NOW.
+# ---------------------------------------------------------------------------
+
+our $FAKE;
+our $NOW = 1_000_000;
+our $SLEPT = 0;
+our $QUORATE = 1;
+our $LOCK_HOOK;
+our (%LOCK_HELD, @LOCKS, @NESTED_LOCKS, @QUORUM, @WARNINGS, @SYSFS_WRITES, %FILES, %CONFIG);
+our (%CORPUS, @VIOLATIONS, @WRITES, @UNLINKED, %BLOCK, @UUIDS);
+
+my $uuid_seq = 0;
+
+sub lock_id($scfg) {
+ return 'zfsnvme-' . ($scfg->{server} =~ s/[^A-Za-z0-9.-]/_/gr);
+}
+
+sub lock_held($scfg) {
+ return (($LOCK_HELD{ lock_id($scfg) } // 0) == $$) ? 1 : 0;
+}
+
+# Whether a step changes the target, by the vocabulary of the fake target,
+# independently of the plugin.
+sub mutating_step($step) {
+ return 1 if ref($step) ne 'ARRAY';
+ my ($command, $subcommand) = $step->@*;
+ return 0 if grep { $command eq $_ } qw(test grep cat printf env);
+ return 0 if $command eq 'zfs' && ($subcommand eq 'get' || $subcommand eq 'list');
+ return 1;
+}
+
+# Whether steps change the target, not counting the module load of a locked read.
+sub changing($steps) {
+ return scalar(grep { mutating_step($_) && !step_is($_, 'modprobe') } $steps->@*);
+}
+
+my $plugin_mock = Test::MockModule->new($PLUGIN);
+my $cluster_mock = Test::MockModule->new('PVE::Cluster');
+my $storage_mock = Test::MockModule->new('PVE::Storage');
+my $sysfs_mock = Test::MockModule->new('PVE::SysFSTools');
+my $tools_mock = Test::MockModule->new('PVE::Tools');
+
+$plugin_mock->redefine(
+ _nvmet_run => sub($scfg, $steps, %opts) {
+ die "test error: no fake target\n" if !$FAKE;
+ my $rendered = nv('_nvmet_render', $steps);
+ $CORPUS{$rendered} //= $opts{op} // '';
+ return $FAKE->run($scfg, $steps, $rendered, %opts);
+ },
+);
+# The target is the only place that runs commands, and only through _nvmet_run.
+# A command is also a violation, in case the caller handles the error.
+my $no_command = sub($cmd, %opts) {
+ my $what = 'unexpected command ' . (ref($cmd) ? $cmd->[0] : $cmd);
+ push @VIOLATIONS, $what;
+ die "test error: $what\n";
+};
+$plugin_mock->redefine(run_command => $no_command);
+$tools_mock->redefine(run_command => $no_command);
+$plugin_mock->redefine(_now => sub () { return $NOW });
+$plugin_mock->redefine(
+ _sleep => sub($seconds) {
+ $NOW += $seconds;
+ $SLEPT += $seconds;
+ return;
+ },
+);
+$plugin_mock->redefine(_block_device => sub($path) { return $BLOCK{$path} });
+$plugin_mock->redefine(
+ _unlink_file => sub($path) {
+ push @UNLINKED, $path;
+ delete $FILES{$path};
+ return 1;
+ },
+);
+$plugin_mock->redefine(log_warn => sub($message) { push @WARNINGS, $message; return });
+my $sysfs_write = sub($path, $data, $allow_existing = undef) {
+ push @SYSFS_WRITES, [$path, $data];
+ return 1;
+};
+$sysfs_mock->redefine(file_write => $sysfs_write);
+# The host side runs its kernel-facing steps in a bounded child; here they
+# run inline.
+$tools_mock->redefine(
+ run_fork_with_timeout => sub($timeout, $code, $opts = undef) {
+ my $res = $code->();
+ return wantarray ? ($res, 0) : $res;
+ },
+);
+# UUIDs come from @UUIDS first, then from a sequence.
+$plugin_mock->redefine(
+ file_read_firstline => sub($path) {
+ if ($path eq '/proc/sys/kernel/random/uuid') {
+ return shift(@UUIDS) if @UUIDS;
+ $uuid_seq++;
+ return sprintf('%08x-7e57-4000-8000-%012x', $uuid_seq, $uuid_seq);
+ }
+ return $FILES{$path};
+ },
+);
+$plugin_mock->redefine(
+ file_set_contents => sub($path, $data, $perm = undef, @rest) {
+ push @WRITES, { path => $path, data => $data, perm => $perm };
+ return;
+ },
+);
+$plugin_mock->redefine(make_path => sub(@args) { push @WRITES, { make_path => [@args] }; return });
+$plugin_mock->redefine(_namespace_openers => sub($nqn) { return [] });
+
+$cluster_mock->redefine(
+ check_cfs_quorum => sub($noerr = undef) {
+ push @QUORUM, $noerr;
+ die "cluster not ready - no quorum?\n" if !$QUORATE && !$noerr;
+ return $QUORATE;
+ },
+);
+# Keeps the cfs_lock contract: the code's result with $@ cleared, or undef
+# with $@ set. pmxcfs locks are not re-entrant: a second request for a held
+# lock times out, like a mkdir on /etc/pve/priv/lock would.
+$cluster_mock->redefine(
+ cfs_lock_domain => sub($name, $timeout, $code, @param) {
+ push @LOCKS, { name => $name, timeout => $timeout, pid => $$, now => $NOW };
+ if ($LOCK_HELD{$name}) {
+ push @NESTED_LOCKS, $name;
+ $@ = "cfs-lock 'domain-$name' error: got lock request timeout\n";
+ return undef;
+ }
+ if (my $hook = $LOCK_HOOK) {
+ local $LOCK_HOOK;
+ $hook->($name);
+ }
+ local $LOCK_HELD{$name} = $$;
+ my $res = eval { $code->(@param) };
+ my $err = $@;
+ if ($err) {
+ $@ = $err;
+ return undef;
+ }
+ $@ = undef;
+ return $res;
+ },
+);
+$storage_mock->redefine(config => sub () { return { ids => {%CONFIG} } });
+
+# ---------------------------------------------------------------------------
+# FakeTarget: ZFS datasets and an nvmet configfs tree with the kernel rules
+# the plugin depends on. It accepts only the command vocabulary of the
+# plugin; anything else is recorded as a protocol violation.
+# ---------------------------------------------------------------------------
+
+package FakeTarget {
+ use Digest::SHA qw(sha256_hex);
+ use Storable qw(dclone);
+
+ my $RR = quotemeta($ROOT);
+
+ # fail => [$match, $error, $times]: the steps that match (an argv prefix,
+ # { write => $suffix, value => $value } or a sub of the step) fail $times
+ # times with $error.
+ sub new($class, %opts) {
+ my $pools = delete($opts{pools}) // ['tank'];
+ if (my $fail = delete($opts{fail})) {
+ my ($match, $error, $times) = $fail->@*;
+ my $matches =
+ ref($match) eq 'CODE' ? $match
+ : ref($match) eq 'HASH'
+ ? sub($step) { main::write_to($step, $match->{write}, $match->{value}) }
+ : sub($step) { main::step_is($step, $match->@*) };
+ my $count = 0;
+ $opts{step_fault} = sub($step, $call) {
+ return undef if !$matches->($step) || $count >= ($times // 99);
+ $count++;
+ return $error;
+ };
+ }
+ my $self = bless {
+ m => {
+ ds => {},
+ snaps => {},
+ seq => 0,
+ subsystems => {},
+ ports => {},
+ hosts => {},
+ mounted => 1,
+ loaded => 1,
+ available => '1073741824',
+ used => '1048576',
+ },
+ calls => [],
+ violations => [],
+ faults => {}, # call index => before | after | late | cut:<step> | lost:<step>
+ late => [], # calls that ssh gave up on, applied by land()
+ # { after => call index, mode => before | unreachable | failed }
+ read_fault => undef,
+ step_fault => undef, # sub ($step, $call) returning an error text
+ before_call => undef,
+ after_call => undef,
+ udev_delay => 0, # failing `test -b` polls after a zvol (re)appears
+ unbindable => {}, # port id => 1: its address is missing on the target
+ unreachable => 0,
+ secrets => [$KEY, $KEY_B],
+ %opts,
+ }, $class;
+ $self->{m}->{ds}->{$_} = { type => 'filesystem', props => {} } for $pools->@*;
+ return $self;
+ }
+
+ sub from_state($class, $state, %opts) {
+ my $self = $class->new(%opts);
+ $self->{m} = dclone($state);
+ return $self;
+ }
+
+ # --- building a state -------------------------------------------------
+
+ sub add_zvol($self, $name, %o) {
+ my %props;
+ @props{qw(proxmox:nvme-subsys proxmox:nvme-nsid proxmox:nvme-uuid)} = $o{identity}->@*
+ if $o{identity};
+ $self->{m}->{ds}->{$name} = {
+ type => 'volume',
+ props => \%props,
+ volsize => $o{volsize} // 1073741824,
+ origin => $o{origin} // '-',
+ devwait => 0,
+ };
+ return $self;
+ }
+
+ sub add_snapshot($self, $snapshot, %o) {
+ my ($name) = split /\@/, $snapshot;
+ my $ds = $self->{m}->{ds}->{$name};
+ $self->{m}->{snaps}->{$snapshot} = {
+ props => $o{props} // dclone($ds->{props}),
+ volsize => $ds->{volsize},
+ seq => ++$self->{m}->{seq},
+ };
+ return $self;
+ }
+
+ sub add_subsystem($self, $nqn, %o) {
+ $self->{m}->{subsystems}->{$nqn} = {
+ attr => {
+ attr_model => $o{model} // $MODEL,
+ attr_serial => $o{serial} // PVE::Storage::ZFSNVMePlugin::_nvmet_serial($nqn),
+ attr_allow_any_host => $o{allow_any_host} // '0',
+ },
+ acl => { map { $_ => 1 } ($o{acl} // [])->@* },
+ ns => {},
+ };
+ return $self;
+ }
+
+ sub add_namespace($self, $nqn, $nsid, $uuid, $dev, $enable = 1) {
+ $self->{m}->{subsystems}->{$nqn}->{ns}->{$nsid} = {
+ enable => "$enable",
+ device_path => $dev,
+ device_uuid => $uuid,
+ buffered_io => '0',
+ };
+ return $self;
+ }
+
+ sub add_port($self, $id, $address, $service, %o) {
+ $self->{m}->{ports}->{$id} = {
+ attr => {
+ addr_trtype => 'tcp',
+ addr_adrfam => $o{family} // 'ipv4',
+ addr_traddr => $address,
+ addr_trsvcid => "$service",
+ },
+ links => { map { $_ => 1 } ($o{links} // [])->@* },
+ };
+ return $self;
+ }
+
+ # nvmet creates the key attributes of a host world-readable
+ sub host_entry($key = undef) {
+ return { key => $key, mode => { map { $_ => '0644' } qw(dhchap_key dhchap_ctrl_key) } };
+ }
+
+ sub add_host($self, $hostnqn, $key = undef) {
+ $self->{m}->{hosts}->{$hostnqn} = host_entry($key);
+ return $self;
+ }
+
+ # A target restart: configfs is empty and, unless $loaded, nvmet is not
+ # loaded yet.
+ sub reboot($self, $loaded = 0) {
+ my $m = $self->{m};
+ $m->{$_} = {} for qw(subsystems ports hosts);
+ $m->{loaded} = $loaded;
+ $_->{devwait} = 0 for grep { $_->{type} eq 'volume' } values $m->{ds}->%*;
+ return $self;
+ }
+
+ # --- execution --------------------------------------------------------
+
+ sub run($self, $scfg, $steps, $rendered, %opts) {
+ my $call = {
+ index => scalar($self->{calls}->@*),
+ op => $opts{op},
+ steps => dclone($steps),
+ rendered => $rendered,
+ timeout => $opts{timeout},
+ input => defined($opts{input}) ? 1 : 0,
+ locked => main::lock_held($scfg),
+ mutating => (grep { main::mutating_step($_) } $steps->@*) ? 1 : 0,
+ changing => main::changing($steps),
+ now => $main::NOW,
+ };
+ push $self->{calls}->@*, $call;
+ $self->audit($call, $opts{input});
+ $self->{before_call}->($self, $call) if $self->{before_call};
+ my $res = $self->execute($call, $steps, $opts{input});
+ $call->{rc} = $res->{rc};
+ $call->{err} = $res->{err};
+ $self->audit_text($call, $res->{err}, $res->{out}->@*);
+ $self->{after_call}->($self, $call) if $self->{after_call};
+ return $res;
+ }
+
+ # Violations are recorded for the fake and for the whole run, which the
+ # last subtest checks.
+ sub flag($self, @what) {
+ push $self->{violations}->@*, @what;
+ push @main::VIOLATIONS, @what;
+ }
+
+ sub violation($self, $what) {
+ $self->flag($what);
+ die "FAKEERR:protocol violation: $what\n";
+ }
+
+ sub err($message) {
+ die "FAKEERR:$message\n";
+ }
+
+ sub audit($self, $call, $input) {
+ my @bad;
+ my $op = $call->{op} // '';
+ push @bad, 'call without an operation label' if $op eq '';
+ for my $secret ($self->{secrets}->@*) {
+ push @bad, "key in the command of '$op'" if index($call->{rendered}, $secret) >= 0;
+ }
+ my $keys = grep { ref($_) eq 'HASH' && exists($_->{key}) } $call->{steps}->@*;
+ push @bad, "key write without stdin in '$op'" if $keys && !defined($input);
+ push @bad, "stdin without a key write in '$op'" if defined($input) && !$keys;
+ push @bad, "stdin is not a configured key in '$op'"
+ if defined($input) && !grep { $input eq "$_\n" } $self->{secrets}->@*;
+ push @bad, "target change without the target lock in '$op'"
+ if $call->{mutating} && !$call->{locked};
+ $self->flag(@bad);
+ }
+
+ sub audit_text($self, $call, @texts) {
+ for my $text (grep { defined } @texts) {
+ for my $secret ($self->{secrets}->@*) {
+ $self->flag("key in the output of '$call->{op}'") if index($text, $secret) >= 0;
+ }
+ }
+ }
+
+ # Applies the calls that ssh gave up on, as the target finally runs them:
+ # each chain stops at its first failing step.
+ sub land($self) {
+ for my $late (splice($self->{late}->@*)) {
+ for my $step ($late->{steps}->@*) {
+ last if !eval { $self->step($step, $late->{input}); 1 };
+ }
+ }
+ return $self;
+ }
+
+ # A fault mode fails a call before it runs (rc 1), after it ran (the
+ # connection breaks: 255), at one of its steps (cut: the connection breaks
+ # and the remote shell stops; lost: it breaks and the rest of the chain
+ # still runs, later), or ssh gives up on it while it runs to its end later
+ # (late: -1).
+ sub execute($self, $call, $steps, $input) {
+ my $closed = 'Connection to 192.0.2.10 closed by remote host.';
+ my $broken = { rc => 255, out => [], err => $closed };
+ my $no_route = 'ssh: connect to host 192.0.2.10: No route to host';
+ return { rc => 255, out => [], err => $no_route } if $self->{unreachable};
+ my $read_fault = $self->{read_fault};
+ if ($read_fault && !$call->{mutating} && $call->{index} > $read_fault->{after}) {
+ $self->{read_fault} = undef;
+ $call->{fault} = "read $read_fault->{mode}";
+ return { rc => 255, out => [], err => $no_route }
+ if $read_fault->{mode} eq 'unreachable';
+ return { rc => -1, out => [], err => 'ssh failed' } if $read_fault->{mode} eq 'failed';
+ return { rc => 1, out => [], err => 'injected read failure' };
+ }
+ my $mode = $self->{faults}->{ $call->{index} } // '';
+ $call->{fault} = $mode if $mode;
+ return { rc => 1, out => [], err => 'injected failure' } if $mode eq 'before';
+ if ($mode eq 'late') {
+ push $self->{late}->@*, { steps => dclone($steps), input => $input };
+ return { rc => -1, out => [], err => 'timeout' };
+ }
+ my ($cut) = $mode =~ /\Acut:([0-9]+)\z/;
+ my ($lost) = $mode =~ /\Alost:([0-9]+)\z/;
+ my @out;
+ for my $index (0 .. $steps->$#*) {
+ return $broken if defined($cut) && $index == $cut;
+ if (defined($lost) && $index == $lost) {
+ my @rest = $steps->@[$index .. $steps->$#*];
+ push $self->{late}->@*, { steps => dclone(\@rest), input => $input };
+ return $broken;
+ }
+ my $step = $steps->[$index];
+ my $error = $self->{step_fault} ? $self->{step_fault}->($step, $call) : undef;
+ if (!defined($error)) {
+ my @lines = eval { $self->step($step, $input) };
+ my $e = $@;
+ if (!$e) {
+ push @out, @lines;
+ next;
+ }
+ if ($e !~ s/\AFAKEERR://) {
+ $self->flag("fake target error: $e");
+ return { rc => 2, out => \@out, err => 'fake target error' };
+ }
+ chomp($error = $e);
+ }
+ return { rc => 1, out => \@out, err => $error eq '' ? 'exit code 1' : $error };
+ }
+ return $broken if $mode eq 'after';
+ return { rc => 0, out => \@out, err => '' };
+ }
+
+ sub step($self, $step, $input) {
+ if (ref($step) eq 'HASH' && exists($step->{write})) {
+ $self->write_attr($step->{write}, $step->{value});
+ return ();
+ }
+ if (ref($step) eq 'HASH') {
+ err('dd: error reading standard input') if !defined($input);
+ (my $key = $input) =~ s/\n\z//;
+ my @failed;
+ for my $path ($step->{key}->@*) {
+ eval { $self->write_key($path, $key) };
+ push @failed, $@ if $@;
+ }
+ die $failed[0] if @failed;
+ return ();
+ }
+ my ($command, @args) = $step->@*;
+ my $method = "cmd_$command";
+ $self->violation("command '$command'") if !$self->can($method);
+ return $self->$method(@args);
+ }
+
+ # --- configfs ---------------------------------------------------------
+
+ sub available($self) {
+ err("find: '$ROOT': No such file or directory")
+ if !$self->{m}->{mounted} || !$self->{m}->{loaded};
+ }
+
+ sub parts($self, $path) {
+ return () if $path !~ m{\A$RR/(.+)\z};
+ return split m{/}, $1;
+ }
+
+ sub zvol_name($self, $dev) {
+ return $dev =~ m{\A/dev/zvol/(.+)\z} ? $1 : undef;
+ }
+
+ sub block_device($self, $dev, $poll) {
+ my $name = $self->zvol_name($dev) // return 0;
+ my $ds = $self->{m}->{ds}->{$name} // return 0;
+ return 0 if $ds->{type} ne 'volume';
+ if ($ds->{devwait} > 0) {
+ $ds->{devwait}-- if $poll;
+ return 0;
+ }
+ return 1;
+ }
+
+ sub busy($self, $name) {
+ for my $subsys (values $self->{m}->{subsystems}->%*) {
+ return 1
+ if grep { $_->{enable} eq '1' && $_->{device_path} eq "/dev/zvol/$name" }
+ values $subsys->{ns}->%*;
+ }
+ return 0;
+ }
+
+ sub new_uuid($self) {
+ my $seq = ++$self->{m}->{seq};
+ return sprintf('%08x-dead-4000-8000-%012x', $seq, $seq);
+ }
+
+ sub write_attr($self, $path, $value) {
+ $self->available;
+ my @p = $self->parts($path) or $self->violation("write '$path'");
+ my $m = $self->{m};
+ if ($p[0] eq 'subsystems' && @p == 3) {
+ my $subsys = $m->{subsystems}->{ $p[1] } // err("$path: No such file or directory");
+ $self->violation("write '$path'")
+ if $p[2] !~ /\A(?:attr_model|attr_serial|attr_allow_any_host)\z/;
+ if ($p[2] eq 'attr_allow_any_host') {
+ err("$path: Invalid argument") if $value !~ /\A[01]\z/;
+ err("$path: Invalid argument") if $value eq '1' && $subsys->{acl}->%*;
+ }
+ $subsys->{attr}->{ $p[2] } = "$value";
+ } elsif ($p[0] eq 'subsystems' && @p == 5 && $p[2] eq 'namespaces') {
+ my $subsys = $m->{subsystems}->{ $p[1] } // err("$path: No such file or directory");
+ my $ns = $subsys->{ns}->{ $p[3] } // err("$path: No such file or directory");
+ my $attr = $p[4];
+ if ($attr eq 'enable') {
+ err("$path: Invalid argument") if $value !~ /\A[01]\z/;
+ err("$path: No such device")
+ if $value eq '1'
+ && $ns->{enable} ne '1'
+ && !$self->block_device($ns->{device_path}, 0);
+ $ns->{enable} = "$value";
+ } elsif ($attr =~ /\A(?:device_path|device_uuid|buffered_io)\z/) {
+ err("$path: Device or resource busy") if $ns->{enable} eq '1';
+ if ($attr eq 'device_uuid') {
+ err("$path: Invalid argument") if $value !~ $UUID_RE;
+ $value = lc($value);
+ }
+ $ns->{$attr} = "$value";
+ } elsif ($attr eq 'revalidate_size') {
+ err("$path: Invalid argument") if $ns->{enable} ne '1' || $value ne '1';
+ $self->{revalidated}->{"$p[1]/$p[3]"}++;
+ } else {
+ $self->violation("write '$path'");
+ }
+ } elsif ($p[0] eq 'ports' && @p == 3) {
+ my $port = $m->{ports}->{ $p[1] } // err("$path: No such file or directory");
+ $self->violation("write '$path'")
+ if $p[2] !~ /\Aaddr_(?:trtype|adrfam|traddr|trsvcid)\z/;
+ err("$path: Permission denied") if $port->{links}->%*;
+ $port->{attr}->{ $p[2] } = "$value";
+ } else {
+ $self->violation("write '$path'");
+ }
+ }
+
+ sub write_key($self, $path, $key) {
+ $self->available;
+ my @p = $self->parts($path);
+ $self->violation("key write to '$path'")
+ if @p != 3 || $p[0] ne 'hosts' || $p[2] ne 'dhchap_key';
+ my $host = $self->{m}->{hosts}->{ $p[1] } // err("tee: $path: No such file or directory");
+ err("tee: $path: Invalid argument") if $key !~ /\ADHHC-1:/;
+ $self->violation("key write to the world-readable '$path'")
+ if ($host->{mode}->{ $p[2] } // '') ne '0600';
+ $host->{key} = $key;
+ }
+
+ sub cmd_mkdir($self, @paths) {
+ $self->available;
+ my $m = $self->{m};
+ for my $path (@paths) {
+ my @p = $self->parts($path) or $self->violation("mkdir '$path'");
+ my $exists = "mkdir: cannot create directory '$path': File exists";
+ my $missing = "mkdir: cannot create directory '$path': No such file or directory";
+ if ($p[0] eq 'subsystems' && @p == 2) {
+ err($exists) if $m->{subsystems}->{ $p[1] };
+ $m->{subsystems}->{ $p[1] } = {
+ attr => {
+ attr_model => 'Linux',
+ attr_serial => substr(sha256_hex("serial $p[1]"), 0, 16),
+ attr_allow_any_host => '0',
+ },
+ acl => {},
+ ns => {},
+ };
+ } elsif ($p[0] eq 'subsystems' && @p == 4 && $p[2] eq 'namespaces') {
+ my $subsys = $m->{subsystems}->{ $p[1] } // err($missing);
+ err("mkdir: cannot create directory '$path': Invalid argument")
+ if $p[3] !~ /\A[0-9]+\z/ || $p[3] == 0 || $p[3] > 0xfffffffe;
+ err($exists) if $subsys->{ns}->{ $p[3] };
+ $subsys->{ns}->{ $p[3] } = {
+ enable => '0',
+ device_path => '(null)',
+ device_uuid => $self->new_uuid,
+ buffered_io => '0',
+ };
+ } elsif ($p[0] eq 'ports' && @p == 2 && $p[1] =~ /\A[0-9]+\z/) {
+ err($exists) if $m->{ports}->{ $p[1] };
+ $m->{ports}->{ $p[1] } = {
+ attr => { map { ("addr_$_" => '') } qw(trtype adrfam traddr trsvcid) },
+ links => {},
+ };
+ } elsif ($p[0] eq 'hosts' && @p == 2) {
+ err($exists) if $m->{hosts}->{ $p[1] };
+ $m->{hosts}->{ $p[1] } = host_entry();
+ } else {
+ $self->violation("mkdir '$path'");
+ }
+ }
+ return ();
+ }
+
+ sub cmd_chmod($self, $mode, @paths) {
+ $self->available;
+ $self->violation("chmod '$mode'") if $mode ne '0600';
+ for my $path (@paths) {
+ my @p = $self->parts($path);
+ $self->violation("chmod '$path'")
+ if @p != 3 || $p[0] ne 'hosts' || $p[2] !~ /\Adhchap_(?:ctrl_)?key\z/;
+ my $host = $self->{m}->{hosts}->{ $p[1] }
+ // err("chmod: cannot access '$path': No such file or directory");
+ $host->{mode}->{ $p[2] } = $mode;
+ }
+ return ();
+ }
+
+ sub cmd_rmdir($self, @paths) {
+ $self->available;
+ my $m = $self->{m};
+ my @errors;
+ for my $path (@paths) {
+ my @p = $self->parts($path) or $self->violation("rmdir '$path'");
+ my $fail = sub($why) { push @errors, "rmdir: failed to remove '$path': $why" };
+ if ($p[0] eq 'subsystems' && @p == 4 && $p[2] eq 'namespaces') {
+ my $subsys = $m->{subsystems}->{ $p[1] };
+ $fail->('No such file or directory')
+ if !$subsys || !delete($subsys->{ns}->{ $p[3] });
+ } elsif ($p[0] eq 'subsystems' && @p == 2) {
+ my $subsys = $m->{subsystems}->{ $p[1] };
+ if (!$subsys) {
+ $fail->('No such file or directory');
+ } elsif ($subsys->{ns}->%* || $subsys->{acl}->%*) {
+ $fail->('Directory not empty');
+ } elsif (grep { $_->{links}->{ $p[1] } } values $m->{ports}->%*) {
+ $fail->('Device or resource busy');
+ } else {
+ delete $m->{subsystems}->{ $p[1] };
+ }
+ } elsif ($p[0] eq 'hosts' && @p == 2) {
+ if (!$m->{hosts}->{ $p[1] }) {
+ $fail->('No such file or directory');
+ } elsif (grep { $_->{acl}->{ $p[1] } } values $m->{subsystems}->%*) {
+ $fail->('Device or resource busy');
+ } else {
+ delete $m->{hosts}->{ $p[1] };
+ }
+ } else {
+ $self->violation("rmdir '$path'");
+ }
+ }
+ err($errors[0]) if @errors;
+ return ();
+ }
+
+ sub cmd_rm($self, @paths) {
+ $self->available;
+ my $m = $self->{m};
+ my @errors;
+ for my $path (@paths) {
+ my @p = $self->parts($path);
+ my $missing = "rm: cannot remove '$path': No such file or directory";
+ if (@p == 4 && $p[0] eq 'ports' && $p[2] eq 'subsystems') {
+ my $port = $m->{ports}->{ $p[1] };
+ push @errors, $missing if !$port || !delete($port->{links}->{ $p[3] });
+ } elsif (@p == 4 && $p[0] eq 'subsystems' && $p[2] eq 'allowed_hosts') {
+ my $subsys = $m->{subsystems}->{ $p[1] };
+ push @errors, $missing if !$subsys || !delete($subsys->{acl}->{ $p[3] });
+ } else {
+ $self->violation("rm '$path'");
+ }
+ }
+ err($errors[0]) if @errors;
+ return ();
+ }
+
+ sub cmd_ln($self, $flag, $target, $link) {
+ $self->available;
+ my $m = $self->{m};
+ my @t = $self->parts($target);
+ my @l = $self->parts($link);
+ my $fail = sub($why) { err("ln: failed to create symbolic link '$link': $why") };
+ if (@l == 4 && $l[0] eq 'ports' && $l[2] eq 'subsystems') {
+ $self->violation("ln '$target' '$link'")
+ if @t != 2 || $t[0] ne 'subsystems' || $t[1] ne $l[3];
+ my $port = $m->{ports}->{ $l[1] } // $fail->('No such file or directory');
+ $fail->('No such file or directory') if !$m->{subsystems}->{ $t[1] };
+ $fail->('File exists') if $port->{links}->{ $l[3] };
+ my $a = $port->{attr};
+ $fail->('Invalid argument')
+ if $a->{addr_trtype} ne 'tcp'
+ || $a->{addr_adrfam} !~ /\Aipv[46]\z/
+ || $a->{addr_traddr} eq ''
+ || $a->{addr_trsvcid} eq '';
+ if (!$port->{links}->%*) { # the first link enables the port
+ $fail->('Cannot assign requested address') if $self->{unbindable}->{ $l[1] };
+ my @fields = qw(addr_trtype addr_adrfam addr_traddr addr_trsvcid);
+ my $address = join(' ', $a->@{@fields});
+ for my $id (keys $m->{ports}->%*) {
+ my $other = $m->{ports}->{$id};
+ next if $id eq $l[1] || !$other->{links}->%*;
+ $fail->('Address already in use')
+ if join(' ', $other->{attr}->@{@fields}) eq $address;
+ }
+ }
+ $port->{links}->{ $l[3] } = 1;
+ } elsif (@l == 4 && $l[0] eq 'subsystems' && $l[2] eq 'allowed_hosts') {
+ $self->violation("ln '$target' '$link'")
+ if @t != 2 || $t[0] ne 'hosts' || $t[1] ne $l[3];
+ my $subsys = $m->{subsystems}->{ $l[1] } // $fail->('No such file or directory');
+ $fail->('No such file or directory') if !$m->{hosts}->{ $t[1] };
+ $fail->('Invalid argument') if $subsys->{attr}->{attr_allow_any_host} eq '1';
+ $fail->('File exists') if $subsys->{acl}->{ $l[3] };
+ $subsys->{acl}->{ $l[3] } = 1;
+ } else {
+ $self->violation("ln '$target' '$link'");
+ }
+ return ();
+ }
+
+ sub cmd_test($self, $flag, $path) {
+ if ($flag eq '-b') {
+ err('') if !$self->block_device($path, 1);
+ return ();
+ }
+ $self->available;
+ my $m = $self->{m};
+ my @p = $self->parts($path);
+ my $found;
+ if (@p == 4 && $p[0] eq 'subsystems' && $p[2] eq 'allowed_hosts') {
+ $found = ($m->{subsystems}->{ $p[1] } // {})->{acl}->{ $p[3] };
+ } elsif (@p == 4 && $p[0] eq 'ports' && $p[2] eq 'subsystems') {
+ $found = ($m->{ports}->{ $p[1] } // {})->{links}->{ $p[3] };
+ } else {
+ $self->violation("test $flag '$path'");
+ }
+ err('') if !$found;
+ return ();
+ }
+
+ sub cmd_grep($self, $flag, $value, $path) {
+ $self->available;
+ my @p = $self->parts($path);
+ $self->violation("grep '$path'")
+ if @p != 3 || $p[0] ne 'subsystems' || $p[2] ne 'attr_allow_any_host';
+ my $subsys = $self->{m}->{subsystems}->{ $p[1] }
+ // err("grep: $path: No such file or directory");
+ err('') if $subsys->{attr}->{attr_allow_any_host} ne $value;
+ return ();
+ }
+
+ sub cmd_cat($self, $path) {
+ $self->violation("cat '$path'") if $path ne '/proc/mounts';
+ return (
+ 'sysfs /sys sysfs rw,nosuid,nodev,noexec,relatime 0 0',
+ 'proc /proc proc rw,nosuid,nodev,noexec,relatime 0 0',
+ (
+ $self->{m}->{mounted}
+ ? ('configfs /sys/kernel/config configfs rw,relatime 0 0')
+ : ()
+ ),
+ );
+ }
+
+ sub cmd_mount($self, @args) {
+ err('mount: /sys/kernel/config: none already mounted on /sys/kernel/config.')
+ if $self->{m}->{mounted};
+ $self->{m}->{mounted} = 1;
+ return ();
+ }
+
+ sub cmd_modprobe($self, $module) {
+ $self->violation("modprobe '$module'") if $module ne 'nvmet_tcp';
+ $self->{m}->{loaded} = 1;
+ return ();
+ }
+
+ sub cmd_printf($self, @args) {
+ $self->violation('printf') if @args != 2 || $args[0] ne '%s\n' || $args[1] ne $MARKER;
+ return ($MARKER);
+ }
+
+ # The configfs read; 'golden commands' pins its argv.
+ sub cmd_env($self, @args) {
+ my @hosts = map { $args[$_ + 1] =~ m{\A$RR/hosts/([^/]+)/dhchap_key\z} ? ($1) : () }
+ grep { $args[$_] eq '-path' } 0 .. $#args;
+ $self->available;
+ return $self->configfs_lines(\@hosts);
+ }
+
+ sub configfs_lines($self, $hosts) {
+ my $m = $self->{m};
+ my (@dirs, @links, @attrs, @digests);
+ push @dirs, "D $ROOT", map { "D $ROOT/$_" } qw(hosts ports subsystems);
+ push @dirs, map { "D $ROOT/hosts/$_" } sort keys $m->{hosts}->%*;
+ for my $id (sort { $a <=> $b } keys $m->{ports}->%*) {
+ my $port = $m->{ports}->{$id};
+ push @dirs, "D $ROOT/ports/$id", "D $ROOT/ports/$id/subsystems";
+ push @links, map { "L $ROOT/ports/$id/subsystems/$_" } sort keys $port->{links}->%*;
+ push @attrs,
+ map { "$ROOT/ports/$id/$_:$port->{attr}->{$_}" } sort keys $port->{attr}->%*;
+ }
+ for my $nqn (sort keys $m->{subsystems}->%*) {
+ my $subsys = $m->{subsystems}->{$nqn};
+ my $S = "$ROOT/subsystems/$nqn";
+ push @dirs, "D $S", "D $S/allowed_hosts", "D $S/namespaces";
+ push @links, map { "L $S/allowed_hosts/$_" } sort keys $subsys->{acl}->%*;
+ push @attrs, map { "$S/$_:$subsys->{attr}->{$_}" } sort keys $subsys->{attr}->%*;
+ for my $id (sort { $a <=> $b } keys $subsys->{ns}->%*) {
+ my $ns = $subsys->{ns}->{$id};
+ push @dirs, "D $S/namespaces/$id";
+ push @attrs,
+ map { "$S/namespaces/$id/$_:$ns->{$_}" }
+ qw(buffered_io enable device_uuid device_path);
+ }
+ }
+ for my $hostnqn ($hosts->@*) {
+ my $host = $m->{hosts}->{$hostnqn} // next;
+ push @digests,
+ sha256_hex(($host->{key} // '') . "\n") . " $ROOT/hosts/$hostnqn/dhchap_key";
+ }
+ return (@dirs, @links, @attrs, @digests);
+ }
+
+ # --- zfs --------------------------------------------------------------
+
+ sub cmd_zfs($self, $subcommand, @args) {
+ my $method = "zfs_$subcommand";
+ $self->violation("zfs $subcommand") if !$self->can($method);
+ return $self->$method(@args);
+ }
+
+ # Options with values are listed with 1, flags with 0.
+ sub options($self, $args, %spec) {
+ my %o;
+ my @rest = $args->@*;
+ while (@rest && $rest[0] =~ /\A-/) {
+ my $option = shift @rest;
+ if ($option eq '-Hp') {
+ $o{H} = $o{p} = 1;
+ next;
+ }
+ if ($option =~ /\A-d([0-9]+)\z/) {
+ $o{d} = [$1];
+ next;
+ }
+ my $name = substr($option, 1);
+ $self->violation("zfs option '$option'") if !exists($spec{$name});
+ if ($spec{$name}) {
+ push $o{$name}->@*, shift(@rest);
+ } else {
+ $o{$name} = 1;
+ }
+ }
+ return (\%o, @rest);
+ }
+
+ sub properties_of($self, $assignments) {
+ my %props;
+ for my $assignment (($assignments // [])->@*) {
+ $self->violation("zfs property '$assignment'") if $assignment !~ /\A([^=]+)=(.*)\z/;
+ $props{$1} = $2;
+ }
+ return \%props;
+ }
+
+ sub datasets_under($self, $target, $depth) {
+ my $m = $self->{m};
+ return () if !$m->{ds}->{$target};
+ return sort grep {
+ my $name = $_;
+ $name eq $target
+ || (index($name, "$target/") == 0
+ && (() = substr($name, length($target)) =~ m{/}g) <= $depth)
+ } keys $m->{ds}->%*;
+ }
+
+ sub property($self, $name, $prop) {
+ my $m = $self->{m};
+ my $ds = $m->{ds}->{$name};
+ return ($ds->{type}, '-') if $prop eq 'type';
+ if ($prop =~ /:/) {
+ return ($ds->{props}->{$prop}, 'local') if defined($ds->{props}->{$prop});
+ my $parent = $name;
+ while ($parent =~ s{/[^/]+\z}{}) {
+ my $value = ($m->{ds}->{$parent} // last)->{props}->{$prop};
+ return ($value, "inherited from $parent") if defined($value);
+ }
+ return ('-', '-');
+ }
+ return ($m->{available}, '-') if $prop eq 'available';
+ return ($m->{used}, '-') if $prop eq 'used';
+ return ($ds->{volsize} // '-', '-') if $prop eq 'volsize';
+ return ('4096', '-') if $prop eq 'usedbydataset';
+ $self->violation("zfs property '$prop'");
+ }
+
+ sub zfs_get($self, @args) {
+ my ($o, $props, @targets) = $self->options(\@args, H => 0, p => 0, d => 1, t => 1, o => 1);
+ $self->violation('zfs get') if @targets != 1 || !$o->{H};
+ $self->violation('zfs get inventory')
+ if $props =~ /\Atype,proxmox:/
+ && !main::same(['get', @args], main::expected_inventory_read($targets[0]));
+ my $target = $targets[0];
+ my @fields = split /,/, ($o->{o} // ['name,property,value,source'])->[0];
+ my %types = map { $_ => 1 } split /,/, ($o->{t} // ['filesystem,volume'])->[0];
+ my @datasets = $self->datasets_under($target, $o->{d} ? $o->{d}->[0] : 0)
+ or err("cannot open '$target': dataset does not exist");
+ my @out;
+
+ for my $name (grep { $types{ $self->{m}->{ds}->{$_}->{type} } } @datasets) {
+ for my $prop (split /,/, $props) {
+ my ($value, $source) = $self->property($name, $prop);
+ my %row = (name => $name, property => $prop, value => $value, source => $source);
+ push @out, join("\t", map { $row{$_} } @fields);
+ }
+ }
+ return @out;
+ }
+
+ sub list_field($self, $row, $field) {
+ my $m = $self->{m};
+ return $row if $field eq 'name';
+ if (my $snap = $m->{snaps}->{$row}) {
+ return 1000 + $snap->{seq} if $field eq 'guid';
+ return 1700000000 + $snap->{seq} if $field eq 'creation';
+ $self->violation("zfs list snapshot field '$field'");
+ }
+ my $ds = $m->{ds}->{$row};
+ return $ds->{type} eq 'volume' ? $ds->{volsize} : '-' if $field eq 'volsize';
+ return $ds->{origin} // '-' if $field eq 'origin';
+ return $ds->{type} if $field eq 'type';
+ $self->violation("zfs list field '$field'");
+ }
+
+ sub zfs_list($self, @args) {
+ my ($o, @targets) =
+ $self->options(\@args, H => 0, p => 0, r => 0, d => 1, t => 1, o => 1, s => 1);
+ $self->violation('zfs list') if @targets != 1 || !$o->{H};
+ my $m = $self->{m};
+ my $target = $targets[0];
+ my @fields = split /,/, ($o->{o} // ['name'])->[0];
+ my %types = map { $_ => 1 } split /,/, ($o->{t} // ['filesystem,volume'])->[0];
+ my @rows;
+ if ($target =~ /\@/) {
+ err("cannot open '$target': dataset does not exist") if !$m->{snaps}->{$target};
+ @rows = ($target);
+ } else {
+ my $depth = $o->{d} ? $o->{d}->[0] : $o->{r} ? 1000 : 0;
+ my @datasets = $self->datasets_under($target, $depth)
+ or err("cannot open '$target': dataset does not exist");
+ if ($types{snapshot}) {
+ my %under = map { $_ => 1 } @datasets;
+ push @rows, sort { $m->{snaps}->{$a}->{seq} <=> $m->{snaps}->{$b}->{seq} }
+ grep { $under{ (split /\@/)[0] } } keys $m->{snaps}->%*;
+ }
+ push @rows, grep { $types{ $m->{ds}->{$_}->{type} } } @datasets;
+ }
+ return map {
+ my $row = $_;
+ join("\t", map { $self->list_field($row, $_) } @fields)
+ } @rows;
+ }
+
+ sub zfs_create($self, @args) {
+ my ($o, @targets) = $self->options(\@args, s => 0, b => 1, o => 1, V => 1);
+ $self->violation('zfs create') if @targets != 1 || !$o->{V};
+ my $m = $self->{m};
+ my $name = $targets[0];
+ my ($parent) = $name =~ m{\A(.+)/[^/]+\z} or $self->violation('zfs create name');
+ $self->violation('zfs create size') if $o->{V}->[0] !~ /\A([0-9]+)k\z/;
+ my $size = $1 * 1024;
+ err("cannot create '$name': parent does not exist") if !$m->{ds}->{$parent};
+ err("cannot create '$name': dataset already exists") if $m->{ds}->{$name};
+ $m->{ds}->{$name} = {
+ type => 'volume',
+ props => $self->properties_of($o->{o}),
+ volsize => $size,
+ origin => '-',
+ devwait => $self->{udev_delay},
+ sparse => $o->{s} ? 1 : 0,
+ blocksize => $o->{b} ? $o->{b}->[0] : undef,
+ };
+ return ();
+ }
+
+ sub zfs_clone($self, @args) {
+ my ($o, $origin, $name, @extra) = $self->options(\@args, o => 1);
+ $self->violation('zfs clone') if @extra || !defined($name);
+ my $m = $self->{m};
+ my $snap = $m->{snaps}->{$origin} // err("cannot open '$origin': dataset does not exist");
+ err("cannot create '$name': dataset already exists") if $m->{ds}->{$name};
+ $m->{ds}->{$name} = {
+ type => 'volume',
+ props => $self->properties_of($o->{o}),
+ volsize => $snap->{volsize},
+ origin => $origin,
+ devwait => $self->{udev_delay},
+ };
+ return ();
+ }
+
+ sub zfs_destroy($self, @args) {
+ my ($o, $target, @extra) = $self->options(\@args, r => 0);
+ $self->violation('zfs destroy') if @extra || !defined($target);
+ my $m = $self->{m};
+ my $clones = sub($snapshot) {
+ return grep { ($_->{origin} // '-') eq $snapshot } values $m->{ds}->%*;
+ };
+ if ($target =~ /\@/) {
+ $self->violation('zfs destroy -r of a snapshot') if $o->{r};
+ err("could not find any snapshots to destroy; check snapshot names.")
+ if !$m->{snaps}->{$target};
+ err("cannot destroy snapshot $target: snapshot has dependent clones")
+ if $clones->($target);
+ delete $m->{snaps}->{$target};
+ return ();
+ }
+ $self->violation('zfs destroy without -r') if !$o->{r};
+ err("cannot open '$target': dataset does not exist") if !$m->{ds}->{$target};
+ err("cannot destroy '$target': dataset is busy") if $self->busy($target);
+ my @snaps = grep { index($_, "$target\@") == 0 } keys $m->{snaps}->%*;
+ err("cannot destroy '$target': filesystem has dependent clones")
+ if grep { $clones->($_) } @snaps;
+ delete $m->{snaps}->{$_} for @snaps;
+ delete $m->{ds}->{$target};
+ return ();
+ }
+
+ # Stricter than ZFS: user properties revert to their value at snapshot
+ # time, so the plugin must set the identity again after a rollback.
+ sub zfs_rollback($self, @args) {
+ $self->violation('zfs rollback') if @args != 1 || $args[0] !~ /\@/;
+ my $m = $self->{m};
+ my ($name) = split /\@/, $args[0];
+ my $snap = $m->{snaps}->{ $args[0] }
+ // err("cannot open '$args[0]': dataset does not exist");
+ err("cannot rollback to '$args[0]': more recent snapshots or bookmarks exist")
+ if grep { index($_, "$name\@") == 0 && $m->{snaps}->{$_}->{seq} > $snap->{seq} }
+ keys $m->{snaps}->%*;
+ err("cannot rollback '$name': dataset is busy") if $self->busy($name);
+ my $ds = $m->{ds}->{$name};
+ $ds->{props} = dclone($snap->{props});
+ $ds->{volsize} = $snap->{volsize};
+ $ds->{devwait} = $self->{udev_delay};
+ return ();
+ }
+
+ sub zfs_set($self, @args) {
+ my $target = pop @args;
+ $self->violation('zfs set') if !@args;
+ my $m = $self->{m};
+ my $object = $target =~ /\@/ ? $m->{snaps}->{$target} : $m->{ds}->{$target};
+ err("cannot open '$target': dataset does not exist") if !$object;
+ for my $assignment (@args) {
+ $self->violation("zfs set '$assignment'") if $assignment !~ /\A([^=]+)=(.*)\z/;
+ my ($prop, $value) = ($1, $2);
+ if ($prop eq 'volsize') {
+ $self->violation('zfs set volsize') if $value !~ /\A([0-9]+)k\z/;
+ $object->{volsize} = $1 * 1024;
+ } elsif ($prop =~ /:/) {
+ $object->{props}->{$prop} = $value;
+ } else {
+ $self->violation("zfs set '$prop'");
+ }
+ }
+ return ();
+ }
+
+ # Like OpenZFS, which renames the minor of a zvol that is open: an
+ # exported zvol is renamed under its namespace.
+ sub zfs_rename($self, @args) {
+ $self->violation('zfs rename') if @args != 2;
+ my ($old, $new) = @args;
+ my $m = $self->{m};
+ err("cannot open '$old': dataset does not exist") if !$m->{ds}->{$old};
+ err("cannot rename to '$new': dataset already exists") if $m->{ds}->{$new};
+ $m->{ds}->{$new} = delete $m->{ds}->{$old};
+ for my $snap (grep { index($_, "$old\@") == 0 } keys $m->{snaps}->%*) {
+ my $renamed = $new . substr($snap, length($old));
+ $m->{snaps}->{$renamed} = delete $m->{snaps}->{$snap};
+ for my $ds (values $m->{ds}->%*) {
+ $ds->{origin} = $renamed if ($ds->{origin} // '-') eq $snap;
+ }
+ }
+ $m->{ds}->{$new}->{devwait} = $self->{udev_delay};
+ return ();
+ }
+
+ sub zfs_snapshot($self, @args) {
+ $self->violation('zfs snapshot') if @args != 1 || $args[0] !~ /\A([^@]+)\@(.+)\z/;
+ my ($name, $snap) = ($1, $2);
+ err("cannot create snapshot '$args[0]': invalid character in name")
+ if $snap !~ /\A[A-Za-z0-9_.: -]+\z/;
+ my $m = $self->{m};
+ my $ds = $m->{ds}->{$name} // err("cannot open '$name': dataset does not exist");
+ err("cannot create snapshot '$args[0]': dataset already exists")
+ if $m->{snaps}->{ $args[0] };
+ $m->{snaps}->{ $args[0] } =
+ { props => dclone($ds->{props}), volsize => $ds->{volsize}, seq => ++$m->{seq} };
+ return ();
+ }
+}
+
+# The inventory read written down independently of the plugin code.
+my $INVENTORY_PROPS =
+ 'type,proxmox:nvme-subsys,proxmox:nvme-nsid,proxmox:nvme-uuid,proxmox:nvme-last-nsid';
+
+sub expected_inventory_read($pool) {
+ return [
+ 'get',
+ '-H',
+ '-p',
+ '-d',
+ '1',
+ '-t',
+ 'filesystem,volume',
+ '-o',
+ 'name,property,value,source',
+ $INVENTORY_PROPS,
+ $pool,
+ ];
+}
+
+# ---------------------------------------------------------------------------
+# Flow helpers and invariants
+# ---------------------------------------------------------------------------
+
+# Runs $code against $fake and returns what happened during the run.
+sub flow($fake, $code) {
+ local $FAKE = $fake;
+ local @LOCKS = ();
+ local @NESTED_LOCKS = ();
+ local @WARNINGS = ();
+ local @SYSFS_WRITES = ();
+
+ my $start = scalar($fake->{calls}->@*);
+ my $result = eval { $code->() };
+ my $error = $@;
+ my @calls = $fake->{calls}->@*;
+ return {
+ error => $error,
+ result => $result,
+ calls => [@calls[$start .. $#calls]],
+ locks => [@LOCKS],
+ nested => [@NESTED_LOCKS],
+ warnings => [@WARNINGS],
+ sysfs_writes => [@SYSFS_WRITES],
+ };
+}
+
+sub ops($res) {
+ return [map { $_->{op} } $res->{calls}->@*];
+}
+
+sub changes($res) {
+ return scalar(grep { $_->{changing} } $res->{calls}->@*);
+}
+
+# Our subsystem published on two ports to two hosts, no volumes yet.
+sub target_fake(%opts) {
+ my $fake = FakeTarget->new(pools => ['tank', 'other'], %opts);
+ $fake->add_host($_, $KEY) for @HOSTS;
+ $fake->add_subsystem($NQN, acl => [@HOSTS]);
+ $fake->add_port(1, '192.0.2.21', 4420, links => [$NQN]);
+ $fake->add_port(2, '192.0.2.22', 4420, links => [$NQN]);
+ return $fake;
+}
+
+sub add_volume($fake, $name, $nsid, $uuid, %o) {
+ $fake->add_zvol("tank/$name", identity => [$NQN, $nsid, $uuid]);
+ $fake->add_namespace($NQN, $nsid, $uuid, "/dev/zvol/tank/$name", $o{enable} // 1)
+ if !$o{unexported};
+ return $fake;
+}
+
+# A converged target with three owned volumes, an unowned leftover, a volume
+# of another subsystem in our pool, and a foreign subsystem that shares the
+# first host and port 1.
+sub lifecycle_fake(%opts) {
+ my $fake = target_fake(%opts);
+ add_volume($fake, 'vm-100-disk-0', 1, $U{1});
+ add_volume($fake, 'vm-101-disk-0', 2, $U{2});
+ add_volume($fake, 'base-102-disk-0', 3, $U{3});
+ $fake->add_snapshot('tank/vm-100-disk-0@snap1');
+ $fake->add_snapshot('tank/base-102-disk-0@__base__');
+ $fake->add_zvol('tank/vm-900-disk-0');
+ $fake->add_zvol('tank/vm-901-disk-0', identity => [$FOREIGN_NQN, 7, $U{7}]);
+ $fake->add_zvol('other/foreign-disk', identity => [$FOREIGN_NQN, 1, $U{8}]);
+ $fake->add_subsystem(
+ $FOREIGN_NQN,
+ model => 'Linux',
+ serial => '0123456789abcdef',
+ acl => [$HOSTS[0]],
+ );
+ $fake->add_namespace($FOREIGN_NQN, 1, $U{8}, '/dev/zvol/other/foreign-disk');
+ $fake->{m}->{ports}->{1}->{links}->{$FOREIGN_NQN} = 1;
+ return $fake;
+}
+
+sub owned_volumes($m, $nqn = $NQN, $pool = 'tank') {
+ my %owned;
+ for my $name (sort keys $m->{ds}->%*) {
+ my $ds = $m->{ds}->{$name};
+ next if $ds->{type} ne 'volume' || $name !~ m{\A\Q$pool\E/[^/]+\z};
+ next if ($ds->{props}->{'proxmox:nvme-subsys'} // '') ne $nqn;
+ $owned{$name} = [map { $ds->{props}->{"proxmox:nvme-$_"} } qw(nsid uuid)];
+ }
+ return \%owned;
+}
+
+# Everything on the target that does not belong to the storage under test.
+# The last NSID handed out, on its pool, belongs to the storage.
+sub foreign_view($m, $port_ids, $nqn, $hostnqns, $pool = 'tank') {
+ my %ours = map { $_ => 1 } $hostnqns->@*;
+ my $dataset = sub($name) {
+ my $ds = $m->{ds}->{$name};
+ return $ds if $name ne $pool;
+ my %props = $ds->{props}->%*;
+ delete $props{'proxmox:nvme-last-nsid'};
+ return { $ds->%*, props => \%props };
+ };
+ return {
+ subsystems =>
+ { map { $_ => $m->{subsystems}->{$_} } grep { $_ ne $nqn } keys $m->{subsystems}->%* },
+ datasets => {
+ map { $_ => $dataset->($_) }
+ grep { ($m->{ds}->{$_}->{props}->{'proxmox:nvme-subsys'} // '') ne $nqn }
+ keys $m->{ds}->%*
+ },
+ hosts => { map { $_ => $m->{hosts}->{$_} } grep { !$ours{$_} } keys $m->{hosts}->%* },
+ ports => {
+ map {
+ my $port = $m->{ports}->{$_};
+ my @links = $port ? sort grep { $_ ne $nqn } keys $port->{links}->%* : ();
+ ($_ => $port ? [$port->{attr}, \@links] : undef)
+ } $port_ids->@*
+ },
+ };
+}
+
+# (i), (ii), (vi) and (vii), checked after every call.
+sub online_violations($fake, $ctx) {
+ my $m = $fake->{m};
+ my $nqn = $ctx->{nqn};
+ my @bad;
+ if (my $subsys = $m->{subsystems}->{$nqn}) {
+ my $owned = owned_volumes($m, $nqn, $ctx->{pool});
+ my %identity = map { ("/dev/zvol/$_" => $owned->{$_}) } keys $owned->%*;
+ my %seen;
+ for my $id (sort keys $subsys->{ns}->%*) {
+ my $ns = $subsys->{ns}->{$id};
+ push @bad, "(ii) UUID $ns->{device_uuid} used by two namespaces"
+ if $seen{ lc($ns->{device_uuid}) }++;
+ next if $ns->{enable} ne '1';
+ my $want = $identity{ $ns->{device_path} };
+ push @bad, "(i) namespace $id exports $ns->{device_path} under another identity"
+ if !$want || $want->[0] ne $id || lc($want->[1]) ne lc($ns->{device_uuid});
+ }
+ if (grep { $_->{links}->{$nqn} } values $m->{ports}->%*) {
+ push @bad, '(vi) published while any host may connect'
+ if $subsys->{attr}->{attr_allow_any_host} ne '0';
+ for my $hostnqn ($ctx->{teardown} ? () : $ctx->{hosts}->@*) {
+ push @bad, "(vi) published without the ACL of $hostnqn"
+ if !$subsys->{acl}->{$hostnqn};
+ push @bad, "(vi) published without the key of $hostnqn"
+ if (($m->{hosts}->{$hostnqn} // {})->{key} // '') ne $ctx->{key};
+ }
+ }
+ }
+ push @bad,
+ '(vii) objects of other storages changed'
+ if !same(
+ foreign_view($m, $ctx->{port_ids}, $nqn, $ctx->{hosts}, $ctx->{pool}),
+ $ctx->{foreign},
+ );
+ return @bad;
+}
+
+# (iii): owned identities stay valid, unique and attached to their volume.
+sub identity_violations($m, $ctx) {
+ my @bad;
+ my $owned = owned_volumes($m, $ctx->{nqn}, $ctx->{pool});
+ my (%nsids, %uuids);
+ for my $name (sort keys $owned->%*) {
+ my ($nsid, $uuid) = $owned->{$name}->@*;
+ push @bad, "(iii) invalid identity on $name"
+ if ($nsid // '') !~ /\A[1-9][0-9]*\z/ || ($uuid // '') !~ $UUID_RE;
+ push @bad, "(iii) NSID $nsid on two volumes" if $nsids{ $nsid // '' }++;
+ push @bad, "(iii) UUID $uuid on two volumes" if $uuids{ lc($uuid // '') }++;
+ }
+ for my $name (sort keys $ctx->{initial_owned}->%*) {
+ my $identity = join('/', $ctx->{initial_owned}->{$name}->@*);
+ my @holders = grep { join('/', $owned->{$_}->@*) eq $identity } keys $owned->%*;
+ next if !@holders && ($ctx->{may_vanish} // '') eq $name;
+ my %allowed = map { $_ => 1 } $name, ($ctx->{renames}->{$name} // ());
+ push @bad, "(iii) identity of $name lost" if @holders != 1 || !$allowed{ $holders[0] };
+ }
+ return @bad;
+}
+
+# A creation cut right after its mkdir leaves an unused subsystem with the
+# kernel's default model, which another tool could have created as well.
+# Activation refuses it until an administrator removes it.
+sub default_subsystem($m, $nqn) {
+ my $subsys = $m->{subsystems}->{$nqn} // return 0;
+ return
+ $subsys->{attr}->{attr_model} eq 'Linux'
+ && !$subsys->{ns}->%*
+ && !$subsys->{acl}->%*
+ && !grep { $_->{links}->{$nqn} } values $m->{ports}->%*;
+}
+
+# (v): one more activation exports every owned volume exactly once.
+sub convergence_violations($fake, $ctx) {
+ my $res = flow($fake, sub { activate($ctx->{scfg}, $ctx->{key}, $ctx->{hosts}) });
+ if (my $error = $res->{error}) {
+ return ()
+ if $error =~ /not created by Proxmox VE/ && default_subsystem($fake->{m}, $ctx->{nqn});
+ return "(v) activation afterwards failed: $error";
+ }
+ my $m = $fake->{m};
+ my $owned = owned_volumes($m, $ctx->{nqn}, $ctx->{pool});
+ my $subsys = $m->{subsystems}->{ $ctx->{nqn} } // return '(v) subsystem missing';
+ my %want = map { $owned->{$_}->[0] => ["/dev/zvol/$_", lc($owned->{$_}->[1]), '1'] }
+ keys $owned->%*;
+ my %have = map {
+ my $ns = $subsys->{ns}->{$_};
+ ($_ => [$ns->{device_path}, lc($ns->{device_uuid}), $ns->{enable}])
+ } keys $subsys->{ns}->%*;
+ my @bad;
+ push @bad, '(v) exported namespaces differ from the owned volumes' if !same(\%want, \%have);
+ # An interrupted port creation leaves an incomplete port that is ignored;
+ # the portal is then served by a new port with the same address.
+ for my $portal ($PORTALS->@*) {
+ my ($address, $service) = $portal->@{qw(address port)};
+ push @bad, "(v) not published on $address"
+ if !grep {
+ $_->{links}->{ $ctx->{nqn} }
+ && $_->{attr}->{addr_traddr} eq $address
+ && $_->{attr}->{addr_trsvcid} eq "$service"
+ } values $m->{ports}->%*;
+ }
+ return @bad;
+}
+
+sub run_case($initial, $ctx, $faults, $read_fault = undef) {
+ my $fake = FakeTarget->from_state($initial);
+ $fake->{faults} = $faults;
+ $fake->{read_fault} = $read_fault;
+ my @online;
+ $fake->{after_call} = sub($f, $call) {
+ push @online,
+ map { "call $call->{index} ($call->{op}): $_" } online_violations($f, $ctx);
+ };
+ my $res = flow($fake, $ctx->{action});
+ $res->{fake} = $fake;
+ $res->{online} = \@online;
+ $res->{faults} = { $faults->%*, ($read_fault ? (read => $read_fault->{mode}) : ()) };
+ return $res;
+}
+
+sub case_violations($ctx, $res) {
+ my $fake = $res->{fake};
+ my @bad;
+ push @bad, 'domain lock taken ' . scalar($res->{locks}->@*) . ' times'
+ if $res->{locks}->@* > 1;
+ push @bad, 'nested domain lock request' if $res->{nested}->@*;
+ push @bad, 'key in the error text' if grep { index($res->{error}, $_) >= 0 } $KEY, $KEY_B;
+
+ # The result must match the outcome: success only when the goal is
+ # reached, and a flow whose only faults are lost replies (create and
+ # clone, which cannot settle one with a read) may fail after reaching its
+ # goal only with an unknown target state.
+ my @missing = $ctx->{done}->($fake->{m}, $res);
+ push @bad, map { "success without reaching the goal: $_" } @missing if $res->{error} eq '';
+ my @modes = values $res->{faults}->%*;
+ push @bad, "failure after reaching the goal: $res->{error}"
+ if $ctx->{lost_replies}
+ && $res->{error} ne ''
+ && $res->{error} !~ /the target state is unknown\n\z/
+ && !@missing
+ && @modes
+ && !grep { $_ ne 'after' } @modes;
+
+ # A call that ssh gave up on completes on the target afterwards. When the
+ # rest of it changes the target, the outcome was unknown.
+ if ($fake->{late}->@*) {
+ if (
+ grep {
+ grep { mutating_step($_) } $_->{steps}->@*
+ } $fake->{late}->@*
+ ) {
+ push @bad, "late call after a success" if $res->{error} eq '';
+ push @bad, "late call reported without an unknown state: $res->{error}"
+ if $res->{error} !~ /the target state is unknown\n\z/;
+ }
+ $fake->land;
+ push @bad, map { "after the late call: $_" } online_violations($fake, $ctx);
+ }
+ push @bad, identity_violations($fake->{m}, $ctx);
+ $fake->{faults} = {};
+ $fake->{read_fault} = undef;
+ if ($ctx->{retry}) {
+ my $retry = flow($fake, $ctx->{retry});
+ push @bad, "retry failed: $retry->{error}" if $retry->{error};
+ push @bad, $ctx->{goal}->($fake->{m}) if !$retry->{error};
+ } else {
+ push @bad, convergence_violations($fake, $ctx);
+ }
+ push @bad, $res->{online}->@*;
+ push @bad, map { "protocol: $_" } $fake->{violations}->@*;
+ return @bad;
+}
+
+# Fault sweep: every mutating call of a flow fails, and then every later
+# mutating call of that run fails as well, or the first read after it. After
+# each run: (i) every enabled namespace of the subsystem exports the zvol of
+# its identity, (ii) no UUID is used twice, (iii) owned identities are kept
+# and unique, (v) one more activation (or a retried removal) converges, (vi)
+# the subsystem is never published before its ACLs, keys and
+# allow_any_host=0, (vii) nothing of other storages changes, every change ran
+# under one domain lock, and the result matches the outcome. (i), (ii), (vi)
+# and (vii) are checked after every call; refusals are covered by their own
+# tests.
+sub sweep($name, $ctx) {
+ my $initial = dclone($ctx->{setup}->()->{m});
+ $ctx->{nqn} //= $NQN;
+ $ctx->{pool} //= 'tank';
+ $ctx->{hosts} //= [@HOSTS];
+ $ctx->{key} //= $KEY;
+ $ctx->{scfg} //= scfg();
+ $ctx->{initial_owned} = owned_volumes($initial, $ctx->{nqn}, $ctx->{pool});
+ $ctx->{port_ids} = [sort keys $initial->{ports}->%*];
+ $ctx->{foreign} =
+ foreign_view($initial, $ctx->{port_ids}, $ctx->{nqn}, $ctx->{hosts}, $ctx->{pool});
+
+ my $baseline = run_case($initial, $ctx, {});
+ is($baseline->{error}, '', "$name: succeeds without faults");
+ is_deeply([$ctx->{done}->($baseline->{fake}->{m}, $baseline)], [], "$name: reaches its goal");
+ is_deeply([case_violations($ctx, $baseline)], [], "$name: invariants hold without faults");
+
+ # A call fails in each mode of FakeTarget::execute, at each step of its
+ # chain. A second failure, in the compensation or the next round, is one
+ # that fails before or after its call, or the first read after the first
+ # failure fails or finds the target unreachable.
+ my $modes = sub($call) {
+ my @steps = 1 .. $call->{steps}->$#*;
+ return ('before', 'after', 'late', (map { "cut:$_" } @steps), map { "lost:$_" } @steps);
+ };
+ my @first = grep { $_->{mutating} } $baseline->{calls}->@*;
+ ok(scalar(@first), "$name: changes the target");
+ my ($runs, @failures) = (0);
+ for my $first (@first) {
+ my $k = $first->{index};
+ for my $mode ($modes->($first)) {
+ my $res = run_case($initial, $ctx, { $k => $mode });
+ $runs++;
+ my @later = grep { $_->{mutating} && $_->{index} > $k } $res->{calls}->@*;
+ push @failures, map { "call $k $mode: $_" } case_violations($ctx, $res);
+ for my $second (@later) {
+ my $j = $second->{index};
+ for my $again_mode (qw(before after)) {
+ my $again = run_case($initial, $ctx, { $k => $mode, $j => $again_mode });
+ $runs++;
+ push @failures,
+ map { "call $k $mode, call $j $again_mode: $_" }
+ case_violations($ctx, $again);
+ }
+ }
+ for my $read_mode (qw(before unreachable)) {
+ my $read_fault = { after => $k, mode => $read_mode };
+ my $again = run_case($initial, $ctx, { $k => $mode }, $read_fault);
+ $runs++;
+ push @failures,
+ map { "call $k $mode, next read $read_mode: $_" } case_violations($ctx, $again);
+ }
+ }
+ }
+ ok(!@failures, "$name: the invariants hold in $runs fault injections")
+ or diag(join("\n", @failures));
+}
+
+# ---------------------------------------------------------------------------
+# Fixtures and the fake target formats
+# ---------------------------------------------------------------------------
+
+my $FIXTURE_ZFS = slurp("$FIXTURES/zfs_inventory.txt");
+my $FIXTURE_CFS = slurp("$FIXTURES/configfs_snapshot.txt");
+my $FIXTURE_DIGESTS = join('', map { "$KEY_SHA $ROOT/hosts/$_/dhchap_key\n" } @FIXTURE_HOSTS);
+
+sub fixture_fake() {
+ my $fake = FakeTarget->new(pools => ['tank']);
+ $fake->add_host($_) for @FIXTURE_HOSTS;
+ $fake->add_subsystem($FIXTURE_NQN, acl => [@FIXTURE_HOSTS]);
+ $fake->add_port($_, "192.0.2.2$_", 4420, links => [$FIXTURE_NQN]) for 1, 2;
+ $fake->add_zvol('tank/vm-100-disk-0', identity => [$FIXTURE_NQN, 5, $FIXTURE_UUID5]);
+ $fake->add_zvol('tank/vm-100-cloudinit', identity => [$FIXTURE_NQN, 6, $FIXTURE_UUID6]);
+ $fake->add_namespace($FIXTURE_NQN, 5, $FIXTURE_UUID5, '/dev/zvol/tank/vm-100-disk-0');
+ $fake->add_namespace($FIXTURE_NQN, 6, $FIXTURE_UUID6, '/dev/zvol/tank/vm-100-cloudinit');
+ return $fake;
+}
+
+subtest 'the fake target answers in the formats of the fixture target' => sub {
+ my $fake = fixture_fake();
+ my @zfs = $fake->zfs_get(expected_inventory_read('tank')->@[1 .. 10]);
+ is(join('', map { "$_\n" } @zfs), $FIXTURE_ZFS, 'the ZFS inventory equals the fixture');
+ my $cfs = join('', map { "$_\n" } $fake->configfs_lines([]));
+ is_deeply(
+ nv('_nvmet_parse_configfs', $cfs),
+ nv('_nvmet_parse_configfs', $FIXTURE_CFS),
+ 'the configfs read parses to the fixture',
+ );
+};
+
+# ---------------------------------------------------------------------------
+# Parsers
+# ---------------------------------------------------------------------------
+
+sub zfs_rows($name, $type, $nqn = '-', $nsid = '-', $uuid = '-', $source = 'local', $last = '-') {
+ my $source_of = sub($value) { $value eq '-' ? '-' : $source };
+ return (
+ "$name\ttype\t$type\t-",
+ "$name\tproxmox:nvme-subsys\t$nqn\t" . $source_of->($nqn),
+ "$name\tproxmox:nvme-nsid\t$nsid\t" . $source_of->($nsid),
+ "$name\tproxmox:nvme-uuid\t$uuid\t" . $source_of->($uuid),
+ "$name\tproxmox:nvme-last-nsid\t$last\t" . $source_of->($last),
+ );
+}
+
+sub inventory(@rows) {
+ return nv('_nvmet_parse_zfs_inventory', join('', map { "$_\n" } @rows), 'tank');
+}
+
+my @POOL = zfs_rows('tank', 'filesystem');
+my @ONE = zfs_rows('tank/vm-100-disk-0', 'volume', $NQN, 1, $U{1});
+my @TWO = zfs_rows('tank/vm-101-disk-0', 'volume', $NQN, 2, $U{2});
+
+subtest 'ZFS inventory parser' => sub {
+ is_deeply(
+ inventory(@POOL, @ONE, @TWO),
+ {
+ types => {
+ tank => 'filesystem',
+ 'tank/vm-100-disk-0' => 'volume',
+ 'tank/vm-101-disk-0' => 'volume',
+ },
+ volumes => {
+ 'tank/vm-100-disk-0' => { nqn => $NQN, nsid => 1, uuid => $U{1} },
+ 'tank/vm-101-disk-0' => { nqn => $NQN, nsid => 2, uuid => $U{2} },
+ },
+ last_nsid => 0,
+ },
+ 'complete inventory',
+ );
+ my $pool_last = sub($last, $source = 'local') {
+ return inventory(
+ zfs_rows('tank', 'filesystem', '-', '-', '-', $source, $last),
+ zfs_rows('tank/vm-100-disk-0', 'volume', $NQN, 1, $U{1}, 'local', '-'),
+ )->{last_nsid};
+ };
+ is($pool_last->(7), 7, 'the last NSID handed out is read from the pool');
+ is($pool_last->(7, 'inherited from tank'), 0, 'but not inherited');
+ is($pool_last->('x'), 0, 'and only when it is a valid NSID');
+ is(
+ inventory(@POOL, zfs_rows('tank/vm-100-disk-0', 'volume', $NQN, 1, $U{1}, 'local', 9))
+ ->{last_nsid},
+ 0,
+ 'a volume does not carry it',
+ );
+ my $short_owner = "tank/vm-101-disk-0\tproxmox:nvme-subsys\t$NQN";
+ my $foreign_source = "tank/vm-101-disk-0\tproxmox:nvme-uuid\t$U{2}\tgarbage";
+ for my $case (
+ [
+ 'missing owner row',
+ [@POOL, @ONE, @TWO[0, 2, 3]],
+ qr/incomplete ZFS properties on 'tank\/vm-101-disk-0'/,
+ ],
+ [
+ 'short row',
+ [@POOL, @ONE, $TWO[0], $short_owner, @TWO[2, 3]],
+ qr/malformed ZFS inventory row/,
+ ],
+ [
+ 'duplicate property',
+ [@POOL, @ONE, @TWO, $TWO[1]],
+ qr/duplicate ZFS property 'proxmox:nvme-subsys'/,
+ ],
+ [
+ 'unknown property',
+ [@POOL, @ONE, @TWO, "tank/vm-101-disk-0\tcompression\ton\tlocal"],
+ qr/unexpected ZFS property on 'tank\/vm-101-disk-0'/,
+ ],
+ [
+ 'extra field',
+ [@POOL, @ONE, @TWO[0 .. 2], "$TWO[3]\textra"],
+ qr/malformed ZFS inventory row/,
+ ],
+ [
+ 'unknown source',
+ [@POOL, @ONE, @TWO[0 .. 2], $foreign_source],
+ qr/invalid ZFS property source/,
+ ],
+ ['missing pool rows', [@ONE, @TWO], qr/missing the configured pool/],
+ [
+ 'missing pool type row',
+ [@POOL[1 .. 4], @ONE],
+ qr/incomplete ZFS properties on 'tank'/,
+ ],
+ [
+ 'pool name prefix',
+ [@POOL, zfs_rows('tank2/vm-1-disk-0', 'volume')],
+ qr/outside configured pool/,
+ ],
+ [
+ 'depth 2 dataset',
+ [@POOL, zfs_rows('tank/sub/vm-1-disk-0', 'volume')],
+ qr/outside configured pool/,
+ ],
+ [
+ 'invalid type',
+ [@POOL, "tank/vm-100-disk-0\ttype\tsnapshot\t-", @ONE[1 .. 3]],
+ qr/invalid ZFS dataset type/,
+ ],
+ [
+ 'value with CR',
+ [@POOL, @ONE[0 .. 2], "tank/vm-100-disk-0\tproxmox:nvme-uuid\t$U{1}\r\tlocal"],
+ qr/malformed ZFS inventory row/,
+ ],
+ [
+ 'inherited from a malformed pool',
+ [
+ @POOL,
+ @ONE[0 .. 2],
+ "tank/vm-100-disk-0\tproxmox:nvme-uuid\t$U{1}\tinherited from bad pool",
+ ],
+ qr/invalid ZFS pool name/,
+ ],
+ ) {
+ my ($name, $rows, $error) = $case->@*;
+ eval { inventory($rows->@*) };
+ like($@, $error, "rejects $name");
+ }
+
+ is_deeply(
+ inventory(@POOL),
+ { types => { tank => 'filesystem' }, volumes => {}, last_nsid => 0 },
+ 'a verified empty pool',
+ );
+ my $received = inventory(
+ @POOL, zfs_rows('tank/vm-100-disk-0', 'volume', $NQN, 1, $U{1}, 'received'),
+ );
+ is($received->{volumes}->{'tank/vm-100-disk-0'}->{nqn}, $NQN, 'received properties count');
+ my $inherited = inventory(
+ zfs_rows('tank', 'filesystem', $NQN, '-', '-'),
+ zfs_rows('tank/vm-100-disk-0', 'volume', $NQN, 1, $U{1}, 'inherited from tank'),
+ );
+ is_deeply(
+ $inherited->{volumes}->{'tank/vm-100-disk-0'},
+ { nqn => '-', nsid => '-', uuid => '-' },
+ 'inherited properties never own a volume',
+ );
+ my $subvol = inventory(@POOL, zfs_rows('tank/subvol-100-disk-0', 'filesystem'));
+ ok(
+ $subvol->{types}->{'tank/subvol-100-disk-0'} eq 'filesystem'
+ && !$subvol->{volumes}->{'tank/subvol-100-disk-0'},
+ 'a filesystem child is a dataset but not a volume',
+ );
+
+ # a direct child with a name PVE never uses is carried as not owned
+ my $spaced = inventory(@POOL, zfs_rows('tank/my data', 'volume', $NQN, 9, $U{9}));
+ is($spaced->{types}->{'tank/my data'}, 'volume',
+ 'a name with a space is kept as a dataset');
+ is($spaced->{volumes}->{'tank/my data'}->{nqn}, '-', 'a name with a space is never owned');
+ is_deeply(nv('_nvmet_desired_namespaces', $spaced, $NQN), {}, 'and never exported');
+ eval { inventory(@POOL, "tank/my data\tcompression\ton\tlocal") };
+ is($@, "unexpected ZFS property\n", 'errors do not echo names PVE never uses');
+
+ my $fixture = nv('_nvmet_parse_zfs_inventory', $FIXTURE_ZFS, 'tank');
+ is_deeply(
+ $fixture->{volumes},
+ {
+ 'tank/vm-100-disk-0' => { nqn => $FIXTURE_NQN, nsid => 5, uuid => $FIXTURE_UUID5 },
+ 'tank/vm-100-cloudinit' =>
+ { nqn => $FIXTURE_NQN, nsid => 6, uuid => $FIXTURE_UUID6 },
+ },
+ 'the fixture inventory',
+ );
+ eval { nv('_nvmet_parse_zfs_inventory', $FIXTURE_ZFS, 'bad pool') };
+ like($@, qr/invalid ZFS pool name/, 'the configured pool is validated');
+};
+
+subtest 'ZFS property values and sources' => sub {
+ for my $case (
+ [undef, 'local'],
+ ['', 'local'],
+ ["a\nb", 'local'],
+ ["a\tb", 'local'],
+ ['x', ''],
+ ['x', undef],
+ ['x', "local\r"],
+ ) {
+ eval { nv('_nvmet_property_value', $case->@*) };
+ like($@, qr/malformed ZFS property value/, 'rejects a malformed value or source');
+ }
+ is(nv('_nvmet_property_value', $NQN, 'local'), $NQN, 'local values count');
+ is(nv('_nvmet_property_value', $NQN, 'received'), $NQN, 'received values count');
+ is(nv('_nvmet_property_value', '-', '-'), '-', 'unset values');
+ is(
+ nv('_nvmet_property_value', $NQN, 'inherited from tank/a'),
+ '-',
+ 'inherited values do not',
+ );
+ eval { nv('_nvmet_property_value', $NQN, 'default') };
+ like($@, qr/invalid ZFS property source/, 'rejects other sources');
+ eval { nv('_nvmet_property_value', $NQN, '-') };
+ like($@, qr/invalid ZFS property source/, 'rejects a value without a source');
+ eval { nv('_nvmet_property_value', $NQN, 'inherited from tank;x') };
+ like($@, qr/invalid ZFS pool name/, 'rejects an inherited source outside pool names');
+};
+
+subtest 'configfs parser' => sub {
+ my $with_digests = nv('_nvmet_parse_configfs', $FIXTURE_CFS . $FIXTURE_DIGESTS);
+ is($with_digests->{hosts}->{$_}->{key_sha256}, $KEY_SHA, 'key digest of a fixture host')
+ for @FIXTURE_HOSTS;
+ my $binary = nv(
+ '_nvmet_parse_configfs',
+ $FIXTURE_CFS . join('', map { "$KEY_SHA *$ROOT/hosts/$_/dhchap_key\n" } @FIXTURE_HOSTS),
+ );
+ is($binary->{hosts}->{ $FIXTURE_HOSTS[0] }->{key_sha256}, $KEY_SHA, 'binary-mode digests');
+
+ my $without = sub($re) { return $FIXTURE_CFS =~ s/^$re\n//mr };
+ for my $case (
+ [
+ 'missing root',
+ $without->(qr{D \Q$ROOT\E}),
+ qr/^NVMe target configfs is unavailable$/,
+ ],
+ [
+ 'a namespace without enable',
+ $without->(qr{.*namespaces/5/enable:1}),
+ qr/incomplete NVMe namespace '5' of subsystem '\Q$FIXTURE_NQN\E'/,
+ ],
+ [
+ 'a port without address',
+ $without->(qr{.*ports/2/addr_traddr.*}),
+ qr/incomplete NVMe port '2'/,
+ ],
+ [
+ 'a duplicate attribute',
+ "$FIXTURE_CFS$ROOT/subsystems/$FIXTURE_NQN/attr_serial:X\n",
+ qr/duplicate NVMe target configfs attribute/,
+ ],
+ [
+ 'an attribute without its directory',
+ "$FIXTURE_CFS$ROOT/subsystems/nqn.x:y/attr_model:Z\n",
+ qr/incomplete NVMe subsystem 'nqn.x:y'/,
+ ],
+ [
+ 'a link to a missing port',
+ "${FIXTURE_CFS}L $ROOT/ports/9/subsystems/$FIXTURE_NQN\n",
+ qr/incomplete NVMe port '9'/,
+ ],
+ [
+ 'a digest for a missing host',
+ "$FIXTURE_CFS$KEY_SHA $ROOT/hosts/nqn.x:gone/dhchap_key\n",
+ qr/key digest for unknown NVMe host 'nqn.x:gone'/,
+ ],
+ [
+ 'a malformed digest',
+ $FIXTURE_CFS . substr($KEY_SHA, 1) . " $ROOT/hosts/$FIXTURE_HOSTS[0]/dhchap_key\n",
+ qr/unexpected NVMe target configfs line/,
+ ],
+ [
+ 'a duplicate digest',
+ "$FIXTURE_CFS$FIXTURE_DIGESTS$FIXTURE_DIGESTS",
+ qr/duplicate DH-HMAC-CHAP key digest/,
+ ],
+ ) {
+ my ($name, $text, $error) = $case->@*;
+ eval { nv('_nvmet_parse_configfs', $text) };
+ like($@, $error, "rejects $name");
+ }
+ eval { nv('_nvmet_parse_configfs', "$FIXTURE_CFS$KEY and more\n") };
+ is($@, "unexpected NVMe target configfs line\n", 'a garbage line is rejected without echo');
+ eval { nv('_nvmet_parse_configfs', "${FIXTURE_CFS}D $ROOT/subsystems/bad name\n") };
+ is($@, "incomplete NVMe subsystem\n", 'malformed names are not echoed');
+
+ my $tolerated = nv(
+ '_nvmet_parse_configfs',
+ join(
+ '',
+ $FIXTURE_CFS,
+ "D $ROOT/ports/1/referrals/x\n",
+ "L $ROOT/subsystems/$FIXTURE_NQN/passthru/x\n",
+ "D $ROOT/subsystems/$FIXTURE_NQN/namespaces/5/ana\n",
+ "$ROOT/ports/1/param_inline_data_size:16384\n",
+ "$ROOT/subsystems/$FIXTURE_NQN/attr_version:1.3\n",
+ ),
+ );
+ is_deeply(
+ $tolerated,
+ nv('_nvmet_parse_configfs', $FIXTURE_CFS),
+ 'other objects are ignored',
+ );
+ my $unset = $FIXTURE_CFS =~ s{(namespaces/6/device_path:).*}{$1(null)}r;
+ $unset =~ s{(namespaces/5/device_path:).*}{$1};
+ my $null = nv('_nvmet_parse_configfs', $unset);
+ is(
+ $null->{subsystems}->{$FIXTURE_NQN}->{namespaces}->{6}->{device_path},
+ '(null)',
+ 'a (null) device',
+ );
+ is(
+ $null->{subsystems}->{$FIXTURE_NQN}->{namespaces}->{5}->{device_path},
+ '',
+ 'an empty device',
+ );
+};
+
+subtest 'state marker and mount table' => sub {
+ my ($zfs, $cfs) = nv('_nvmet_split_state', "a\tb\n$MARKER\nD x\n");
+ is_deeply([$zfs, $cfs], ["a\tb\n", "D x\n"], 'splits at the marker');
+ for my $text ("a\n", "a\n$MARKER\nb\n$MARKER\n", "a\n $MARKER\n", "a\n${MARKER}x\n") {
+ eval { nv('_nvmet_split_state', $text) };
+ like(
+ $@,
+ qr/malformed NVMe target state/,
+ 'rejects a missing, repeated or altered marker',
+ );
+ }
+ my $mounts = "sysfs /sys sysfs rw 0 0\nnone /sys/kernel/config configfs rw 0 0\n";
+ ok(nv('_nvmet_parse_mounts', $mounts), 'configfs is mounted');
+ ok(!nv('_nvmet_parse_mounts', "sysfs /sys sysfs rw 0 0\n"), 'configfs is not mounted');
+ ok(!nv('_nvmet_parse_mounts', "none /mnt configfs rw 0 0\n"), 'configfs elsewhere');
+ ok(!nv('_nvmet_parse_mounts', "x /sys/kernel/config tmpfs rw 0 0\n"),
+ 'another file system');
+ ok(!nv('_nvmet_parse_mounts', ''), 'empty mount table');
+};
+
+# ---------------------------------------------------------------------------
+# Planners
+# ---------------------------------------------------------------------------
+
+my $S = "$ROOT/subsystems/$NQN";
+sub write_step($path, $value) { return { write => $path, value => $value } }
+
+# The step that makes the key attributes of hosts readable by root only.
+sub chmod_step(@hosts) {
+ return ['chmod', '0600', map { ("$_/dhchap_key", "$_/dhchap_ctrl_key") } @hosts];
+}
+
+sub cfs_model(%o) {
+ my $model = { subsystems => {}, ports => {}, hosts => {} };
+ if (!$o{no_subsystem}) {
+ $model->{subsystems}->{$NQN} = {
+ attr_model => $MODEL,
+ attr_serial => nv('_nvmet_serial', $NQN),
+ attr_allow_any_host => '0',
+ acl => {},
+ namespaces => {},
+ ($o{subsystem} // {})->%*,
+ };
+ }
+ return $model;
+}
+
+sub ns_model($uuid, $dev, $enable = '1') {
+ return { enable => $enable, device_path => $dev, device_uuid => $uuid, buffered_io => '0' };
+}
+
+sub build_steps($nsid, $uuid, $dev) {
+ my $ns = "$S/namespaces/$nsid";
+ return [
+ ['test', '-b', $dev],
+ ['mkdir', $ns],
+ write_step("$ns/device_path", $dev),
+ write_step("$ns/device_uuid", $uuid),
+ write_step("$ns/buffered_io", 0),
+ write_step("$ns/enable", 1),
+ ];
+}
+
+subtest 'owned identity' => sub {
+ my $inv = inventory(
+ @POOL,
+ @ONE,
+ @TWO,
+ zfs_rows('tank/vm-102-disk-0', 'filesystem'),
+ zfs_rows('tank/vm-103-disk-0', 'volume'),
+ zfs_rows('tank/vm-104-disk-0', 'volume', $FOREIGN_NQN, 1, $U{1}),
+ );
+ is_deeply(
+ [nv('_nvmet_owned_identity', $inv, $NQN, 'tank/vm-100-disk-0')],
+ [1, $U{1}],
+ 'identity of an owned volume; the same identity under another NQN is allowed',
+ );
+ for my $case (
+ ['tank/vm-999-disk-0', qr/does not exist/],
+ ['tank/vm-102-disk-0', qr/is not a ZFS volume/],
+ ['tank/vm-103-disk-0', qr/is not owned by NVMe subsystem/],
+ ['tank/vm-104-disk-0', qr/is not owned by NVMe subsystem/],
+ ) {
+ eval { nv('_nvmet_owned_identity', $inv, $NQN, $case->[0]) };
+ like($@, $case->[1], "refuses $case->[0]");
+ }
+ for my $nsid (0, 0xffffffff, 'x', '01') {
+ my $bad =
+ inventory(@POOL, zfs_rows('tank/vm-100-disk-0', 'volume', $NQN, $nsid, $U{1}));
+ eval { nv('_nvmet_owned_identity', $bad, $NQN, 'tank/vm-100-disk-0') };
+ like($@, qr/invalid NSID/, "refuses NSID '$nsid'");
+ }
+ my $bad = inventory(@POOL, zfs_rows('tank/vm-100-disk-0', 'volume', $NQN, 1, 'nope'));
+ eval { nv('_nvmet_owned_identity', $bad, $NQN, 'tank/vm-100-disk-0') };
+ like($@, qr/invalid namespace UUID/, 'refuses a malformed UUID');
+ for my $other ([1, $U{2}], [2, $U{1}]) {
+ my $dup = inventory(
+ @POOL, @ONE, zfs_rows('tank/vm-101-disk-0', 'volume', $NQN, $other->@*),
+ );
+ eval { nv('_nvmet_owned_identity', $dup, $NQN, 'tank/vm-100-disk-0') };
+ like($@, qr/duplicate NVMe identity on 'tank\/vm-100-disk-0'/, 'refuses a duplicate');
+ }
+ is(nv('_nvmet_serial', $NQN), 'PVEZFS' . substr(sha256_hex($NQN), 0, 14), 'serial formula');
+ is(
+ nv('_nvmet_parse_configfs', $FIXTURE_CFS)->{subsystems}->{$FIXTURE_NQN}->{attr_serial},
+ nv('_nvmet_serial', $FIXTURE_NQN),
+ 'the fixture subsystem carries the serial of its NQN',
+ );
+};
+
+subtest 'namespace ID allocation' => sub {
+ my $rows = sub(@ids) {
+ my $uuid = sub($id) { sprintf('%08d-0000-4000-8000-000000000000', $id) };
+ return inventory(
+ @POOL,
+ map { zfs_rows("tank/vm-$_-disk-0", 'volume', $NQN, $_, $uuid->($_)) } @ids,
+ );
+ };
+ my $cfs = cfs_model();
+ is(nv('_nvmet_allocate_nsid', inventory(@POOL), $cfs, $NQN), 1, 'first NSID');
+ is(nv('_nvmet_allocate_nsid', $rows->(1, 3), $cfs, $NQN), 4, 'a gap is not reused');
+ my $dirs =
+ cfs_model(subsystem => { namespaces => { 7 => ns_model($U{7}, '(null)', '0') } });
+ is(nv('_nvmet_allocate_nsid', $rows->(1), $dirs, $NQN), 8,
+ 'configfs-only namespaces count');
+ my $foreign =
+ inventory(@POOL, zfs_rows('tank/vm-9-disk-0', 'volume', $FOREIGN_NQN, 9, $U{9}));
+ my $foreign_cfs = cfs_model();
+ $foreign_cfs->{subsystems}->{$FOREIGN_NQN} = { namespaces => { 12 => {} } };
+ is(nv('_nvmet_allocate_nsid', $foreign, $foreign_cfs, $NQN), 1, 'others do not count');
+ my $invalid = inventory(@POOL, zfs_rows('tank/vm-1-disk-0', 'volume', $NQN, 'x', $U{1}));
+ is(nv('_nvmet_allocate_nsid', $invalid, $cfs, $NQN), 1, 'invalid NSIDs do not count');
+ my $top = inventory(
+ @POOL,
+ zfs_rows('tank/vm-1-disk-0', 'volume', $NQN, 1, $U{1}),
+ zfs_rows('tank/vm-2-disk-0', 'volume', $NQN, 0xfffffffe, $U{2}),
+ );
+ is(
+ nv('_nvmet_allocate_nsid', $top, $cfs, $NQN),
+ 2,
+ 'after the last NSID the lowest free one',
+ );
+ my $reserved = $rows->(1);
+ $reserved->{last_nsid} = 9;
+ is(nv('_nvmet_allocate_nsid', $reserved, $cfs, $NQN), 10, 'a reserved NSID is not reused');
+ $reserved->{last_nsid} = 0xfffffffe;
+ is(
+ nv('_nvmet_allocate_nsid', $reserved, $cfs, $NQN),
+ 2,
+ 'until the last NSID was handed out',
+ );
+};
+
+subtest 'export planner' => sub {
+ my $dev = '/dev/zvol/tank/vm-100-disk-0';
+ my $case = sub($namespaces) {
+ return cfs_model(subsystem => { namespaces => $namespaces });
+ };
+ my $build = build_steps(1, $U{1}, $dev);
+ my @plans = (
+ ['absent', {}, [$build], 'absent'],
+ ['present', { 1 => ns_model($U{1}, $dev) }, [], 'present'],
+ [
+ 'disabled',
+ { 1 => ns_model($U{1}, $dev, '0') },
+ [[['test', '-b', $dev], write_step("$S/namespaces/1/enable", 1)]],
+ 'disabled',
+ ],
+ [
+ 'disabled and incomplete',
+ { 1 => ns_model($U{9}, '(null)', '0') },
+ [[['rmdir', "$S/namespaces/1"], $build->@*]],
+ 'reclaim',
+ ],
+ [
+ 'enabled on the template name of the zvol',
+ { 1 => ns_model($U{1}, '/dev/zvol/tank/base-100-disk-0') },
+ [[
+ write_step("$S/namespaces/1/enable", 0),
+ ['rmdir', "$S/namespaces/1"],
+ $build->@*,
+ ]],
+ 'stale',
+ ],
+ );
+ for my $plan (@plans) {
+ my ($name, $namespaces, $units, $state) = $plan->@*;
+ is_deeply(
+ [nv('_nvmet_plan_export', $case->($namespaces), $NQN, 1, $U{1}, $dev)],
+ [$units, $state],
+ "plan for a namespace that is $name",
+ );
+ }
+ my $incomplete = $case->({ 1 => ns_model($U{9}, '(null)', '0') });
+ is(
+ nv('_nvmet_export_state', $incomplete, $NQN, 1, $U{1}, $dev),
+ 'incomplete',
+ 'a disabled namespace with another identity is incomplete',
+ );
+ is(
+ nv('_nvmet_export_state', cfs_model(no_subsystem => 1), $NQN, 1, $U{1}, $dev),
+ 'nosubsys',
+ 'no subsystem',
+ );
+ for my $refusal (
+ [
+ 'enabled with another UUID',
+ { 1 => ns_model($U{9}, $dev) },
+ qr/NVMe namespace ID '1' has a different identity/,
+ ],
+ [
+ 'enabled with another device',
+ { 1 => ns_model($U{1}, '/dev/zvol/tank/other') },
+ qr/different identity/,
+ ],
+ [
+ 'enabled with another UUID on the template name',
+ { 1 => ns_model($U{9}, '/dev/zvol/tank/base-100-disk-0') },
+ qr/different identity/,
+ ],
+ [
+ 'UUID in use at another NSID',
+ { 2 => ns_model($U{1}, '/dev/zvol/tank/other') },
+ qr/namespace UUID '\Q@{[ $U{1} ]}\E' is already in use/,
+ ],
+ ) {
+ my ($name, $namespaces, $error) = $refusal->@*;
+ eval { nv('_nvmet_plan_export', $case->($namespaces), $NQN, 1, $U{1}, $dev) };
+ like($@, $error, "refuses a namespace $name");
+ }
+ eval { nv('_nvmet_plan_export', cfs_model(no_subsystem => 1), $NQN, 1, $U{1}, $dev) };
+ like($@, qr/NVMe subsystem does not exist/, 'refuses without subsystem');
+};
+
+subtest 'unexport planner' => sub {
+ my $dev = '/dev/zvol/tank/vm-100-disk-0';
+ is_deeply(
+ [nv('_nvmet_plan_unexport', cfs_model(), $NQN, $U{1})],
+ [[], undef],
+ 'nothing to remove',
+ );
+ is_deeply(
+ [nv('_nvmet_plan_unexport', cfs_model(no_subsystem => 1), $NQN, $U{1})],
+ [[], undef],
+ 'no subsystem',
+ );
+ my ($units, $pre) = nv(
+ '_nvmet_plan_unexport',
+ cfs_model(subsystem => { namespaces => { 4 => ns_model($U{1}, $dev) } }),
+ $NQN,
+ $U{1},
+ );
+ is_deeply(
+ $units,
+ [[write_step("$S/namespaces/4/enable", 0), ['rmdir', "$S/namespaces/4"]]],
+ 'an enabled namespace is disabled, then removed',
+ );
+ ok($pre->{nsid} == 4 && $pre->{enabled}, 'and was enabled');
+ ($units, $pre) = nv(
+ '_nvmet_plan_unexport',
+ cfs_model(subsystem => { namespaces => { 4 => ns_model($U{1}, $dev, '0') } }),
+ $NQN,
+ $U{1},
+ );
+ is_deeply($units, [[['rmdir', "$S/namespaces/4"]]], 'a disabled namespace is only removed');
+ ok($pre->{nsid} == 4 && !$pre->{enabled}, 'and was disabled');
+ eval {
+ nv(
+ '_nvmet_plan_unexport',
+ cfs_model(
+ subsystem => {
+ namespaces =>
+ { 4 => ns_model($U{1}, $dev), 5 => ns_model($U{1}, $dev, '0') },
+ },
+ ),
+ $NQN,
+ $U{1},
+ );
+ };
+ like($@, qr/duplicate namespace UUID/, 'refuses a duplicate UUID');
+};
+
+subtest 'port lookup' => sub {
+ my $port = sub($family, $address, $service, %links) {
+ return {
+ addr_trtype => 'tcp',
+ addr_adrfam => $family,
+ addr_traddr => $address,
+ addr_trsvcid => $service,
+ links => {%links},
+ };
+ };
+ my $cfs = {
+ ports => {
+ 3 => $port->('ipv4', '192.0.2.21', '4420'),
+ 5 => $port->('ipv4', '192.0.2.21', '4420', $NQN => 1),
+ 6 => $port->('ipv4', '192.0.2.22', '4420'),
+ 7 => $port->('ipv6', '2001:db8:0:0::11', '4420'),
+ 8 => { $port->('ipv4', '192.0.2.23', '4420')->%*, addr_trtype => '' },
+ 9 => { $port->('ipv4', '', '4420')->%* },
+ },
+ };
+ my $find =
+ sub($nqn, @portal) { return nv('_nvmet_find_port', $cfs, $nqn, portal(@portal)) };
+ is($find->($NQN, 'ipv4', '192.0.2.21', 4420), 5, 'the linked port wins');
+ is($find->($FOREIGN_NQN, 'ipv4', '192.0.2.21', 4420), 3, 'else the lowest');
+ is($find->($NQN, 'ipv4', '192.0.2.22', 4420), 6, 'exact match');
+ ok(!defined($find->($NQN, 'ipv6', '192.0.2.22', 4420)), 'wrong family');
+ ok(!defined($find->($NQN, 'ipv4', '192.0.2.22', 4421)), 'wrong service');
+ ok(!defined($find->($NQN, 'ipv4', '192.0.2.23', 4420)), 'incomplete port');
+ is($find->($NQN, 'ipv6', '2001:db8::11', 4420), 7, 'IPv6 spellings match');
+};
+
+subtest 'activation planner' => sub {
+ my $conf = { nqn => $NQN, portals => $PORTALS, hostnqns => [@HOSTS], keysha => $KEY_SHA };
+ my $inv = inventory(@POOL, @ONE);
+ my $empty = { subsystems => {}, ports => {}, hosts => {} };
+ my $dev = '/dev/zvol/tank/vm-100-disk-0';
+ my $serial = nv('_nvmet_serial', $NQN);
+
+ # A creation interrupted after the model write is completed while the
+ # subsystem is still unused ('an interrupted subsystem creation' below).
+ my $unfinished = sub(%attr) {
+ return {
+ $empty->%*,
+ subsystems => {
+ $NQN => {
+ attr_serial => '0f1e2d3c4b5a6978',
+ attr_allow_any_host => '0',
+ acl => {},
+ namespaces => {},
+ %attr,
+ },
+ },
+ };
+ };
+ is_deeply(
+ nv(
+ '_nvmet_plan_activation',
+ $inv,
+ $unfinished->(attr_model => $MODEL, attr_allow_any_host => '1'),
+ $conf,
+ )->{prepublish}->[0],
+ [write_step("$S/attr_serial", $serial), write_step("$S/attr_allow_any_host", 0)],
+ 'an unused subsystem with our model is completed and denies unknown hosts',
+ );
+ for my $case (
+ [
+ 'a namespace',
+ sub($c) { $c->{subsystems}->{$NQN}->{namespaces}->{1} = ns_model($U{1}, $dev) },
+ ],
+ ['an ACL', sub($c) { $c->{subsystems}->{$NQN}->{acl}->{ $HOSTS[0] } = 1 }],
+ [
+ 'a port link',
+ sub($c) {
+ $c->{ports}->{7} = {
+ addr_trtype => 'tcp',
+ addr_adrfam => 'ipv4',
+ addr_traddr => '192.0.2.99',
+ addr_trsvcid => '4420',
+ links => { $NQN => 1 },
+ };
+ },
+ ],
+ ) {
+ my ($name, $change) = $case->@*;
+ my $used = $unfinished->(attr_model => $MODEL);
+ $change->($used);
+ eval { nv('_nvmet_plan_activation', $inv, $used, $conf) };
+ like(
+ $@,
+ qr/refusing to take over existing NVMe subsystem/,
+ "a subsystem with our model and another serial and $name is refused",
+ );
+ }
+ # the identities of every owned volume are validated before any unit
+ for my $case (
+ ['a duplicate NSID', [$NQN, 1, $U{2}], qr/duplicate NSID '1'/],
+ ['a duplicate UUID', [$NQN, 2, $U{1}], qr/duplicate namespace UUID/],
+ ['an invalid UUID', [$NQN, 2, 'bad'], qr/invalid namespace UUID/],
+ ) {
+ my ($name, $identity, $error) = $case->@*;
+ my $bad =
+ inventory(@POOL, @ONE, zfs_rows('tank/vm-101-disk-0', 'volume', $identity->@*));
+ eval { nv('_nvmet_desired_namespaces', $bad, $NQN) };
+ like($@, $error, "$name is refused");
+ }
+
+ my $state = sub(%o) {
+ my $cfs = cfs_model(
+ subsystem => {
+ acl => { map { $_ => 1 } @HOSTS },
+ namespaces => { 1 => ns_model($U{1}, $dev) },
+ ($o{subsystem} // {})->%*,
+ },
+ );
+ $cfs->{ports} = {
+ map {
+ $_ => {
+ addr_trtype => 'tcp',
+ addr_adrfam => 'ipv4',
+ addr_traddr => "192.0.2.2$_",
+ addr_trsvcid => '4420',
+ links => { $NQN => 1 },
+ }
+ } 1,
+ 2,
+ };
+ $cfs->{hosts} = { map { $_ => { key_sha256 => $KEY_SHA } } @HOSTS };
+ $o{change}->($cfs) if $o{change};
+ return $cfs;
+ };
+ is_deeply(
+ nv('_nvmet_plan_activation', $inv, $state->(), $conf),
+ { prepublish => [], port_ids => [1, 2] },
+ 'a converged target plans nothing',
+ );
+ is_deeply(
+ nv('_nvmet_plan_publish', $state->(), $NQN, [1, 2], [@HOSTS]),
+ [],
+ 'nor publishes',
+ );
+ is_deeply(
+ nv(
+ '_nvmet_plan_activation',
+ $inv,
+ $state->(subsystem => { attr_allow_any_host => '1' }),
+ $conf,
+ )->{prepublish},
+ [[write_step("$S/attr_allow_any_host", 0)]],
+ 'allow_any_host is reset',
+ );
+
+ # the key matrix for the second host
+ my $h = $HOSTS[1];
+ my $H = "$ROOT/hosts/$h";
+ my $other = sha256_hex("$KEY_B\n");
+ my $foreign_acl = sub($cfs) {
+ $cfs->{subsystems}->{$FOREIGN_NQN} = { acl => { $h => 1 }, namespaces => {} };
+ };
+ for my $case (
+ [
+ 'absent',
+ sub($c) { delete $c->{hosts}->{$h} },
+ [[['mkdir', $H]], [chmod_step($H), { key => ["$H/dhchap_key"] }]],
+ ],
+ ['matching', sub($c) { }, []],
+ [
+ 'different and unlinked',
+ sub($c) {
+ $c->{hosts}->{$h}->{key_sha256} = $other;
+ delete $c->{subsystems}->{$NQN}->{acl}->{$h};
+ },
+ [
+ [chmod_step($H), { key => ["$H/dhchap_key"] }],
+ [['ln', '-s', $H, "$S/allowed_hosts/$h"]],
+ ],
+ ],
+ [
+ 'empty and unlinked',
+ sub($c) {
+ $c->{hosts}->{$h}->{key_sha256} = $EMPTY_KEY_SHA;
+ delete $c->{subsystems}->{$NQN}->{acl}->{$h};
+ },
+ [
+ [chmod_step($H), { key => ["$H/dhchap_key"] }],
+ [['ln', '-s', $H, "$S/allowed_hosts/$h"]],
+ ],
+ ],
+ ) {
+ my ($name, $change, $units) = $case->@*;
+ my $plan = nv('_nvmet_plan_activation', $inv, $state->(change => $change), $conf);
+ is_deeply($plan->{prepublish}, $units, "key of a host that is $name");
+ }
+ for my $case (
+ [
+ 'different and linked by us',
+ sub($c) { $c->{hosts}->{$h}->{key_sha256} = $other },
+ qr/refusing to replace an in-use DH-HMAC-CHAP key/,
+ ],
+ [
+ 'different and linked elsewhere',
+ sub($c) {
+ $c->{hosts}->{$h}->{key_sha256} = $other;
+ delete $c->{subsystems}->{$NQN}->{acl}->{$h};
+ $foreign_acl->($c);
+ },
+ qr/in-use DH-HMAC-CHAP key/,
+ ],
+ [
+ 'empty and linked',
+ sub($c) { $c->{hosts}->{$h}->{key_sha256} = $EMPTY_KEY_SHA },
+ qr/in-use DH-HMAC-CHAP key/,
+ ],
+ [
+ 'without a digest',
+ sub($c) { $c->{hosts}->{$h}->{key_sha256} = undef },
+ qr/does not expose a DH-HMAC-CHAP key for host/,
+ ],
+ ) {
+ my ($name, $change, $error) = $case->@*;
+ eval { nv('_nvmet_plan_activation', $inv, $state->(change => $change), $conf) };
+ like($@, $error, "refuses the key of a host that is $name");
+ }
+
+ my $undesired = $state->(
+ change => sub($c) {
+ $c->{subsystems}->{$NQN}->{namespaces}->{7} =
+ ns_model($U{7}, '/dev/zvol/tank/gone');
+ $c->{subsystems}->{$NQN}->{namespaces}->{8} = ns_model($U{8}, '(null)', '0');
+ },
+ );
+ is_deeply(
+ nv('_nvmet_plan_activation', $inv, $undesired, $conf)->{prepublish},
+ [
+ [write_step("$S/namespaces/7/enable", 0), ['rmdir', "$S/namespaces/7"]],
+ [['rmdir', "$S/namespaces/8"]],
+ ],
+ 'undesired namespaces are disabled and removed',
+ );
+};
+
+subtest 'publish planner' => sub {
+ my $cfs = cfs_model(subsystem => { acl => { map { $_ => 1 } @HOSTS } });
+ $cfs->{ports} = {
+ 1 => { links => { $NQN => 1 } },
+ 2 => { links => {} },
+ 5 => { links => { $NQN => 1, $FOREIGN_NQN => 1 } },
+ 6 => { links => { $FOREIGN_NQN => 1 } },
+ };
+ my @guards = (
+ ['grep', '-qx', '0', "$S/attr_allow_any_host"],
+ map { ['test', '-L', "$S/allowed_hosts/$_"] } @HOSTS,
+ );
+ is_deeply(
+ nv('_nvmet_plan_publish', $cfs, $NQN, [1, 2], [@HOSTS]),
+ [
+ {
+ port => 2,
+ action => 'link',
+ steps => [@guards, ['ln', '-s', $S, "$ROOT/ports/2/subsystems/$NQN"]],
+ },
+ {
+ port => 5,
+ action => 'unlink',
+ steps => [['rm', "$ROOT/ports/5/subsystems/$NQN"]],
+ },
+ ],
+ 'links a missing port behind the guards and unlinks a stray port',
+ );
+ $cfs->{ports}->{2}->{links}->{$NQN} = 1;
+ delete $cfs->{ports}->{5};
+ is_deeply(nv('_nvmet_plan_publish', $cfs, $NQN, [1, 2], [@HOSTS]), [], 'converged');
+};
+
+subtest 'target removal planners' => sub {
+ my $conf_hosts =
+ [@HOSTS, 'nqn.2014-08.org.nvmexpress:uuid:00000000-0000-4000-8000-000000000003'];
+ my $cfs = cfs_model(subsystem => { acl => { map { $_ => 1 } @HOSTS } });
+ $cfs->{ports} = {
+ 2 => { links => { $NQN => 1 } },
+ 1 => { links => { $NQN => 1 } },
+ 3 => { links => {} },
+ };
+ is_deeply(
+ [nv('_nvmet_plan_delete_target', inventory(@POOL), $cfs, $NQN, $conf_hosts)],
+ [
+ [[
+ [
+ 'rm',
+ "$ROOT/ports/1/subsystems/$NQN",
+ "$ROOT/ports/2/subsystems/$NQN",
+ map { "$S/allowed_hosts/$_" } @HOSTS,
+ ],
+ ['rmdir', $S],
+ ]],
+ [sort $conf_hosts->@*],
+ ],
+ 'teardown removes port links first, then ACLs, then the subsystem',
+ );
+ # The candidates are the ACL and the configured hosts: a retry after a
+ # teardown that removed the ACLs but not the subsystem still finds them.
+ my $stale_host = 'nqn.2026-01.com.example:stale';
+ my $stale = cfs_model(subsystem => { acl => { $stale_host => 1 } });
+ is_deeply(
+ (nv('_nvmet_plan_delete_target', inventory(@POOL), $stale, $NQN, [@HOSTS]))[1],
+ [sort $stale_host, @HOSTS],
+ 'a host that is only in the ACL is a candidate too',
+ );
+ is_deeply(
+ [nv('_nvmet_plan_delete_target', inventory(@POOL), cfs_model(), $NQN, [@HOSTS])],
+ [[[['rmdir', $S]]], [sort @HOSTS]],
+ 'without ACLs the configured hosts are candidates',
+ );
+
+ my $orphans = {
+ subsystems => { $FOREIGN_NQN => { acl => { $HOSTS[0] => 1 } } },
+ hosts => { map { $_ => {} } @HOSTS },
+ };
+ is_deeply(
+ nv('_nvmet_plan_orphan_hosts', $orphans, [@HOSTS, @HOSTS, 'nqn.x:gone']),
+ [[['rmdir', "$ROOT/hosts/$HOSTS[1]"]]],
+ 'orphans exclude linked and missing hosts',
+ );
+ is_deeply(nv('_nvmet_plan_orphan_hosts', $orphans, [$HOSTS[0]]), [], 'no orphans');
+ $orphans->{hosts}->{'..'} = {};
+ is_deeply(
+ nv('_nvmet_plan_orphan_hosts', $orphans, ['..']),
+ [],
+ 'a host name read from the target is only used when it is an NQN',
+ );
+ my $odd = cfs_model(subsystem => { acl => { $HOSTS[0] => 1, '..' => 1 } });
+ eval { nv('_nvmet_plan_delete_target', inventory(@POOL), $odd, $NQN, [@HOSTS]) };
+ like($@, qr/it allows a malformed host name/, 'an ACL with a malformed name is refused');
+};
+
+subtest 'template name and chunking' => sub {
+ is(nv('_nvmet_template_name', 'tank/vm-100-disk-0'), 'tank/base-100-disk-0', 'vm to base');
+ is(nv('_nvmet_template_name', 'tank/sub/vm-1-disk-2'), 'tank/sub/base-1-disk-2', 'nested');
+ for my $name ('tank/base-100-disk-0', 'tank/subvol-100-disk-0', 'tank/xvm-1-disk-0') {
+ eval { nv('_nvmet_template_name', $name) };
+ like($@, qr/only VM zvols can become templates/, "refuses $name");
+ }
+
+ my @units = map { build_steps($_, $U{1}, "/dev/zvol/tank/vm-$_-disk-0") } 1 .. 400;
+ my $chunks = nv('_nvmet_chunk', \@units);
+ ok($chunks->@* > 1, 'large plans are split');
+ ok(
+ !grep({ length(nv('_nvmet_render', $_)) > 65536 } $chunks->@*),
+ 'no call exceeds 64 KiB',
+ );
+ ok(!grep({ $_->[0]->[0] ne 'test' || $_->@* % 6 } $chunks->@*), 'chunks hold whole units');
+ is_deeply([map { $_->@* } $chunks->@*], [map { $_->@* } @units], 'order is preserved');
+ is_deeply(nv('_nvmet_chunk', []), [], 'no units, no calls');
+ my $small = nv(
+ '_nvmet_chunk',
+ [@units[0 .. 3]],
+ length(nv('_nvmet_render', [map { $_->@* } @units[0 .. 1]])),
+ );
+ is(scalar($small->@*), 2, 'the limit is configurable');
+ eval { nv('_nvmet_chunk', [$units[0]], 10) };
+ like($@, qr/NVMe target command too long/, 'a unit over the limit is refused');
+ my $key_unit = [{ key => ["$ROOT/hosts/$HOSTS[0]/dhchap_key"] }];
+ my $mixed = nv('_nvmet_chunk', [@units[0 .. 199], $key_unit, @units[200 .. 399]]);
+ is(
+ scalar(
+ grep {
+ grep { ref($_) eq 'HASH' && $_->{key} } $_->@*
+ } $mixed->@*
+ ),
+ 1,
+ 'the key step is in exactly one call',
+ );
+};
+
+# ---------------------------------------------------------------------------
+# Flows against the fake target
+# ---------------------------------------------------------------------------
+
+sub secret_file($storeid) {
+ return "/etc/pve/priv/storage/$storeid.nvme-dhchap";
+}
+
+# The target part of a storage activation, with the key stored for the storage.
+sub activate($scfg = scfg(), $key = $KEY, $hostnqns = [@HOSTS], $portals = $PORTALS) {
+ local $FILES{ secret_file('st') } = $key;
+ return nv('_nvmet_activate_target', 'st', $scfg, $portals, $hostnqns, $key);
+}
+
+# The public operations of the flows, on lifecycle_fake().
+my %ACT = (
+ activate => sub { activate() },
+ alloc => sub { $PLUGIN->alloc_image('st', scfg(), 200, 'raw', undef, 1024) },
+ clone => sub { $PLUGIN->clone_image(scfg(), 'st', 'base-102-disk-0', 201) },
+ free => sub { $PLUGIN->free_image('st', scfg(), 'vm-101-disk-0') },
+ rollback =>
+ sub { $PLUGIN->volume_snapshot_rollback(scfg(), 'st', 'vm-100-disk-0', 'snap1') },
+ template => sub { $PLUGIN->create_base('st', scfg(), 'vm-101-disk-0') },
+ resize => sub { $PLUGIN->volume_resize(scfg(), 'st', 'vm-100-disk-0', 2 * 1024**3) },
+ remove => sub { nv('_nvmet_delete_target', scfg()) },
+);
+
+# Runs an action against a new lifecycle_fake(%opts).
+sub run_on($action, %opts) {
+ my $fake = lifecycle_fake(%opts);
+ return ($fake, flow($fake, $ACT{$action}));
+}
+
+# An action (a name of %ACT or a sub) that is refused without changing the target.
+sub refused($fake, $action, $error, $name) {
+ my $before = dclone($fake->{m});
+ my $res = flow($fake, ref($action) ? $action : $ACT{$action});
+ like($res->{error}, $error, "refuses $name");
+ ok(!changes($res) && same($before, $fake->{m}), "without changes for $name");
+ return $res;
+}
+
+sub ns_of($fake, $nsid, $nqn = $NQN) {
+ return ($fake->{m}->{subsystems}->{$nqn} // {})->{ns}->{$nsid};
+}
+
+sub exported($fake, $nsid, $uuid, $dev) {
+ my $ns = ns_of($fake, $nsid) // return 0;
+ return $ns->{enable} eq '1' && $ns->{device_uuid} eq $uuid && $ns->{device_path} eq $dev;
+}
+
+sub step_is($step, @prefix) {
+ return
+ ref($step) eq 'ARRAY'
+ && $step->@* >= @prefix
+ && join("\0", $step->@[0 .. $#prefix]) eq join("\0", @prefix);
+}
+
+sub write_to($step, $suffix, $value = undef) {
+ return
+ ref($step) eq 'HASH'
+ && exists($step->{write})
+ && $step->{write} =~ m{\Q$suffix\E\z}
+ && (!defined($value) || $step->{value} eq $value);
+}
+
+subtest 'path() is one read-only ZFS query' => sub {
+ my $fake = lifecycle_fake();
+ my $res = flow($fake, sub { [$PLUGIN->path(scfg(), 'vm-100-disk-0', 'st')] });
+ is($res->{error}, '', 'lookup succeeds');
+ is_deeply(
+ $res->{result},
+ ["/dev/disk/by-id/nvme-uuid.$U{1}", 100, 'images'],
+ 'the UUID link',
+ );
+ is_deeply(ops($res), ['read ZFS inventory'], 'one call');
+ is(scalar($res->{locks}->@*), 0, 'without the lock');
+ $res = flow($fake, sub { [$PLUGIN->path(scfg(), 'vm-100-disk-0', 'st', 'snap1')] });
+ like($res->{error}, qr/direct access to snapshots not implemented/, 'snapshots');
+ is(scalar($res->{calls}->@*), 0, 'are refused locally');
+};
+
+subtest 'a failed read is retried' => sub {
+ local $SLEPT = 0;
+ my $fake = lifecycle_fake(read_fault => { after => -1, mode => 'before' });
+ my $res = flow($fake, sub { [$PLUGIN->path(scfg(), 'vm-100-disk-0', 'st')] });
+ is($res->{error}, '', 'a read that fails once');
+ is_deeply(ops($res), ['read ZFS inventory', 'read ZFS inventory'], 'is repeated');
+ is($SLEPT, 0.2, 'after 200 ms');
+
+ $fake = lifecycle_fake(fail => [['zfs', 'get'], 'I/O error']);
+ $res = flow($fake, sub { [$PLUGIN->path(scfg(), 'vm-100-disk-0', 'st')] });
+ is($res->{error}, "cannot read NVMe target state: I/O error\n",
+ 'a read that keeps failing');
+ is(scalar($res->{calls}->@*), 3, 'is tried three times');
+
+ $fake = lifecycle_fake(read_fault => { after => -1, mode => 'unreachable' });
+ $res = flow($fake, sub { [$PLUGIN->path(scfg(), 'vm-100-disk-0', 'st')] });
+ like($res->{error}, qr/NVMe target '192\.0\.2\.10' is unreachable/,
+ 'an unreachable target');
+ is(scalar($res->{calls}->@*), 1, 'is not tried again');
+};
+
+subtest 'activate_volume' => sub {
+ my %forced;
+ my $activations = 0;
+ my $activation = sub($class, $storeid, $scfg, $cache = undef) {
+ $activations++;
+ $forced{$storeid} = $cache->{'zfsnvme-force-reconcile'}->{$storeid};
+ die "activation attempted\n";
+ };
+ $plugin_mock->redefine(activate_storage => $activation);
+
+ my $fake = lifecycle_fake();
+ delete $fake->{m}->{subsystems}->{$NQN}->{ns}->{2};
+ my $res = flow($fake, sub { $PLUGIN->activate_volume('st', scfg(), 'vm-101-disk-0') });
+ like($res->{error}, qr/activation attempted/, 'a missing namespace forces an activation');
+ is($forced{st}, 1, 'the activation is forced past its fast path');
+ is_deeply(ops($res), ['read target state'], 'after one read');
+ ok(!$res->{calls}->[0]->{mutating} && !$res->{locks}->@*, 'without a change or the lock');
+
+ $fake = lifecycle_fake();
+ ns_of($fake, 2)->{device_uuid} = $U{9};
+ $activations = 0;
+ $res = flow($fake, sub { $PLUGIN->activate_volume('st', scfg(), 'vm-101-disk-0') });
+ like(
+ $res->{error},
+ qr/NVMe namespace ID '2' has a different identity/,
+ 'refuses an enabled namespace with another identity',
+ );
+ is($activations, 0, 'without activating');
+ $res = flow($fake, sub { $PLUGIN->activate_volume('st', scfg(), 'vm-101-disk-0', 'snap') });
+ like($res->{error}, qr/unable to activate snapshot/, 'snapshots are refused');
+
+ # The local device: its udev link and whether it is the namespace.
+ my $ok = 1;
+ $plugin_mock->redefine(_nvmet_local_namespace_ok => sub(@args) { return $ok });
+ my $link = "/dev/disk/by-id/nvme-uuid.$U{1}";
+ local %BLOCK = ($link => 1);
+ $fake = target_fake();
+ add_volume($fake, 'vm-100-disk-0', 1, $U{1});
+ $activations = 0;
+ $res = flow($fake, sub { $PLUGIN->activate_volume('st', scfg(), 'vm-100-disk-0') });
+ is($res->{error}, '', 'a converged volume activates');
+ is_deeply(ops($res), ['read target state'], 'with one read');
+ is(scalar($res->{locks}->@*) + $activations, 0, 'without lock or activation');
+
+ for my $case (
+ ['disabled', sub($f) { ns_of($f, 1)->{enable} = '0' }],
+ ['absent', sub($f) { delete $f->{m}->{subsystems}->{$NQN}->{ns}->{1} }],
+ ['without subsystem', sub($f) { delete $f->{m}->{subsystems}->{$NQN} }],
+ [
+ 'on the template name of the zvol',
+ sub($f) { ns_of($f, 1)->{device_path} = '/dev/zvol/tank/base-100-disk-0' },
+ ],
+ [
+ 'being built',
+ sub($f) {
+ $f->{m}->{subsystems}->{$NQN}->{ns}->{1} = ns_model($U{9}, '(null)', '0');
+ },
+ ],
+ ) {
+ my ($name, $change) = $case->@*;
+ $fake = target_fake();
+ add_volume($fake, 'vm-100-disk-0', 1, $U{1});
+ $change->($fake);
+ %forced = ();
+ $res = flow($fake, sub { $PLUGIN->activate_volume('st', scfg(), 'vm-100-disk-0') });
+ is($forced{st}, 1, "a namespace that is $name on the target is exported again");
+ }
+
+ $fake = target_fake();
+ add_volume($fake, 'vm-100-disk-0', 1, $U{1});
+ $ok = 0;
+ $activations = 0;
+ $plugin_mock->redefine(activate_storage => sub(@args) { $activations++; return 1 });
+ $res = flow($fake, sub { $PLUGIN->activate_volume('st', scfg(), 'vm-100-disk-0') });
+ like(
+ $res->{error},
+ qr/NVMe namespace for 'vm-100-disk-0' has an unexpected identity/,
+ 'a local device of another namespace is never used',
+ );
+ is($activations, 1, 'after one forced activation');
+ is_deeply(ops($res), ['read target state'], 'and one read');
+ $plugin_mock->unmock('_nvmet_local_namespace_ok');
+
+ # While the device is missing, live controllers are rescanned.
+ my $file_mock = Test::MockModule->new('PVE::File');
+ $file_mock->redefine(
+ dir_glob_foreach => sub($dir, $regex, $func) {
+ return if $dir ne '/sys/class/nvme';
+ $func->($_) for qw(nvme7 nvme8 nvme9);
+ },
+ );
+ local %FILES = (
+ '/sys/class/nvme/nvme7/subsysnqn' => $NQN,
+ '/sys/class/nvme/nvme7/state' => 'live',
+ '/sys/class/nvme/nvme8/subsysnqn' => $NQN,
+ '/sys/class/nvme/nvme8/state' => 'connecting',
+ '/sys/class/nvme/nvme9/subsysnqn' => $FOREIGN_NQN,
+ '/sys/class/nvme/nvme9/state' => 'live',
+ );
+ $plugin_mock->redefine(activate_storage => sub(@args) { return 1 });
+ $fake = lifecycle_fake();
+ delete $fake->{m}->{subsystems}->{$NQN}->{ns}->{2};
+ my $start = $NOW;
+ $res = flow($fake, sub { $PLUGIN->activate_volume('st', scfg(), 'vm-101-disk-0') });
+ like(
+ $res->{error},
+ qr/NVMe namespace for 'vm-101-disk-0' did not appear/,
+ 'a missing device',
+ );
+ is($NOW - $start, 10, 'after 10 seconds');
+ my $rescan = ['/sys/class/nvme/nvme7/rescan_controller', "1\n"];
+ is_deeply(
+ $res->{sysfs_writes},
+ [($rescan) x 5],
+ 'the live controller is rescanned every 2 s',
+ );
+ is_deeply($res->{warnings}, [], 'without a warning');
+
+ # The stock file_write on files that fail: a rescan that fails is a task
+ # warning with the errno text, and a controller that went away since it
+ # was listed needs none.
+ my $file_write = $sysfs_mock->original('file_write');
+ my $dir = tempdir(CLEANUP => 1);
+ my @perl_warnings;
+ local $SIG{__WARN__} = sub($warning) { push @perl_warnings, $warning };
+ for my $case (
+ ['/dev/full', 'No space left on device'],
+ [$dir, 'Is a directory'],
+ ["$dir/gone/rescan_controller", undef],
+ ) {
+ my ($path, $reason) = $case->@*;
+ $sysfs_mock->redefine(file_write => sub($file, @args) { $file_write->($path, @args) });
+ $res = flow($fake, sub { $PLUGIN->activate_volume('st', scfg(), 'vm-101-disk-0') });
+ is_deeply(
+ $res->{warnings},
+ defined($reason) ? [("cannot rescan NVMe controller 'nvme7': $reason") x 5] : [],
+ defined($reason)
+ ? "a failed rescan is a task warning with the errno text ($reason)"
+ : 'a controller that went away is not',
+ );
+ }
+ is_deeply(\@perl_warnings, [], 'and file_write warns about none');
+ $sysfs_mock->redefine(file_write => $sysfs_write);
+ $file_mock->unmock_all();
+ $plugin_mock->unmock('activate_storage');
+};
+
+subtest 'activation of an empty target' => sub {
+ my $fake = FakeTarget->new(pools => ['tank']);
+ $fake->add_zvol('tank/vm-100-disk-0', identity => [$NQN, 1, $U{1}]);
+ my $res = flow($fake, $ACT{activate});
+ is($res->{error}, '', 'the target is configured');
+ is_deeply(
+ [map { $_->{locked} } $res->{calls}->@*],
+ [0, 1, 1, 1, 1, 1],
+ 'read, then lock, read, apply, verify and publish',
+ );
+ is_deeply(
+ [map { [$_->@{qw(name timeout)}] } $res->{locks}->@*],
+ [['zfsnvme-192.0.2.10', 30]],
+ 'under the domain lock of the target, waiting up to 30 seconds for it',
+ );
+ my $m = $fake->{m};
+ my $subsys = $m->{subsystems}->{$NQN};
+ is_deeply(
+ $subsys->{attr},
+ {
+ attr_model => $MODEL,
+ attr_serial => nv('_nvmet_serial', $NQN),
+ attr_allow_any_host => '0',
+ },
+ 'the subsystem carries our marker',
+ );
+ is_deeply($subsys->{acl}, { map { $_ => 1 } @HOSTS }, 'every host has an ACL');
+ is_deeply([map { $m->{hosts}->{$_}->{key} } @HOSTS], [$KEY, $KEY], 'and the key');
+ ok(exported($fake, 1, $U{1}, '/dev/zvol/tank/vm-100-disk-0'), 'the volume is exported');
+
+ $res = flow($fake, $ACT{activate});
+ is_deeply(ops($res), ['read target state'], 'a converged target costs one read');
+ ok(!$res->{locks}->@* && !$res->{calls}->[0]->{mutating}, 'without lock or change');
+};
+
+subtest 'activation refusals change nothing' => sub {
+ for my $case (
+ [
+ 'a foreign subsystem at our NQN',
+ sub($f) { $f->add_subsystem($NQN, model => 'Other Storage', serial => 'abc') },
+ qr/refusing to take over existing NVMe subsystem/,
+ ],
+ [
+ 'a used subsystem with the default model',
+ sub($f) {
+ $f->add_host($HOSTS[0], $KEY);
+ $f->add_subsystem($NQN, model => 'Linux', serial => 'abc', acl => [$HOSTS[0]]);
+ },
+ qr/refusing to take over existing NVMe subsystem/,
+ ],
+ [
+ 'a duplicate UUID on a published target',
+ sub($f) {
+ $f->{m} = lifecycle_fake()->{m};
+ $f->add_zvol('tank/vm-150-disk-0', identity => [$NQN, 7, $U{5}]);
+ $f->add_zvol('tank/vm-151-disk-0', identity => [$NQN, 8, $U{5}]);
+ },
+ qr/duplicate namespace UUID '\Q@{[ $U{5} ]}\E'/,
+ ],
+ [
+ 'an enabled namespace of another volume',
+ sub($f) {
+ $f->add_zvol('tank/vm-100-disk-0', identity => [$NQN, 1, $U{1}]);
+ $f->add_subsystem($NQN);
+ $f->add_namespace($NQN, 1, $U{9}, '/dev/zvol/tank/other');
+ },
+ qr/NVMe namespace ID '1' has a different identity/,
+ ],
+ ) {
+ my ($name, $setup, $error) = $case->@*;
+ my $fake = FakeTarget->new(pools => ['tank']);
+ $setup->($fake);
+ my $res = refused($fake, 'activate', $error, $name);
+ unlike($res->{error}, qr/DHHC-1/, 'without key material in the message');
+ }
+};
+
+subtest 'nothing is published unless a fresh read converges' => sub {
+ my $acl = "$S/allowed_hosts/$HOSTS[1]";
+ my $failing = sub($limit) {
+ my $failures = 0;
+ return sub($step, $call) {
+ return undef
+ if !step_is($step, 'ln', '-s') || $step->[3] ne $acl || $failures >= $limit;
+ $failures++;
+ return "ln: failed to create symbolic link '$acl': Invalid argument";
+ };
+ };
+ my $fake = FakeTarget->new(pools => ['tank'], step_fault => $failing->(99));
+ my $res = flow($fake, $ACT{activate});
+ like(
+ $res->{error},
+ qr/NVMe target did not converge: ln: failed to create symbolic link/,
+ 'a plan that fails twice stops',
+ );
+ is(
+ scalar(grep { $_->{op} eq 'configure NVMe target' } $res->{calls}->@*),
+ 2,
+ 'in two rounds',
+ );
+ is(scalar(grep { $_->{op} =~ /\Alink / } $res->{calls}->@*), 0, 'and publishes nothing');
+ ok(!grep({ $_->{links}->%* } values $fake->{m}->{ports}->%*),
+ 'no port links the subsystem');
+
+ $fake = FakeTarget->new(pools => ['tank'], step_fault => $failing->(1));
+ $res = flow($fake, $ACT{activate});
+ is($res->{error}, '', 'a transient failure is repaired by the second round');
+ is(scalar(grep { $_->{op} =~ /\Alink / } $res->{calls}->@*), 2, 'which then publishes');
+
+ $fake = FakeTarget->new(pools => ['tank']);
+ $fake->{m}->{mounted} = 0;
+ $res = flow($fake, $ACT{activate});
+ is($res->{error}, '', 'configfs is mounted when it is missing');
+ my @seen = grep { /mount/ } ops($res)->@*;
+ is_deeply(\@seen, ['read mounts', 'mount configfs'], 'after reading the mount table');
+ is(
+ $res->{calls}->[-1]->{op},
+ 'link NVMe subsystem on port 2',
+ 'and the target is published',
+ );
+
+ $fake = FakeTarget->new(pools => ['tank']);
+ $fake->{step_fault} = sub($step, $call) {
+ return step_is($step, 'env') ? "find: '$ROOT/hosts': Permission denied" : undef;
+ };
+ $res = flow($fake, $ACT{activate});
+ like(
+ $res->{error},
+ qr/cannot read NVMe target state: find: '\Q$ROOT\E\/hosts': Permission denied/,
+ 'an unreadable, mounted configfs stops the activation',
+ );
+ ok(!grep({ $_ eq 'mount configfs' } ops($res)->@*), 'without mounting');
+ is(changes($res), 0, 'and without changes');
+};
+
+subtest 'an interrupted subsystem creation' => sub {
+ # The chain that creates the subsystem stops after its mkdir or after the
+ # model write, for example when the remote shell is killed. Call 2 is the
+ # configuration. Its outcome is unknown; the next activation repairs it.
+ my $marker = {
+ attr_model => $MODEL,
+ attr_serial => nv('_nvmet_serial', $NQN),
+ attr_allow_any_host => '0',
+ };
+ my $completed = sub($fake, $what) {
+ is_deeply($fake->{m}->{subsystems}->{$NQN}->{attr}, $marker, "$what: our marker");
+ ok(
+ exported($fake, 1, $U{1}, '/dev/zvol/tank/vm-100-disk-0'),
+ "$what: exports the volume",
+ );
+ ok(
+ $fake->{m}->{ports}->{1}->{links}->{$NQN}
+ && $fake->{m}->{ports}->{2}->{links}->{$NQN},
+ "$what: and publishes it",
+ );
+ is_deeply($fake->{violations}, [], "$what: no protocol violation");
+ };
+ my $new = sub(%faults) {
+ my $fake = FakeTarget->new(pools => ['tank'], faults => {%faults});
+ $fake->add_zvol('tank/vm-100-disk-0', identity => [$NQN, 1, $U{1}]);
+ return $fake;
+ };
+
+ # After the model write, only this plugin can have created the subsystem.
+ my $fake = $new->(2 => 'cut:2');
+ my $res = flow($fake, $ACT{activate});
+ like(
+ $res->{error},
+ qr/'configure NVMe target' did not complete .* state is unknown\n\z/,
+ 'a cut after the model write',
+ );
+ ok(!grep({ $_->{links}->%* } values $fake->{m}->{ports}->%*), 'is unpublished');
+ $fake->{faults} = {};
+ $res = flow($fake, $ACT{activate});
+ is($res->{error}, '', 'the next activation completes the subsystem');
+ my ($configure) = grep { $_->{op} eq 'configure NVMe target' } $res->{calls}->@*;
+ is_deeply(
+ $configure->{steps}->[0],
+ write_step("$S/attr_serial", $marker->{attr_serial}),
+ 'by writing the serial',
+ );
+ $completed->($fake, 'next activation after the model write');
+
+ # Right after the mkdir, the subsystem looks like one of another tool.
+ $fake = $new->(2 => 'cut:1');
+ $res = flow($fake, $ACT{activate});
+ like($res->{error}, qr/state is unknown\n\z/, 'a cut right after the mkdir');
+ $fake->{faults} = {};
+ $res = flow($fake, $ACT{activate});
+ like(
+ $res->{error},
+ qr/refusing to take over existing NVMe subsystem .* not created by Proxmox VE/,
+ 'is refused by the next activation',
+ );
+ is($fake->{m}->{subsystems}->{$NQN}->{attr}->{attr_model}, 'Linux',
+ 'the subsystem is kept');
+ ok(!grep({ $_->{links}->%* } values $fake->{m}->{ports}->%*), 'and unpublished');
+ delete $fake->{m}->{subsystems}->{$NQN};
+ $res = flow($fake, $ACT{activate});
+ is($res->{error}, '', 'once an administrator removed it, the next activation creates it');
+ $completed->($fake, 'activation after the removal');
+};
+
+subtest 'quorum and reachability' => sub {
+ my $fake = FakeTarget->new(pools => ['tank']);
+ local $QUORATE = 0;
+ local @QUORUM = ();
+ my $res = flow($fake, $ACT{activate});
+ like($res->{error}, qr/cluster not quorate - refusing NVMe target changes/, 'no quorum');
+ is_deeply(\@QUORUM, [1], 'with a quorum check that does not wait');
+ is(scalar($res->{locks}->@*), 0, 'before requesting the lock');
+ is_deeply(ops($res), ['read target state'], 'after the lockless read');
+ $QUORATE = 1;
+
+ $fake->{unreachable} = 1;
+ $res = flow($fake, $ACT{activate});
+ like(
+ $res->{error},
+ qr/NVMe target '192\.0\.2\.10' is unreachable: ssh: connect/,
+ 'an unreachable target fails in the lockless read',
+ );
+ ok(!$res->{locks}->@* && $res->{calls}->@* == 1, 'without lock or retry');
+
+ $fake = lifecycle_fake();
+ local $QUORATE = 0;
+ $res = flow($fake, $ACT{activate});
+ is($res->{error}, '', 'a converged target needs no quorum');
+};
+
+subtest 'create' => sub {
+ my ($fake, $res) = run_on('alloc');
+ is($res->{error}, '', 'allocation succeeds');
+ is($res->{result}, 'vm-200-disk-0', 'with the next free name');
+ is_deeply(
+ [map { $_->{locked} } $res->{calls}->@*],
+ [0, 1, 1, 1, 1, 1, 1],
+ 'the name lookup without the lock, the rest under it',
+ );
+ my $uuid = $fake->{m}->{ds}->{'tank/vm-200-disk-0'}->{props}->{'proxmox:nvme-uuid'};
+ like($uuid, $UUID_RE, 'a new UUID');
+ ok(exported($fake, 4, $uuid, '/dev/zvol/tank/vm-200-disk-0'), 'exported at the next NSID');
+
+ my $thick = scfg();
+ delete $thick->@{qw(sparse blocksize)};
+ $fake = lifecycle_fake();
+ $res =
+ flow($fake, sub { $PLUGIN->alloc_image('st', $thick, 200, 'raw', 'vm-200-disk-1', 4000) });
+ is($res->{result}, 'vm-200-disk-1', 'an explicit name is used');
+ my ($create) = grep { $_->{op} eq 'create zvol' } $res->{calls}->@*;
+ is_deeply(
+ [grep { /\A(?:-[sb]|[0-9]+k)\z/ } $create->{steps}->[0]->@*],
+ ['4096k'],
+ 'the size is rounded up to KiB; without sparse and blocksize the zvol is thick',
+ );
+
+ ($fake, $res) = run_on('alloc', udev_delay => 3);
+ is($res->{error}, '', 'a slow udev');
+ is(scalar(grep { $_ eq 'wait for zvol' } ops($res)->@*), 4, 'adds polls');
+
+ local @UUIDS = ($U{1}, $U{9});
+ ($fake, $res) = run_on('alloc');
+ is(
+ $fake->{m}->{ds}->{'tank/vm-200-disk-0'}->{props}->{'proxmox:nvme-uuid'},
+ $U{9},
+ 'a UUID that is in use is drawn again',
+ );
+ local @UUIDS = ($U{1}) x 10;
+ ($fake, $res) = run_on('alloc');
+ is($res->{error}, "cannot generate an unused namespace UUID\n", 'but not forever');
+ is(changes($res), 0, 'and nothing changes');
+ @UUIDS = ();
+
+ ($fake, $res) = run_on('alloc', faults => { 3 => 'after' });
+ like(
+ $res->{error},
+ qr/'create zvol' did not complete .*state is unknown\n\z/,
+ 'a lost reply aborts even when the create reached the target',
+ );
+ is($res->{calls}->[-1]->{op}, 'create zvol', 'no read tries to settle it');
+ ok(!ns_of($fake, 4), 'the abandoned operation does not export it');
+ $fake->{faults} = {};
+ $res = flow($fake, $ACT{activate});
+ ok(!$res->{error} && ns_of($fake, 4), 'the next activation exports the owned zvol');
+
+ ($fake, $res) = run_on('alloc', faults => { 5 => 'before' });
+ like($res->{error}, qr/'export namespace' failed: injected failure/, 'a failed export');
+ ok(!$fake->{m}->{ds}->{'tank/vm-200-disk-0'} && !ns_of($fake, 4), 'destroys the new zvol');
+ is($res->{locks}->@* + $res->{nested}->@*, 1, 'the compensation reuses the lock');
+
+ # The export may still run on the target after its connection broke.
+ ($fake, $res) = run_on('alloc', faults => { 5 => 'after' });
+ like(
+ $res->{error},
+ qr/\A.*'export namespace' did not complete \(Connection.*state is unknown\n\z/,
+ 'a lost reply of the export leaves the state unknown',
+ );
+ ok(!grep({ /destroy|remove/ } ops($res)->@*), 'and is not compensated');
+ ok($fake->{m}->{ds}->{'tank/vm-200-disk-0'} && ns_of($fake, 4), 'so the disk is kept');
+
+ ($fake, $res) = run_on(
+ 'alloc',
+ before_call => sub($f, $call) {
+ $f->add_zvol('tank/vm-777-disk-0', identity => [$NQN, 4, $U{9}])
+ if $call->{index} == 6 && !$f->{m}->{ds}->{'tank/vm-777-disk-0'};
+ },
+ );
+ like(
+ $res->{error},
+ qr/duplicate NVMe identity on 'tank\/vm-200-disk-0'/,
+ 'a duplicate that appears before the verify read is refused',
+ );
+ ok(
+ !$fake->{m}->{ds}->{'tank/vm-200-disk-0'} && !ns_of($fake, 4),
+ 'the new zvol is destroyed',
+ );
+ ok($fake->{m}->{ds}->{'tank/vm-777-disk-0'}, 'the other zvol is kept');
+
+ ($fake, $res) = run_on(
+ 'alloc',
+ before_call => sub($f, $call) {
+ my $ds = $f->{m}->{ds}->{'tank/vm-200-disk-0'};
+ $ds->{props}->{'proxmox:nvme-uuid'} = $U{9} if $ds && $call->{index} == 6;
+ },
+ );
+ is(
+ $res->{error},
+ "NVMe identity of zvol 'tank/vm-200-disk-0' changed\n",
+ 'an identity that changes before the verify read is refused',
+ );
+ ok(!ns_of($fake, 4), 'the new namespace is removed');
+ ok($fake->{m}->{ds}->{'tank/vm-200-disk-0'}, 'but a zvol without our identity is kept');
+
+ # A namespace of another zvol that was renamed away while exported uses
+ # the device path of the new zvol: only what this call built is removed.
+ ($fake, $res) = run_on(
+ 'alloc',
+ faults => { 5 => 'before' },
+ before_call => sub($f, $call) {
+ $f->add_namespace($NQN, 9, $U{9}, '/dev/zvol/tank/vm-200-disk-0')
+ if $call->{index} == 0;
+ },
+ );
+ like($res->{error}, qr/'export namespace' failed: injected failure/, 'a failed export');
+ ok(
+ exported($fake, 9, $U{9}, '/dev/zvol/tank/vm-200-disk-0'),
+ 'keeps an enabled namespace of another UUID with the same device path',
+ );
+ ok(
+ !grep({ grep { write_to($_, '/namespaces/9/enable') } $_->{steps}->@* }
+ $res->{calls}->@*),
+ 'without touching it',
+ );
+
+ ($fake, $res) = run_on(
+ 'alloc',
+ faults => { 5 => 'before' },
+ fail => [['zfs', 'destroy'], "cannot destroy 'tank/vm-200-disk-0': dataset is busy"],
+ );
+ like(
+ $res->{error},
+ qr/'export namespace' failed: injected failure/,
+ 'keeps the first error',
+ );
+ my $cleanup = qr/\Afailed to clean up zvol 'vm-200-disk-0': NVMe target operation/;
+ is(
+ scalar(grep { /$cleanup 'destroy zvol' failed: cannot destroy/ } $res->{warnings}->@*),
+ 1,
+ 'a failed cleanup is reported',
+ );
+
+ # ssh gives up on a call, which the target still runs to its end later
+ ($fake, $res) = run_on('alloc', faults => { 3 => 'late' });
+ is(
+ $res->{error},
+ "NVMe target operation 'create zvol' did not complete (timeout);"
+ . " the target state is unknown\n",
+ 'a create that did not complete in time',
+ );
+ is($res->{calls}->[-1]->{op}, 'create zvol', 'is not followed by anything');
+ my $second =
+ flow($fake, sub { $PLUGIN->alloc_image('st', scfg(), 201, 'raw', undef, 1024) });
+ is($second->{error}, '', 'another allocation meanwhile');
+ $fake->land;
+ is_deeply(
+ [map { owned_volumes($fake->{m})->{"tank/vm-$_-disk-0"}->[0] } 200, 201],
+ [4, 5],
+ 'never gets the NSID of the create that completes afterwards',
+ );
+ $res = flow($fake, $ACT{activate});
+ ok(
+ !$res->{error} && ns_of($fake, 4) && ns_of($fake, 5),
+ 'and the next activation exports both',
+ );
+
+ ($fake, $res) = run_on(
+ 'alloc',
+ before_call =>
+ sub($f, $call) { die "received interrupt\n" if $call->{op} eq 'export namespace' },
+ );
+ is($res->{error}, "received interrupt\n", 'a stopped task');
+ ok(!grep({ /destroy|remove/ } ops($res)->@*), 'is not compensated either');
+
+ ($fake, $res) = run_on('alloc', read_fault => { after => 5, mode => 'failed' });
+ is(
+ $res->{error},
+ "NVMe target operation 'read target state' did not complete (ssh failed);"
+ . " the target state is unknown\n",
+ 'a read under the lock that ssh did not run to its end',
+ );
+ ok(!grep({ /destroy|remove/ } ops($res)->@*), 'is not followed by a compensation');
+
+ $fake = lifecycle_fake();
+ $fake->add_zvol('tank/vm-200-disk-0');
+ my $named = sub { $PLUGIN->alloc_image('st', scfg(), 200, 'raw', 'vm-200-disk-0', 1024) };
+ refused($fake, $named, qr/zvol 'tank\/vm-200-disk-0' already exists/, 'an existing zvol');
+ $fake = lifecycle_fake();
+ delete $fake->{m}->{subsystems}->{$NQN};
+ refused(
+ $fake,
+ 'alloc',
+ qr/NVMe subsystem does not exist; activate the storage first/,
+ 'an allocation without the subsystem',
+ );
+ $res =
+ flow($fake, sub { $PLUGIN->alloc_image('st', scfg(), 200, 'raw', 'vm-201-disk-0', 1024) });
+ like($res->{error}, qr/illegal name 'vm-201-disk-0'/, 'names must belong to the VM');
+};
+
+subtest 'clone' => sub {
+ my ($fake, $res) = run_on('clone');
+ is($res->{error}, '', 'a linked clone');
+ is($res->{result}, 'base-102-disk-0/vm-201-disk-0', 'is named after its base');
+ my $clone = $fake->{m}->{ds}->{'tank/vm-201-disk-0'};
+ is($clone->{origin}, 'tank/base-102-disk-0@__base__', 'from the base snapshot');
+ my $uuid = $clone->{props}->{'proxmox:nvme-uuid'};
+ ok(exported($fake, 4, $uuid, '/dev/zvol/tank/vm-201-disk-0'), 'with its own identity');
+
+ $fake = lifecycle_fake();
+ $res =
+ flow($fake, sub { $PLUGIN->clone_image(scfg(), 'st', 'base-102-disk-0', 201, 'missing') });
+ like(
+ $res->{error},
+ qr/'create zvol' failed: cannot open 'tank\/base-102-disk-0\@missing'/,
+ 'a missing snapshot',
+ );
+ ok(!$fake->{m}->{ds}->{'tank/vm-201-disk-0'}, 'leaves nothing behind');
+ $res = flow($fake, sub { $PLUGIN->clone_image(scfg(), 'st', 'base-102-disk-0', 201, '') });
+ is($res->{error}, '', 'an empty snapshot name');
+ is(
+ $fake->{m}->{ds}->{'tank/vm-201-disk-0'}->{origin},
+ 'tank/base-102-disk-0@__base__',
+ 'means the base snapshot',
+ );
+
+ for my $case (
+ ['basevol-102-disk-0', qr/clone_image requires a ZFS volume/],
+ ['vm-100-disk-0', qr/clone_image only works on base images/],
+ ) {
+ $res = flow($fake, sub { $PLUGIN->clone_image(scfg(), 'st', $case->[0], 201) });
+ like($res->{error}, $case->[1], "refuses $case->[0]");
+ is(scalar($res->{calls}->@*), 0, 'locally');
+ }
+ $res =
+ flow($fake, sub { $PLUGIN->clone_image(scfg(), 'st', 'base-102-disk-0', 201, 'a b') });
+ like($res->{error}, qr/invalid snapshot name/, 'refuses a malformed snapshot name');
+};
+
+subtest 'destroy' => sub {
+ my $dev = '/dev/zvol/tank/vm-101-disk-0';
+ my ($fake, $res) = run_on('free');
+ is($res->{error}, '', 'a volume is freed');
+ is_deeply(ops($res), ['read target state', 'destroy zvol'], 'in two calls');
+ ok(!$fake->{m}->{ds}->{'tank/vm-101-disk-0'} && !ns_of($fake, 2), 'both are gone');
+ $res = flow($fake, $ACT{free});
+ is($res->{error}, '', 'a missing volume is already freed');
+ is(changes($res), 0, 'without changes');
+
+ for my $case (
+ [
+ 'a volume that is not owned',
+ 'vm-900-disk-0',
+ undef,
+ qr/is not owned by NVMe subsystem/,
+ ],
+ [
+ 'a UUID exported at another NSID',
+ 'vm-101-disk-0',
+ sub($f) { $f->add_namespace($NQN, 9, $U{2}, '/dev/zvol/tank/x', 0) },
+ qr/namespace UUID '\Q@{[ $U{2} ]}\E' is already in use/,
+ ],
+ [
+ 'another identity at its NSID',
+ 'vm-101-disk-0',
+ sub($f) { ns_of($f, 2)->{device_uuid} = $U{9} },
+ qr/NVMe namespace ID '2' has a different identity/,
+ ],
+ [
+ 'a duplicate identity',
+ 'vm-101-disk-0',
+ sub($f) { $f->add_zvol('tank/vm-104-disk-0', identity => [$NQN, 2, $U{9}]) },
+ qr/duplicate NVMe identity/,
+ ],
+ ) {
+ my ($name, $volume, $change, $error) = $case->@*;
+ $fake = lifecycle_fake();
+ $change->($fake) if $change;
+ refused($fake, sub { $PLUGIN->free_image('st', scfg(), $volume) }, $error, $name);
+ }
+ $res = flow($fake, sub { $PLUGIN->free_image('st', scfg(), 'subvol-100-disk-0') });
+ like($res->{error}, qr/free_image requires a ZFS volume/, 'subvolumes are refused');
+
+ my $busy = 'Device or resource busy';
+ ($fake, $res) =
+ run_on('free', fail => [{ write => '/namespaces/2/enable', value => 0 }, $busy]);
+ like($res->{error}, qr/cannot disable namespace '2': .*busy/, 'a failed disable');
+ ok(exported($fake, 2, $U{2}, $dev), 'keeps the namespace');
+
+ ($fake, $res) = run_on('free', fail => [['rmdir', "$S/namespaces/2"], $busy]);
+ like(
+ $res->{error},
+ qr/cannot remove namespace '2': Device or resource busy\n\z/,
+ 'a failed rmdir',
+ );
+ ok(exported($fake, 2, $U{2}, $dev), 're-enables the namespace');
+ is($res->{calls}->[-1]->{op}, 'enable namespace', 'from a fresh read');
+
+ my $rmdir_and_enable = sub($s) {
+ step_is($s, 'rmdir', "$S/namespaces/2") || write_to($s, '/namespaces/2/enable', 1);
+ };
+ ($fake, $res) = run_on('free', fail => [$rmdir_and_enable, $busy]);
+ like(
+ $res->{error},
+ qr/cannot remove namespace '2': .*; cannot restore namespace: /,
+ 'a failed restore is reported too',
+ );
+
+ my $destroy = ['zfs', 'destroy', '-r', 'tank/vm-101-disk-0'];
+ local $SLEPT = 0;
+ ($fake, $res) = run_on('free', fail => [$destroy, $busy]);
+ like(
+ $res->{error},
+ qr/\Afailed to destroy 'tank\/vm-101-disk-0': Device or resource busy; namespace restored/,
+ 'a zvol that stays busy',
+ );
+ is(scalar(grep { $_ eq 'destroy zvol' } ops($res)->@*), 6, 'is destroyed six times');
+ is($SLEPT, 5, 'a second apart');
+ ok(exported($fake, 2, $U{2}, $dev), 'and its namespace is restored');
+
+ ($fake, $res) = run_on('free', fail => [$destroy, $busy, 3]);
+ is($res->{error}, '', 'a zvol that becomes free is destroyed by a retry');
+ ok(!$fake->{m}->{ds}->{'tank/vm-101-disk-0'}, 'and is gone');
+
+ ($fake, $res) = run_on('free', fail => [$destroy, $busy, 1], faults => { 3 => 'after' });
+ like(
+ $res->{error},
+ qr/'destroy zvol' did not complete .* state is unknown\n\z/,
+ 'a retry whose reply is lost',
+ );
+ is($res->{calls}->[-1]->{op}, 'destroy zvol', 'is not followed by anything');
+
+ ($fake, $res) = run_on(
+ 'free',
+ fail => [$destroy, $busy],
+ read_fault => { after => 1, mode => 'unreachable' },
+ );
+ is(
+ $res->{error},
+ "failed to destroy 'tank/vm-101-disk-0': Device or resource busy\n",
+ 'a failed destroy whose outcome cannot be read is a failure',
+ );
+ ok($fake->{m}->{ds}->{'tank/vm-101-disk-0'}, 'and the zvol is still there');
+
+ ($fake, $res) = run_on(
+ 'free',
+ step_fault => sub($step, $call) {
+ return 'dataset is busy' if step_is($step, $destroy->@*);
+ return 'File exists' if step_is($step, 'mkdir', "$S/namespaces/2");
+ return undef;
+ },
+ );
+ like(
+ $res->{error},
+ qr/failed to destroy 'tank\/vm-101-disk-0': dataset is busy; namespace restoration also/,
+ 'a failed restore after the retries',
+ );
+
+ $fake = lifecycle_fake(fail => [$destroy, $busy]);
+ delete $fake->{m}->{subsystems}->{$NQN}->{ns}->{2};
+ $res = flow($fake, $ACT{free});
+ like(
+ $res->{error},
+ qr/\Afailed to destroy 'tank\/vm-101-disk-0': Device or resource busy\n\z/,
+ 'an unexported busy zvol is reported without restore',
+ );
+ ok(!ns_of($fake, 2), 'and stays unexported');
+
+ $fake = lifecycle_fake();
+ delete $fake->{m}->{subsystems}->{$NQN};
+ delete $_->{links}->{$NQN} for values $fake->{m}->{ports}->%*;
+ $res = flow($fake, $ACT{free});
+ is($res->{error}, '', 'a volume is freed without subsystem');
+ ok(!$fake->{m}->{ds}->{'tank/vm-101-disk-0'}, 'nothing exports it');
+};
+
+subtest 'rollback' => sub {
+ my $dev = '/dev/zvol/tank/vm-100-disk-0';
+ my ($fake, $res) = run_on('rollback');
+ is($res->{error}, '', 'a rollback');
+ is_deeply(
+ ops($res),
+ ['read target state', 'rollback zvol', 'wait for zvol', 'restore namespace'],
+ 'unexports, rolls back and exports again',
+ );
+ ok(exported($fake, 1, $U{1}, $dev), 'with the same identity');
+
+ # A snapshot without the identity properties: the rollback loses them.
+ my $bare = sub(%opts) {
+ my $f = lifecycle_fake(%opts);
+ $f->add_snapshot('tank/vm-100-disk-0@bare', props => {});
+ return $f;
+ };
+ for my $case (
+ ['the disable', { write => '/namespaces/1/enable', value => 0 }],
+ ['the rmdir', ['rmdir', "$S/namespaces/1"]],
+ ['the rollback', ['zfs', 'rollback']],
+ ['the identity', ['zfs', 'set']],
+ ) {
+ my ($name, $match) = $case->@*;
+ $fake = $bare->(fail => [$match, "injected $name failure", 1]);
+ $res = flow(
+ $fake,
+ sub { $PLUGIN->volume_snapshot_rollback(scfg(), 'st', 'vm-100-disk-0', 'bare') },
+ );
+ like(
+ $res->{error},
+ qr/\ANVMe target operation 'rollback zvol' failed: injected \Q$name\E failure\n\z/,
+ "a failure at $name is reported",
+ );
+ ok(exported($fake, 1, $U{1}, $dev), "the namespace is restored after $name");
+ is_deeply(
+ owned_volumes($fake->{m})->{'tank/vm-100-disk-0'},
+ [1, $U{1}],
+ 'with identity',
+ );
+ }
+ my @restore = grep { $_->{op} eq 'restore namespace' } $res->{calls}->@*;
+ ok(step_is($restore[0]->{steps}->[0], 'zfs', 'set'), 'a lost identity is written first');
+
+ ($fake, $res) = run_on(
+ 'rollback',
+ step_fault => sub($step, $call) {
+ return 'rollback failed' if step_is($step, 'zfs', 'rollback');
+ return 'mkdir failed' if step_is($step, 'mkdir');
+ return undef;
+ },
+ );
+ like(
+ $res->{error},
+ qr/
+ \Arollback\ failed:\ NVMe\ target\ operation\ 'rollback\ zvol'\ failed:
+ \ rollback\ failed;
+ \ namespace\ restoration\ failed:\ NVMe\ target\ operation\ 'restore\ namespace'
+ \ failed:\ mkdir\ failed\n\z
+ /x,
+ 'both failures are reported',
+ );
+ ($fake, $res) = run_on('rollback', fail => [['mkdir'], 'mkdir failed']);
+ like(
+ $res->{error},
+ qr/\Anamespace restoration failed after rollback: .*mkdir failed\n\z/,
+ 'a failed export after a good rollback is reported',
+ );
+
+ $fake = lifecycle_fake();
+ for my $case (
+ ['a malformed snapshot name', 'vm-100-disk-0', 'a b', qr/invalid snapshot name/],
+ [
+ 'a subvolume',
+ 'subvol-100-disk-0',
+ 'snap1',
+ qr/snapshot rollback requires a ZFS volume/,
+ ],
+ ['a volume that is not owned', 'vm-900-disk-0', 'snap1', qr/is not owned/],
+ ) {
+ my ($name, $volume, $snap, $error) = $case->@*;
+ my $rollback = sub { $PLUGIN->volume_snapshot_rollback(scfg(), 'st', $volume, $snap) };
+ refused($fake, $rollback, $error, $name);
+ }
+ my $taken = lifecycle_fake();
+ $taken->add_namespace($NQN, 1, $U{9}, '/dev/zvol/tank/vm-900-disk-0');
+ refused(
+ $taken,
+ 'rollback',
+ qr/NVMe namespace ID '1' has a different identity/,
+ 'a namespace ID that holds another identity',
+ );
+ delete $fake->{m}->{subsystems}->{$NQN};
+ refused(
+ $fake,
+ 'rollback',
+ qr/NVMe subsystem does not exist/,
+ 'a rollback without the subsystem',
+ );
+};
+
+subtest 'template' => sub {
+ my $old = '/dev/zvol/tank/vm-101-disk-0';
+ my $new = '/dev/zvol/tank/base-101-disk-0';
+ my ($fake, $res) = run_on('template');
+ is($res->{result}, 'base-101-disk-0', 'a template');
+ is_deeply(
+ ops($res),
+ ['read target state', 'rename zvol', 'wait for zvol', 'create template'],
+ 'is renamed, exported and snapshotted',
+ );
+ ok(exported($fake, 2, $U{2}, $new), 'under the same identity');
+ ok($fake->{m}->{snaps}->{'tank/base-101-disk-0@__base__'}, 'with its base snapshot');
+
+ my $restored = sub($f) {
+ return
+ exported($f, 2, $U{2}, $old)
+ && $f->{m}->{ds}->{'tank/vm-101-disk-0'}
+ && !$f->{m}->{ds}->{'tank/base-101-disk-0'};
+ };
+ for my $case (
+ ['the unexport', sub($s, $c) { write_to($s, '/namespaces/2/enable', 0) }],
+ ['the rename', sub($s, $c) { step_is($s, 'zfs', 'rename', 'tank/vm-101-disk-0') }],
+ ['the export', sub($s, $c) { step_is($s, 'mkdir') && $c->{op} eq 'create template' }],
+ ['the snapshot', sub($s, $c) { step_is($s, 'zfs', 'snapshot') }],
+ ) {
+ my ($name, $match) = $case->@*;
+ my $fault =
+ sub($step, $call) { $match->($step, $call) ? "injected $name failure" : undef };
+ ($fake, $res) = run_on('template', step_fault => $fault);
+ my $prefix = qr/\Atemplate conversion failed: NVMe target operation '[a-z ]+' failed:/;
+ like(
+ $res->{error},
+ qr/$prefix injected \Q$name\E failure\n\z/,
+ "a failure at $name is reported",
+ );
+ ok($restored->($fake), "the old name is exported again after $name");
+ }
+
+ ($fake, $res) = run_on(
+ 'template',
+ after_call => sub($f, $call) {
+ $f->{m}->{ds}->{'tank/base-101-disk-0'}->{devwait} = 1000
+ if $call->{op} eq 'rename zvol';
+ },
+ );
+ like(
+ $res->{error},
+ qr/\Atemplate conversion failed: zvol '\Q$new\E' is not a block device after 10 seconds\n/,
+ 'a device that never appears',
+ );
+ ok($restored->($fake), 'is renamed back');
+
+ ($fake, $res) = run_on('template', faults => { 1 => 'late' });
+ like(
+ $res->{error},
+ qr/\ANVMe target operation 'rename zvol' did not complete .* state is unknown\n\z/,
+ 'a rename that does not complete in time is passed on as it is',
+ );
+ ok(!grep({ $_->{op} =~ /back|restore/ } $res->{calls}->@*), 'is not compensated');
+ $fake->land;
+ $res = flow($fake, $ACT{activate});
+ ok(!$res->{error} && exported($fake, 2, $U{2}, $new), 'the next activation exports it');
+
+ my $taken = lifecycle_fake();
+ $taken->add_namespace($NQN, 2, $U{9}, '/dev/zvol/tank/vm-900-disk-0');
+ refused(
+ $taken,
+ 'template',
+ qr/NVMe namespace ID '2' has a different identity/,
+ 'a namespace ID that holds another identity',
+ );
+ $fake = lifecycle_fake();
+ $fake->add_zvol('tank/base-101-disk-0');
+ refused(
+ $fake,
+ 'template',
+ qr/template zvol 'tank\/base-101-disk-0' already exists/,
+ 'a taken name',
+ );
+ $res = flow($fake, sub { $PLUGIN->create_base('st', scfg(), 'base-102-disk-0') });
+ like($res->{error}, qr/create_base not possible with base image/, 'refuses a base');
+ refused(
+ $fake,
+ sub { $PLUGIN->create_base('st', scfg(), 'vm-900-disk-0') },
+ qr/is not owned/,
+ 'a volume that is not owned',
+ );
+};
+
+subtest 'resize' => sub {
+ my ($fake, $res) = run_on('resize');
+ is($res->{error}, '', 'an exported volume grows');
+ is($res->{result}, 2 * 1024**2, 'the new size in KiB is returned');
+ is_deeply(ops($res), ['read target state', 'resize zvol'], 'read, then resize');
+ ok(!grep({ !$_->{locked} } $res->{calls}->@*), 'under the lock');
+ is($fake->{revalidated}->{"$NQN/1"}, 1,
+ 'the namespace reads the new size in the same call');
+
+ $fake = lifecycle_fake();
+ delete $fake->{m}->{subsystems}->{$NQN}->{ns}->{1};
+ $res = flow($fake, $ACT{resize});
+ is($res->{error}, '', 'an unexported volume grows');
+ is(scalar($res->{calls}->[1]->{steps}->@*), 1, 'without revalidation');
+ $fake = lifecycle_fake();
+ ns_of($fake, 1)->{enable} = '0';
+ $res = flow($fake, $ACT{resize});
+ is(scalar($res->{calls}->[1]->{steps}->@*), 1, 'a disabled namespace is not revalidated');
+ $fake = lifecycle_fake()->reboot(0);
+ $res = flow($fake, $ACT{resize});
+ is($res->{error}, '', 'a volume grows after a target restart');
+ is($fake->{m}->{ds}->{'tank/vm-100-disk-0'}->{volsize}, 2 * 1024**3,
+ 'once nvmet is loaded');
+
+ for my $case (
+ [
+ 'a duplicate UUID',
+ 'vm-100-disk-0',
+ sub($f) { $f->add_namespace($NQN, 9, $U{1}, '/dev/zvol/tank/x', 0) },
+ qr/duplicate namespace UUID/,
+ ],
+ ['a volume that is not owned', 'vm-900-disk-0', sub($f) { }, qr/is not owned/],
+ ['a volume of another subsystem', 'vm-901-disk-0', sub($f) { }, qr/is not owned/],
+ ) {
+ my ($name, $volname, $change, $error) = $case->@*;
+ $fake = lifecycle_fake();
+ $change->($fake);
+ my $resize = sub { $PLUGIN->volume_resize(scfg(), 'st', $volname, 2 * 1024**3) };
+ refused($fake, $resize, $error, $name);
+ }
+ ($fake, $res) = run_on('resize', fail => [['zfs', 'set'], 'out of space']);
+ like($res->{error}, qr/'resize zvol' failed: out of space/, 'a failed resize');
+ is_deeply(ops($res), ['read target state', 'resize zvol'], 'stops there');
+ ($fake, $res) =
+ run_on('resize', fail => [{ write => 'revalidate_size' }, 'Invalid argument']);
+ like($res->{error}, qr/'resize zvol' failed: Invalid argument/, 'a failed revalidation');
+ is($fake->{m}->{ds}->{'tank/vm-100-disk-0'}->{volsize}, 2 * 1024**3, 'keeps the new size');
+};
+
+sub removal_fake(%opts) {
+ my $fake = target_fake(%opts);
+ $fake->add_zvol('other/foreign-disk', identity => [$FOREIGN_NQN, 1, $U{8}]);
+ $fake->add_subsystem(
+ $FOREIGN_NQN,
+ model => 'Linux',
+ serial => '0123456789abcdef',
+ acl => [$HOSTS[0]],
+ );
+ $fake->add_namespace($FOREIGN_NQN, 1, $U{8}, '/dev/zvol/other/foreign-disk');
+ $fake->{m}->{ports}->{1}->{links}->{$FOREIGN_NQN} = 1;
+ return $fake;
+}
+
+subtest 'target removal' => sub {
+ my $fake = removal_fake();
+ my $res = flow($fake, $ACT{remove});
+ is($res->{error}, '', 'the subsystem is removed');
+ is_deeply(
+ ops($res),
+ [
+ 'read target state',
+ 'remove NVMe subsystem',
+ 'read target state',
+ 'remove orphan NVMe hosts',
+ ],
+ 'teardown, then orphans from a fresh read',
+ );
+ my $m = $fake->{m};
+ ok(
+ !$m->{subsystems}->{$NQN} && $m->{hosts}->{ $HOSTS[0] } && !$m->{hosts}->{ $HOSTS[1] },
+ 'subsystem and orphan host are gone, the host linked by another subsystem stays',
+ );
+ ok($m->{ports}->{1} && $m->{ports}->{2}, 'ports are never removed');
+ $res = flow($fake, $ACT{remove});
+ is($res->{error}, '', 'a removed subsystem is already removed');
+ is_deeply(ops($res), ['read target state'], 'with nothing left to do');
+
+ $fake = removal_fake();
+ delete $fake->{m}->{subsystems}->{$NQN};
+ delete $_->{links}->{$NQN} for values $fake->{m}->{ports}->%*;
+ $res = flow($fake, $ACT{remove});
+ is_deeply(
+ ops($res),
+ ['read target state', 'remove orphan NVMe hosts'],
+ 'without subsystem the configured hosts are candidates',
+ );
+ ok(
+ $fake->{m}->{hosts}->{ $HOSTS[0] } && !$fake->{m}->{hosts}->{ $HOSTS[1] },
+ 'only unlinked hosts go',
+ );
+
+ for my $case (
+ [
+ 'an owned volume',
+ sub($f) { add_volume($f, 'vm-100-disk-0', 1, $U{1}, unexported => 1) },
+ qr/owned ZFS volume 'tank\/vm-100-disk-0' exists/,
+ ],
+ [
+ 'a foreign subsystem',
+ sub($f) { $f->{m}->{subsystems}->{$NQN}->{attr}->{attr_model} = 'Linux' },
+ qr/refusing to delete foreign NVMe subsystem/,
+ ],
+ [
+ 'remaining namespaces',
+ sub($f) { $f->add_namespace($NQN, 3, $U{3}, '(null)', 0) },
+ qr/namespaces remain/,
+ ],
+ ) {
+ my ($name, $change, $error) = $case->@*;
+ $fake = removal_fake();
+ $change->($fake);
+ refused($fake, 'remove', $error, $name);
+ }
+
+ $fake =
+ removal_fake(fail => [['rmdir', "$ROOT/hosts/$HOSTS[1]"], 'Device or resource busy']);
+ $res = flow($fake, $ACT{remove});
+ is($res->{error}, '', 'a failed orphan removal');
+ like(
+ $res->{warnings}->[0],
+ qr/could not remove orphan NVMe host\(s\): Device or resource busy/,
+ 'only warns',
+ );
+
+ $fake = removal_fake(fail => [['rmdir', $S], 'Directory not empty']);
+ $res = flow($fake, $ACT{remove});
+ like(
+ $res->{error},
+ qr/cannot remove NVMe subsystem '\Q$NQN\E': Directory not empty/,
+ 'a failed teardown is reported',
+ );
+ ok($fake->{m}->{subsystems}->{$NQN}, 'and the subsystem stays');
+};
+
+subtest 'snapshots and the generic ZFS calls' => sub {
+ my $fake = lifecycle_fake();
+ my $res =
+ flow($fake, sub { $PLUGIN->volume_snapshot(scfg(), 'st', 'vm-100-disk-0', 'snap_2.0') });
+ is($res->{error}, '', 'a snapshot');
+ ok($fake->{m}->{snaps}->{'tank/vm-100-disk-0@snap_2.0'}, 'is taken');
+ is(scalar($res->{locks}->@*), 1, 'under the target lock');
+ $res = flow(
+ $fake,
+ sub { $PLUGIN->volume_rollback_is_possible(scfg(), 'st', 'vm-100-disk-0', 'snap1') },
+ );
+ like(
+ $res->{error},
+ qr/'snap1' is not most recent snapshot/,
+ 'only from the latest snapshot',
+ );
+ $res = flow(
+ $fake,
+ sub { $PLUGIN->volume_snapshot_delete(scfg(), 'st', 'vm-100-disk-0', 'snap_2.0') },
+ );
+ ok(
+ !$res->{error} && !$fake->{m}->{snaps}->{'tank/vm-100-disk-0@snap_2.0'},
+ 'deleted by name',
+ );
+ is(scalar($res->{locks}->@*), 1, 'under the target lock too');
+ $res = flow(
+ $fake,
+ sub { $PLUGIN->volume_rollback_is_possible(scfg(), 'st', 'vm-100-disk-0', 'snap1') },
+ );
+ is($res->{result}, 1, 'now the rollback is possible');
+ $res = flow(
+ $fake,
+ sub { $PLUGIN->volume_rollback_is_possible(scfg(), 'st', 'vm-100-disk-0', 'gone') },
+ );
+ like($res->{error}, qr/snapshot 'gone' does not exist/, 'and needs an existing snapshot');
+ $res = flow($fake, sub { $PLUGIN->volume_snapshot_info(scfg(), 'st', 'vm-100-disk-0') });
+ is_deeply([sort keys $res->{result}->%*], ['snap1'], 'snapshot information');
+
+ $res = flow($fake, sub { [$PLUGIN->volume_size_info(scfg(), 'st', 'vm-100-disk-0', 7)] });
+ is_deeply($res->{result}, [1073741824, 'raw', 4096, undef], 'size information');
+ is($res->{calls}->[0]->{timeout}, 7, 'with the timeout of the caller');
+ $res =
+ flow($fake, sub { scalar($PLUGIN->volume_size_info(scfg(), 'st', 'vm-100-disk-0')) });
+ is($res->{result}, 1073741824, 'the size alone in scalar context');
+
+ # Snapshot names are validated before zfs reads '%' and ',' in them as a
+ # range and a list, and the pool name wherever a call uses it.
+ for my $case (
+ ['volume_snapshot', 'a b'],
+ ['volume_snapshot_delete', 'snap1%'],
+ ['volume_rollback_is_possible', 'snap1,snap2'],
+ ) {
+ my ($method, $snap) = $case->@*;
+ $res = refused(
+ $fake,
+ sub { $PLUGIN->$method(scfg(), 'st', 'vm-100-disk-0', $snap) },
+ qr/\Ainvalid snapshot name\n\z/,
+ "$method with the snapshot name '$snap'",
+ );
+ is(scalar($res->{calls}->@*), 0, 'locally');
+ }
+ my $bad_pool = scfg(pool => '-rH');
+ for my $case (
+ ['list_images', sub { $PLUGIN->list_images('st', $bad_pool) }],
+ [
+ 'volume_size_info',
+ sub { $PLUGIN->volume_size_info($bad_pool, 'st', 'vm-100-disk-0') },
+ ],
+ [
+ 'volume_snapshot_info',
+ sub { $PLUGIN->volume_snapshot_info($bad_pool, 'st', 'vm-100-disk-0') },
+ ],
+ [
+ 'volume_snapshot',
+ sub { $PLUGIN->volume_snapshot($bad_pool, 'st', 'vm-100-disk-0', 'snap2') },
+ ],
+ ) {
+ $res = refused(
+ $fake,
+ $case->[1],
+ qr/\Ainvalid ZFS pool name\n\z/,
+ "$case->[0] with an invalid pool name",
+ );
+ is(scalar($res->{calls}->@*), 0, 'locally');
+ }
+ my @warnings;
+ {
+ local $SIG{__WARN__} = sub($warning) { push @warnings, $warning };
+ $res = flow($fake, sub { [$PLUGIN->status('st', $bad_pool)] });
+ }
+ is_deeply(
+ [$res->{result}, $res->{calls}, \@warnings],
+ [[0, 0, 0, 0], [], ["storage 'st': invalid ZFS pool name\n"]],
+ 'status with an invalid pool name is inactive, locally',
+ );
+
+ # The plugin creates zvols only.
+ for my $case (
+ [
+ 'volume_snapshot',
+ sub { $PLUGIN->volume_snapshot(scfg(), 'st', 'subvol-100-disk-0', 's') },
+ ],
+ [
+ 'volume_size_info',
+ sub { $PLUGIN->volume_size_info(scfg(), 'st', 'subvol-100-disk-0') },
+ ],
+ ) {
+ $res = refused(
+ $fake,
+ $case->[1],
+ qr/\A$case->[0] requires a ZFS volume\n\z/,
+ "$case->[0] of a subvolume",
+ );
+ is(scalar($res->{calls}->@*), 0, 'locally');
+ }
+};
+
+subtest 'lock ownership' => sub {
+ my $fake = lifecycle_fake();
+ pipe(my $reader, my $writer) or die "pipe: $!\n";
+ my $child;
+ $fake->{before_call} = sub($f, $call) {
+ return if $child || $call->{op} ne 'create zvol';
+ $child = fork() // die "fork: $!\n";
+ return if $child;
+ close($reader);
+ my $res = flow(
+ $f, sub { $PLUGIN->volume_resize(scfg(), 'st', 'vm-100-disk-0', 2 * 1024**3, 0) },
+ );
+ my $locked = grep { $_->{pid} == $$ } $res->{locks}->@*;
+ my @counts = ($locked, scalar($res->{nested}->@*), scalar($res->{calls}->@*));
+ print {$writer} "@counts\n";
+ close($writer);
+ POSIX::_exit(0);
+ };
+ my $res = flow($fake, sub { $PLUGIN->alloc_image('st', scfg(), 200, 'raw', undef, 1024) });
+ close($writer);
+ my $line = <$reader>;
+ waitpid($child, 0);
+ is($res->{error}, '', 'the parent keeps its lock');
+ is($line, "1 1 0\n", 'a child requests the lock itself, waits and changes nothing');
+
+ local $QUORATE = 0;
+ $fake = lifecycle_fake();
+ for my $action (qw(alloc free rollback template resize remove)) {
+ $res = flow($fake, $ACT{$action});
+ like($res->{error}, qr/cluster not quorate/, "$action needs quorum");
+ ok(!$res->{locks}->@* && !changes($res), 'and fails before the lock');
+ }
+};
+
+subtest 'target call timeouts' => sub {
+ my ($fake, $res) = run_on('alloc', after_call => sub($f, $call) { $NOW += 100 });
+ is($res->{error}, '', 'slow target calls do not end an operation');
+ is_deeply(
+ [map { $_->{timeout} } grep { $_->{locked} } $res->{calls}->@*],
+ [15, 60, 60, 15, 30, 15],
+ 'reads take up to 15 s, zfs changes 60 s and configfs changes 30 s',
+ );
+
+ # in a worker, a zfs change may take an hour, like in the other ZFS plugins
+ my $rpcenv_mock = Test::MockModule->new('PVE::RPCEnvironment');
+ $rpcenv_mock->redefine(is_worker => sub(@args) { return 1 });
+ my %timeouts;
+ for my $action (qw(alloc template resize rollback free)) {
+ (undef, $res) = run_on($action);
+ $timeouts{ $_->{op} } = $_->{timeout} for $res->{calls}->@*;
+ }
+ is_deeply(
+ [
+ @timeouts{
+ 'reserve namespace ID',
+ 'create zvol',
+ 'rename zvol',
+ 'create template',
+ 'resize zvol',
+ 'rollback zvol',
+ 'restore namespace',
+ 'destroy zvol',
+ },
+ ],
+ [(3600) x 8],
+ 'every zfs change',
+ );
+ $rpcenv_mock->unmock_all();
+};
+
+# ---------------------------------------------------------------------------
+# Fault sweep
+# ---------------------------------------------------------------------------
+
+# The goals of the swept flows. They only read the state: a lookup that
+# created a missing namespace would change the outcome of the next check.
+sub exported_ns($m, $nsid) {
+ my $subsys = $m->{subsystems}->{$NQN} // return undef;
+ return $subsys->{ns}->{$nsid};
+}
+
+sub template_done($m, $res) {
+ my $ds = owned_volumes($m)->{'tank/base-101-disk-0'} // return 'template missing';
+ my $ns = exported_ns($m, 2) // return 'template not exported';
+ return 'template not exported' if $ns->{device_path} ne '/dev/zvol/tank/base-101-disk-0';
+ return $m->{snaps}->{'tank/base-101-disk-0@__base__'} ? () : 'base snapshot missing';
+}
+
+sub new_volume_done($name) {
+ return sub($m, $res) {
+ my $owned = owned_volumes($m)->{"tank/$name"} // return "$name missing";
+ my $ns = exported_ns($m, $owned->[0]) // return "$name not exported";
+ return () if $ns->{enable} eq '1' && $ns->{device_uuid} eq $owned->[1];
+ return "$name not exported";
+ };
+}
+
+sub removal_goal($m) {
+ my @bad;
+ push @bad, 'subsystem left' if $m->{subsystems}->{$NQN};
+ push @bad, 'orphan host left' if $m->{hosts}->{ $HOSTS[1] };
+ push @bad, 'shared host removed' if !$m->{hosts}->{ $HOSTS[0] };
+ push @bad, 'port removed' if !$m->{ports}->{1} || !$m->{ports}->{2};
+ return @bad;
+}
+
+subtest 'fault sweep' => sub {
+ sweep(
+ 'create',
+ {
+ setup => \&lifecycle_fake,
+ lost_replies => 1,
+ action => $ACT{alloc},
+ done => new_volume_done('vm-200-disk-0'),
+ },
+ );
+ sweep(
+ 'clone',
+ {
+ setup => \&lifecycle_fake,
+ lost_replies => 1,
+ action => $ACT{clone},
+ done => new_volume_done('vm-201-disk-0'),
+ },
+ );
+ sweep(
+ 'destroy',
+ {
+ setup => \&lifecycle_fake,
+ action => $ACT{free},
+ may_vanish => 'tank/vm-101-disk-0',
+ done => sub($m, $res) { $m->{ds}->{'tank/vm-101-disk-0'} ? 'volume left' : () },
+ },
+ );
+ sweep(
+ 'rollback',
+ {
+ setup => \&lifecycle_fake,
+ action => $ACT{rollback},
+ done => sub($m, $res) {
+ my $ns = exported_ns($m, 1) // return 'not exported';
+ return $ns->{enable} eq '1' ? () : 'not exported';
+ },
+ },
+ );
+ sweep(
+ 'template',
+ {
+ setup => \&lifecycle_fake,
+ action => $ACT{template},
+ renames => { 'tank/vm-101-disk-0' => 'tank/base-101-disk-0' },
+ done => \&template_done,
+ },
+ );
+ sweep(
+ 'resize',
+ {
+ setup => \&lifecycle_fake,
+ action => $ACT{resize},
+ done => sub($m, $res) {
+ my $ds = $m->{ds}->{'tank/vm-100-disk-0'} // return 'volume missing';
+ return 'not resized' if $ds->{volsize} != 2 * 1024**3;
+ return $res->{fake}->{revalidated}->{"$NQN/1"} ? () : 'not revalidated';
+ },
+ },
+ );
+ sweep(
+ 'activation after a configfs loss',
+ {
+ setup => sub () { lifecycle_fake()->reboot(1) },
+ action => $ACT{activate},
+ done => sub($m, $res) { () },
+ },
+ );
+ sweep(
+ 'target removal',
+ {
+ setup => \&removal_fake,
+ action => $ACT{remove},
+ retry => $ACT{remove},
+ goal => \&removal_goal,
+ teardown => 1,
+ done => sub($m, $res) { $m->{subsystems}->{$NQN} ? 'subsystem left' : () },
+ },
+ );
+};
+
+# ---------------------------------------------------------------------------
+# Two storages on one target, serialized by the lock mock
+# ---------------------------------------------------------------------------
+
+subtest 'two storages on one target' => sub {
+ my $nqn_b = 'nqn.2026-01.com.example:zfsnvme-b';
+ my $scfg_a = scfg(pool => 'tank/a');
+ my $scfg_b = scfg(pool => 'tank/b', subsysnqn => $nqn_b);
+ my $new = sub () { FakeTarget->new(pools => [qw(tank tank/a tank/b)]) };
+
+ # B, with another key, takes the lock first, between A's lockless read
+ # and A's locked read
+ my $fake = $new->();
+ my ($res, $b_res);
+ {
+ local $LOCK_HOOK = sub($name) {
+ $b_res = flow($fake, sub { activate($scfg_b, $KEY_B) });
+ };
+ $res = flow($fake, sub { activate($scfg_a, $KEY) });
+ }
+ is($b_res->{error}, '', 'B configures the target first');
+ like($res->{error}, qr/in-use DH-HMAC-CHAP key/, 'A is refused by its locked read');
+ ok(!$fake->{m}->{subsystems}->{$NQN}, 'without creating anything');
+ is_deeply(
+ [map { $fake->{m}->{hosts}->{$_}->{key} } @HOSTS],
+ [$KEY_B, $KEY_B],
+ "B's key is kept",
+ );
+
+ # (a) a target restart between the verify read and the publish
+ for my $loaded (0, 1) {
+ $fake = FakeTarget->new(pools => ['tank']);
+ my $restarted = 0;
+ my @online;
+ my $ctx =
+ { nqn => $NQN, pool => 'tank', hosts => [@HOSTS], key => $KEY, port_ids => [] };
+ $ctx->{foreign} = foreign_view($fake->{m}, [], $NQN, [@HOSTS]);
+ $fake->{before_call} = sub($f, $call) {
+ $f->reboot($loaded) if $call->{op} =~ /\Alink / && !$restarted++;
+ };
+ $fake->{after_call} = sub($f, $call) { push @online, online_violations($f, $ctx) };
+ $res = flow($fake, $ACT{activate});
+ like(
+ $res->{error},
+ qr/cannot publish NVMe subsystem '\Q$NQN\E'/,
+ 'a restart before the publish fails it',
+ );
+ ok(!grep({ $_->{links}->%* } values $fake->{m}->{ports}->%*), 'nothing is linked');
+ $fake->{before_call} = undef;
+ $res = flow($fake, $ACT{activate});
+ is($res->{error}, '', 'the next activation restores the target');
+ is_deeply(\@online, [], 'never published before the ACLs and keys');
+ ok(
+ $fake->{m}->{ports}->{1}->{links}->{$NQN}
+ && $fake->{m}->{ports}->{2}->{links}->{$NQN},
+ 'and publishes it',
+ );
+ }
+
+ # ports are shared, whoever creates them
+ $fake = $new->();
+ {
+ local $LOCK_HOOK = sub($name) {
+ $b_res = flow($fake, sub { activate($scfg_b, $KEY) });
+ };
+ $res = flow($fake, sub { activate($scfg_a, $KEY) });
+ }
+ ok(!$res->{error} && !$b_res->{error}, 'two storages with one key on the same portals');
+ is_deeply(
+ [sort map { $_->{attr}->{addr_traddr} } values $fake->{m}->{ports}->%*],
+ ['192.0.2.21', '192.0.2.22'],
+ 'share one port per address',
+ );
+ ok(
+ !grep({ !($_->{links}->{$NQN} && $_->{links}->{$nqn_b}) }
+ values $fake->{m}->{ports}->%*),
+ 'both subsystems are published on both',
+ );
+
+ # lifecycle operations of both storages
+ my ($a_vol, $b_vol);
+ {
+ local $LOCK_HOOK = sub($name) {
+ $b_vol = flow(
+ $fake, sub { $PLUGIN->alloc_image('b', $scfg_b, 300, 'raw', undef, 1024) },
+ );
+ };
+ $a_vol =
+ flow($fake, sub { $PLUGIN->alloc_image('a', $scfg_a, 300, 'raw', undef, 1024) });
+ }
+ is_deeply(
+ [$a_vol->{result}, $b_vol->{result}],
+ ['vm-300-disk-0', 'vm-300-disk-0'],
+ 'volumes in both pools',
+ );
+ is_deeply(
+ [
+ sort keys $fake->{m}->{subsystems}->{$NQN}->{ns}->%*,
+ sort keys $fake->{m}->{subsystems}->{$nqn_b}->{ns}->%*,
+ ],
+ [1, 1],
+ 'with independent namespace IDs',
+ );
+};
+
+# ---------------------------------------------------------------------------
+# Target details
+# ---------------------------------------------------------------------------
+
+subtest 'each port is published on its own' => sub {
+ my $fake = FakeTarget->new(pools => ['tank'], unbindable => { 2 => 1 });
+ my $res = flow($fake, $ACT{activate});
+ is($res->{error}, '', 'a port that cannot bind its address');
+ like(
+ join("\n", $res->{warnings}->@*),
+ qr/cannot publish NVMe subsystem '\Q$NQN\E' on port 2: .*Cannot assign requested address/,
+ 'only warns',
+ );
+ ok(
+ $fake->{m}->{ports}->{1}->{links}->{$NQN} && !$fake->{m}->{ports}->{2}->{links}->{$NQN},
+ 'while the other port serves the subsystem',
+ );
+ $res = flow($fake, $ACT{activate});
+ ok(grep({ $_ eq 'link NVMe subsystem on port 2' } ops($res)->@*), 'and is retried');
+
+ $fake = FakeTarget->new(pools => ['tank'], unbindable => { 1 => 1, 2 => 1 });
+ $res = flow($fake, $ACT{activate});
+ like(
+ $res->{error},
+ qr/\Acannot publish NVMe subsystem '\Q$NQN\E': ln: .*Cannot assign requested address\n\z/,
+ 'without any port the activation fails',
+ );
+};
+
+subtest 'UUIDs are compared in lower case' => sub {
+ my $upper = uc($U{9} =~ tr/9/a/r);
+ my $lower = lc($upper);
+ is(
+ inventory(@POOL, zfs_rows('tank/vm-100-disk-0', 'volume', $NQN, 1, $upper))->{volumes}
+ ->{'tank/vm-100-disk-0'}->{uuid},
+ $lower,
+ 'the inventory normalizes the UUID',
+ );
+ my $dup = inventory(
+ @POOL,
+ zfs_rows('tank/vm-100-disk-0', 'volume', $NQN, 1, $upper),
+ zfs_rows('tank/vm-101-disk-0', 'volume', $NQN, 2, $lower),
+ );
+ eval { nv('_nvmet_owned_identity', $dup, $NQN, 'tank/vm-100-disk-0') };
+ like($@, qr/duplicate NVMe identity/, 'a UUID that differs only in case is a duplicate');
+
+ my $fake = target_fake();
+ $fake->add_zvol('tank/vm-100-disk-0', identity => [$NQN, 1, $upper]);
+ my $res = flow($fake, $ACT{activate});
+ is($res->{error}, '', 'the volume is exported');
+ ok(exported($fake, 1, $lower, '/dev/zvol/tank/vm-100-disk-0'), 'under its UUID');
+ $res = flow($fake, $ACT{activate});
+ is_deeply(ops($res), ['read target state'], 'which then is converged');
+};
+
+subtest 'removing the storage is best effort' => sub {
+ my $file_mock = Test::MockModule->new('PVE::File');
+ $file_mock->redefine(
+ dir_glob_foreach => sub($dir, $regex, $func) {
+ $func->('nvme7') if $dir eq '/sys/class/nvme';
+ },
+ );
+ local %FILES = (%FILES, '/sys/class/nvme/nvme7/subsysnqn' => $NQN);
+ my $remove = sub($fake) {
+ local @UNLINKED = ();
+ my $res = flow($fake, sub { $PLUGIN->on_delete_hook('st', scfg()) });
+ $res->{unlinked} = [sort @UNLINKED];
+ $res->{disconnects} =
+ scalar(grep { $_->[0] eq '/sys/class/nvme/nvme7/delete_controller' }
+ $res->{sysfs_writes}->@*);
+ return $res;
+ };
+
+ my $fake = lifecycle_fake();
+ my $before = dclone($fake->{m});
+ my $res = $remove->($fake);
+ is($res->{error}, '', 'a storage that owns volumes can be removed');
+ my $owned = "refusing to delete NVMe subsystem '$NQN': owned ZFS volume"
+ . " 'tank/base-102-disk-0' exists";
+ like(
+ join("\n", $res->{warnings}->@*),
+ qr/keeping the NVMe target configuration of storage 'st': \Q$owned\E/,
+ 'its volumes and subsystem are kept',
+ );
+ ok(!changes($res) && same($before, $fake->{m}), 'the target is untouched');
+ is_deeply($res->{unlinked}, [secret_file('st')], 'the key file is removed');
+ is($res->{disconnects}, 1, 'the local paths are disconnected');
+
+ $fake = removal_fake();
+ local $FILES{ secret_file('st') } = $KEY;
+ my $key_at_first_call;
+ $fake->{before_call} =
+ sub($f, $call) { $key_at_first_call //= $FILES{ secret_file('st') } };
+ $res = $remove->($fake);
+ is($res->{error}, '', 'an empty storage is removed');
+ ok(
+ !$fake->{m}->{subsystems}->{$NQN} && !$fake->{m}->{hosts}->{ $HOSTS[1] },
+ 'with its subsystem and orphan hosts',
+ );
+ ok(scalar($res->{calls}->@*) && !defined($key_at_first_call), 'after its key is removed');
+ $fake->{before_call} = undef;
+
+ # Another node read the key before the removal, and its activation waited
+ # for the target lock until the removal released it.
+ $res =
+ flow($fake, sub { nv('_nvmet_activate_target', 'st', scfg(), $PORTALS, [@HOSTS], $KEY) });
+ is($res->{error}, "storage 'st' is being removed or its key changed\n",
+ 'a late activation');
+ ok(
+ !changes($res)
+ && !$fake->{m}->{subsystems}->{$NQN}
+ && !$fake->{m}->{hosts}->{ $HOSTS[1] },
+ 'does not restore the target',
+ );
+
+ $plugin_mock->redefine(
+ _namespace_openers => sub($nqn) { return ['qemu (PID 7, /dev/nvme0n1p1)'] });
+ $fake = removal_fake();
+ $res = $remove->($fake);
+ is($res->{error}, '', 'a storage with open namespaces can be removed');
+ like(
+ $res->{warnings}->[0],
+ qr/not disconnecting NVMe storage 'st': .*in use by qemu \(PID 7, \/dev\/nvme0n1p1/,
+ 'without disconnecting',
+ );
+ is($res->{disconnects}, 0, 'the paths stay connected');
+ is_deeply($res->{unlinked}, [secret_file('st')], 'and its key file is removed');
+ $plugin_mock->redefine(_namespace_openers => sub($nqn) { return [] });
+};
+
+# ---------------------------------------------------------------------------
+# Renderer
+# ---------------------------------------------------------------------------
+
+my $INVENTORY_TEXT = 'zfs get -H -p -d 1 -t filesystem,volume -o name,property,value,source'
+ . ' type,proxmox:nvme-subsys,proxmox:nvme-nsid,proxmox:nvme-uuid,proxmox:nvme-last-nsid tank';
+my $MARKER_TEXT = q{printf '%s\n' ZFSNVME-CONFIGFS};
+my $CONFIGFS_TEXT =
+ q{env 'LC_ALL=C' find /sys/kernel/config/nvmet}
+ . q{ '(' -path '/sys/kernel/config/nvmet/subsystems/*/passthru'}
+ . q{ -o -path '/sys/kernel/config/nvmet/ports/*/ana_groups'}
+ . q{ -o -path '/sys/kernel/config/nvmet/ports/*/referrals' ')' -type d -prune}
+ . q{ -o -type d -exec printf 'D %s\n' '{}' +}
+ . q{ -o -type l -exec printf 'L %s\n' '{}' +}
+ . q{ -o -type f '(' -name enable -o -name device_path -o -name device_uuid -o -name buffered_io}
+ . q{ -o -name attr_model -o -name attr_serial -o -name attr_allow_any_host}
+ . q{ -o -name addr_trtype -o -name addr_adrfam -o -name addr_traddr -o -name addr_trsvcid ')'}
+ . q{ -exec grep '' /dev/null '{}' +};
+my $CONFIGFS_KEYS_TEXT =
+ $CONFIGFS_TEXT
+ . q{ -o -type f '('}
+ . join(' -o', map { " -path $ROOT/hosts/$_/dhchap_key" } @HOSTS)
+ . q{ ')' -exec sha256sum '{}' +};
+my $STATE_TEXT = "$INVENTORY_TEXT && $MARKER_TEXT && $CONFIGFS_TEXT";
+my $LOCKED_STATE_TEXT = "modprobe nvmet_tcp && $STATE_TEXT";
+
+sub w_text($path, $value) {
+ return "printf '%s\\n' $value > $path";
+}
+
+sub build_text($nsid, $uuid, $dev) {
+ my $ns = "$S/namespaces/$nsid";
+ return join(
+ ' && ',
+ "test -b $dev", "mkdir $ns",
+ w_text("$ns/device_path", $dev),
+ w_text("$ns/device_uuid", $uuid),
+ w_text("$ns/buffered_io", 0),
+ w_text("$ns/enable", 1),
+ );
+}
+
+sub rendered($res) {
+ return [map { $_->{rendered} } $res->{calls}->@*];
+}
+
+subtest 'golden commands' => sub {
+ my @H = map { "$ROOT/hosts/$_" } @HOSTS;
+ my $serial = 'PVEZFS' . substr(sha256_hex($NQN), 0, 14);
+ my $port = sub($id, $address) {
+ my $P = "$ROOT/ports/$id";
+ return join(
+ ' && ',
+ "mkdir $P",
+ w_text("$P/addr_trtype", 'tcp'),
+ w_text("$P/addr_adrfam", 'ipv4'),
+ w_text("$P/addr_traddr", $address),
+ w_text("$P/addr_trsvcid", 4420),
+ );
+ };
+ my $link = sub($id) {
+ return join(
+ ' && ',
+ "grep -qx 0 $S/attr_allow_any_host",
+ (map { "test -L $S/allowed_hosts/$_" } @HOSTS),
+ "ln -s $S $ROOT/ports/$id/subsystems/$NQN",
+ );
+ };
+
+ my $fake = FakeTarget->new(pools => ['tank']);
+ $fake->add_zvol('tank/vm-100-disk-0', identity => [$NQN, 1, $U{1}]);
+ is_deeply(
+ rendered(flow($fake, $ACT{activate})),
+ [
+ "$INVENTORY_TEXT && $MARKER_TEXT && $CONFIGFS_KEYS_TEXT",
+ "modprobe nvmet_tcp && $INVENTORY_TEXT && $MARKER_TEXT && $CONFIGFS_KEYS_TEXT",
+ join(
+ ' && ',
+ "mkdir $S",
+ w_text("$S/attr_model", q{'Proxmox ZFS NVMe'}),
+ w_text("$S/attr_serial", $serial),
+ w_text("$S/attr_allow_any_host", 0),
+ $port->(1, '192.0.2.21'),
+ $port->(2, '192.0.2.22'),
+ "mkdir $H[0]",
+ "mkdir $H[1]",
+ "chmod 0600 $H[0]/dhchap_key $H[0]/dhchap_ctrl_key $H[1]/dhchap_key"
+ . " $H[1]/dhchap_ctrl_key",
+ "dd ibs=4096 obs=8192 2>/dev/null | tee $H[0]/dhchap_key $H[1]/dhchap_key"
+ . ' >/dev/null',
+ "ln -s $H[0] $S/allowed_hosts/$HOSTS[0]",
+ "ln -s $H[1] $S/allowed_hosts/$HOSTS[1]",
+ build_text(1, $U{1}, '/dev/zvol/tank/vm-100-disk-0'),
+ ),
+ "$INVENTORY_TEXT && $MARKER_TEXT && $CONFIGFS_KEYS_TEXT",
+ $link->(1),
+ $link->(2),
+ ],
+ 'read, configure with a key from stdin, verify and publish',
+ );
+ $fake = lifecycle_fake();
+ is_deeply(
+ rendered(flow($fake, sub { $PLUGIN->path(scfg(), 'vm-100-disk-0', 'st') })),
+ [$INVENTORY_TEXT],
+ 'the ZFS inventory',
+ );
+ my $res = flow($fake, sub { $PLUGIN->alloc_image('st', scfg(), 200, 'raw', undef, 1024) });
+ my $uuid = $fake->{m}->{ds}->{'tank/vm-200-disk-0'}->{props}->{'proxmox:nvme-uuid'};
+ is_deeply(
+ rendered($res),
+ [
+ 'zfs list -o name,volsize,origin,type -t volume,filesystem -d1 -Hp tank',
+ $LOCKED_STATE_TEXT,
+ "zfs set 'proxmox:nvme-last-nsid=4' tank",
+ "zfs create -s -b 16k -o 'proxmox:nvme-subsys=$NQN' -o 'proxmox:nvme-nsid=4'"
+ . " -o 'proxmox:nvme-uuid=$uuid' -V 1024k tank/vm-200-disk-0",
+ 'test -b /dev/zvol/tank/vm-200-disk-0',
+ build_text(4, $uuid, '/dev/zvol/tank/vm-200-disk-0'),
+ $STATE_TEXT,
+ ],
+ 'read, reserve, create, wait, export and verify',
+ );
+ is_deeply(
+ rendered(flow(
+ $fake,
+ sub { $PLUGIN->volume_resize(scfg(), 'st', 'vm-200-disk-0', 2 * 1024**3, 0) },
+ )),
+ [
+ $LOCKED_STATE_TEXT,
+ "zfs set 'volsize=2097152k' tank/vm-200-disk-0 && "
+ . w_text("$S/namespaces/4/revalidate_size", 1),
+ ],
+ 'read, then resize and revalidate',
+ );
+ is_deeply(
+ rendered(flow($fake, sub { $PLUGIN->free_image('st', scfg(), 'vm-101-disk-0') })),
+ [
+ $LOCKED_STATE_TEXT,
+ join(
+ ' && ',
+ w_text("$S/namespaces/2/enable", 0),
+ "rmdir $S/namespaces/2",
+ 'zfs destroy -r tank/vm-101-disk-0',
+ ),
+ ],
+ 'unexport and destroy',
+ );
+ is_deeply(
+ rendered(flow(
+ $fake,
+ sub { $PLUGIN->volume_snapshot_rollback(scfg(), 'st', 'vm-100-disk-0', 'snap1') },
+ )),
+ [
+ $LOCKED_STATE_TEXT,
+ join(' && ',
+ w_text("$S/namespaces/1/enable", 0),
+ "rmdir $S/namespaces/1",
+ 'zfs rollback tank/vm-100-disk-0@snap1',
+ "zfs set 'proxmox:nvme-subsys=$NQN' 'proxmox:nvme-nsid=1'"
+ . " 'proxmox:nvme-uuid=$U{1}' tank/vm-100-disk-0"),
+ 'test -b /dev/zvol/tank/vm-100-disk-0',
+ build_text(1, $U{1}, '/dev/zvol/tank/vm-100-disk-0'),
+ ],
+ 'unexport, roll back, write the identity and export again',
+ );
+ is_deeply(
+ rendered(flow($fake, sub { $PLUGIN->create_base('st', scfg(), 'vm-100-disk-0') })),
+ [
+ $LOCKED_STATE_TEXT,
+ join(' && ',
+ w_text("$S/namespaces/1/enable", 0),
+ "rmdir $S/namespaces/1",
+ 'zfs rename tank/vm-100-disk-0 tank/base-100-disk-0'),
+ 'test -b /dev/zvol/tank/base-100-disk-0',
+ build_text(1, $U{1}, '/dev/zvol/tank/base-100-disk-0')
+ . ' && zfs snapshot tank/base-100-disk-0@__base__',
+ ],
+ 'unexport, rename, export and snapshot',
+ );
+ $res = flow($fake, sub { $PLUGIN->clone_image(scfg(), 'st', 'base-102-disk-0', 201) });
+ $uuid = $fake->{m}->{ds}->{'tank/vm-201-disk-0'}->{props}->{'proxmox:nvme-uuid'};
+ is_deeply(
+ [map { $_->{rendered} } $res->{calls}->@[2, 3]],
+ [
+ "zfs set 'proxmox:nvme-last-nsid=5' tank",
+ "zfs clone -o 'proxmox:nvme-subsys=$NQN' -o 'proxmox:nvme-nsid=5'"
+ . " -o 'proxmox:nvme-uuid=$uuid' tank/base-102-disk-0\@__base__ tank/vm-201-disk-0",
+ ],
+ 'reserve and clone',
+ );
+
+ $fake = lifecycle_fake(
+ step_fault => sub($s, $c) { step_is($s, 'rmdir', "$S/namespaces/2") ? 'busy' : undef },
+ );
+ $res = flow($fake, sub { $PLUGIN->free_image('st', scfg(), 'vm-101-disk-0') });
+ is(
+ $res->{calls}->[-1]->{rendered},
+ 'test -b /dev/zvol/tank/vm-101-disk-0 && ' . w_text("$S/namespaces/2/enable", 1),
+ 'enable again',
+ );
+
+ $fake = removal_fake();
+ is_deeply(
+ rendered(flow($fake, sub { nv('_nvmet_delete_target', scfg()) })),
+ [
+ $LOCKED_STATE_TEXT,
+ "rm $ROOT/ports/1/subsystems/$NQN $ROOT/ports/2/subsystems/$NQN"
+ . " $S/allowed_hosts/$HOSTS[0] $S/allowed_hosts/$HOSTS[1] && rmdir $S",
+ $STATE_TEXT,
+ "rmdir $H[1]",
+ ],
+ 'teardown and orphan hosts',
+ );
+
+ $fake = FakeTarget->new(pools => ['tank']);
+ $fake->{m}->{mounted} = 0;
+ $res = flow($fake, $ACT{activate});
+ ok(grep({ $_ eq 'cat /proc/mounts' } rendered($res)->@*), 'the mount table');
+ ok(
+ grep({ $_ eq 'mount -t configfs none /sys/kernel/config' } rendered($res)->@*),
+ 'the mount',
+ );
+
+ $fake = target_fake();
+ $fake->add_port(5, '192.0.2.99', 4420, links => [$NQN]);
+ is(
+ rendered(flow($fake, $ACT{activate}))->[-1],
+ "rm $ROOT/ports/5/subsystems/$NQN",
+ 'unlink',
+ );
+
+ $fake = lifecycle_fake();
+ $fake->{m}->{subsystems}->{$NQN}->{attr}->{attr_allow_any_host} = '1';
+ $fake->{m}->{subsystems}->{$NQN}->{acl} = {};
+ is(
+ rendered(flow($fake, $ACT{activate}))->[2],
+ join(
+ ' && ',
+ w_text("$S/attr_allow_any_host", 0),
+ map { "ln -s $ROOT/hosts/$_ $S/allowed_hosts/$_" } @HOSTS,
+ ),
+ 'allow_any_host before the ACLs',
+ );
+
+ $fake = lifecycle_fake();
+ my @generic = (
+ [sub { [$PLUGIN->status('st', scfg())] }, 'zfs get -o value -Hp available,used tank'],
+ [
+ sub { $PLUGIN->volume_snapshot_info(scfg(), 'st', 'vm-100-disk-0') },
+ 'zfs list -Hp -r -t snapshot -o name,guid,creation tank/vm-100-disk-0',
+ ],
+ [
+ sub { $PLUGIN->volume_rollback_is_possible(scfg(), 'st', 'vm-100-disk-0', 'snap1') }
+ ,
+ 'zfs list -H -r -t snapshot -o name -s creation tank/vm-100-disk-0',
+ ],
+ [
+ sub { [$PLUGIN->volume_size_info(scfg(), 'st', 'vm-100-disk-0')] },
+ 'zfs get -o value -Hp volsize,usedbydataset tank/vm-100-disk-0',
+ ],
+ [
+ sub { $PLUGIN->volume_snapshot(scfg(), 'st', 'vm-100-disk-0', 'snap2') },
+ 'zfs snapshot tank/vm-100-disk-0@snap2',
+ ],
+ [
+ sub { $PLUGIN->volume_snapshot_delete(scfg(), 'st', 'vm-100-disk-0', 'snap2') },
+ 'zfs destroy tank/vm-100-disk-0@snap2',
+ ],
+ );
+
+ for my $case (@generic) {
+ is(rendered(flow($fake, $case->[0]))->[0], $case->[1], $case->[1]);
+ }
+ is(
+ rendered(flow($fake, sub { $PLUGIN->list_images('st', scfg()) }))->[1],
+ 'zfs get -H -d 1 -o name,value,source proxmox:nvme-subsys tank',
+ 'the ownership query',
+ );
+};
+
+sub sh($command, $input = undef) {
+ my $err = gensym;
+ my $pid = open3(my $in, my $out, $err, '/bin/sh', '-c', $command);
+ print {$in} $input if defined($input);
+ close($in);
+ my $stdout = do { local $/; <$out> }
+ // '';
+ my $stderr = do { local $/; <$err> }
+ // '';
+ waitpid($pid, 0) == $pid or die "wait for test shell: $!\n";
+ my $status = $?;
+ return ($status & 127 ? 128 + ($status & 127) : $status >> 8, $stdout, $stderr);
+}
+
+# The subtests that run rendered commands need these tools. A Debian build
+# has them, so a missing one is an error, never a silently skipped test.
+{
+ my @tools = qw(find grep sha256sum dd tee);
+ BAIL_OUT('needs /bin/sh with ' . join(', ', @tools))
+ if !-x '/bin/sh' || (sh('command -v ' . join(' ', @tools)))[0] != 0;
+}
+
+subtest 'quoting' => sub {
+ my $dir = tempdir(CLEANUP => 1);
+ for my $value (
+ 'a b',
+ q{it's},
+ "\$(touch $dir/pwned)",
+ "`touch $dir/pwned`",
+ '-n',
+ '*',
+ '%s\n',
+ q{"x"},
+ 'x;y',
+ 'a && b',
+ '|',
+ '~root',
+ "tab\there",
+ '',
+ '$HOME',
+ 'a\\b',
+ '#c',
+ ) {
+ my $file = "$dir/value";
+ my ($rc) = sh(nv('_nvmet_render', [write_step($file, $value)]));
+ is($rc . slurp($file), "0$value\n", 'a write keeps ' . ($value =~ s/\t/\\t/r));
+ my (undef, $out) = sh(nv('_nvmet_render', [['printf', '%s\n', $value]]));
+ is($out, "$value\n", 'an argument keeps ' . ($value =~ s/\t/\\t/r));
+ }
+ ok(!-e "$dir/pwned", 'no value is executed');
+};
+
+subtest 'renderer rejections' => sub {
+ for my $case (
+ ['a newline in a value', [write_step("$S/x", "a\nb")]],
+ ['a NUL in a value', [write_step("$S/x", "a\0b")]],
+ ['a CR in an argument', [['zfs', 'get', "a\rb"]]],
+ ['an undefined argument', [['zfs', 'get', undef]]],
+ ['an undefined value', [write_step("$S/x", undef)]],
+ ['a reference as argument', [['zfs', 'get', []]]],
+ ['a relative write', [write_step('x', 1)]],
+ ['a relative key path', [{ key => ['x'] }]],
+ ['an empty key step', [{ key => [] }]],
+ ['two key steps', [{ key => ['/a'] }, { key => ['/b'] }]],
+ ['bash', [['bash', '-c', 'true']]],
+ ['sh', [['sh', '-c', 'true']]],
+ ['perl', [['perl', '-e', '1']]],
+ ['an absolute command', [['/bin/rm', '/a']]],
+ ['an empty command', [[]]],
+ ['a write with extra keys', [{ write => '/a', value => 1, mode => 1 }]],
+ ['an unknown step', [{ run => '/a' }]],
+ ['a string step', ['mkdir /a']],
+ ['no steps', []],
+ ['a string', 'mkdir /a'],
+ ) {
+ eval { nv('_nvmet_render', $case->[1]) };
+ is($@, "internal error: invalid NVMe target step\n", "rejects $case->[0]");
+ }
+};
+
+subtest 'the rendered chains on a configfs stand-in' => sub {
+ my $dir = tempdir(CLEANUP => 1);
+ my $root = "$dir/nvmet";
+ mkdir($_) or die "mkdir $_: $!\n" for $root, map { "$root/$_" } qw(hosts ports subsystems);
+ my $device = "$dir/device";
+ symlink('/dev/null', $device) or die "symlink: $!\n";
+ my $box = sub($path) { $path =~ s{\A\Q$ROOT\E(?=/|\z)}{$root}r };
+ my $kind = sub($path) {
+ my $rel = substr($path, length($root) + 1);
+ return 'subsystem' if $rel =~ m{\Asubsystems/[^/]+\z};
+ return 'namespace' if $rel =~ m{\Asubsystems/[^/]+/namespaces/[0-9]+\z};
+ return 'port' if $rel =~ m{\Aports/[0-9]+\z};
+ return 'host' if $rel =~ m{\Ahosts/[^/]+\z};
+ die "unexpected directory $path\n";
+ };
+ my %attrs = (
+ subsystem => {
+ attr_model => 'Linux',
+ attr_serial => '0123456789abcdef',
+ attr_allow_any_host => 0,
+ },
+ namespace => {
+ enable => 0,
+ device_path => '(null)',
+ device_uuid => $U{9},
+ buffered_io => 0,
+ },
+ port => { addr_trtype => '', addr_adrfam => '', addr_traddr => '', addr_trsvcid => '' },
+ host => { dhchap_key => '', dhchap_ctrl_key => '' },
+ );
+ my %groups = (subsystem => ['namespaces', 'allowed_hosts'], port => ['subsystems']);
+ # configfs creates attributes and default groups with each directory and
+ # removes them with it; the stand-in does both with the same commands.
+ my $shim = sub($steps) {
+ my @out;
+ for my $step ($steps->@*) {
+ if (ref($step) eq 'HASH' && exists($step->{write})) {
+ push @out, write_step($box->($step->{write}), $step->{value});
+ } elsif (ref($step) eq 'HASH') {
+ push @out, { key => [map { $box->($_) } $step->{key}->@*] };
+ } elsif ($step->[0] eq 'test' && $step->[1] eq '-b') {
+ push @out, ['test', '-L', $device];
+ } elsif ($step->[0] eq 'mkdir') {
+ for my $path (map { $box->($_) } $step->@[1 .. $step->$#*]) {
+ my $type = $kind->($path);
+ push @out, ['mkdir', $path, map { "$path/$_" } ($groups{$type} // [])->@*];
+ push @out, map { write_step("$path/$_", $attrs{$type}->{$_}) }
+ sort keys $attrs{$type}->%*;
+ }
+ } elsif ($step->[0] eq 'rmdir') {
+ for my $path (map { $box->($_) } $step->@[1 .. $step->$#*]) {
+ my $type = $kind->($path);
+ push @out, ['rm', map { "$path/$_" } sort keys $attrs{$type}->%*];
+ push @out, ['rmdir', map { "$path/$_" } $groups{$type}->@*]
+ if $groups{$type};
+ push @out, ['rmdir', $path];
+ }
+ } else {
+ push @out, [map { $box->($_) } $step->@*];
+ }
+ }
+ return \@out;
+ };
+ my $run = sub($steps, $input = undef) {
+ my ($rc, undef, $err) = sh(nv('_nvmet_render', $shim->($steps)), $input);
+ return ($rc, $err);
+ };
+
+ my $capture = FakeTarget->new(pools => ['tank']);
+ flow($capture, $ACT{activate});
+ my ($find) = grep { step_is($_, 'env') } $capture->{calls}->[0]->{steps}->@*;
+ my $read = sub () {
+ my ($rc, $out, $err) = sh(nv('_nvmet_render', $shim->([$find])));
+ die "configfs read failed: $err\n" if $rc;
+ $out =~ s{\Q$root\E}{$ROOT}g;
+ return nv('_nvmet_parse_configfs', $out);
+ };
+ my $inv = inventory(@POOL, @ONE);
+ my $conf = { nqn => $NQN, portals => $PORTALS, hostnqns => [@HOSTS], keysha => $KEY_SHA };
+ my $plan = nv('_nvmet_plan_activation', $inv, $read->(), $conf);
+ my ($rc, $err) = $run->([map { $_->@* } $plan->{prepublish}->@*], "$KEY\n");
+ is($rc, 0, 'the whole configuration runs as one chain') or diag($err);
+ my $cfs = $read->();
+ is_deeply(
+ nv('_nvmet_plan_activation', $inv, $cfs, $conf)->{prepublish},
+ [],
+ 'and converges',
+ );
+ is_deeply(
+ [map { $cfs->{hosts}->{$_}->{key_sha256} } @HOSTS],
+ [$KEY_SHA, $KEY_SHA],
+ 'the digests read back equal sha256 of the key and a newline',
+ );
+ is_deeply(
+ [map { (stat("$root/hosts/$_/dhchap_key"))[2] & 07777 } @HOSTS],
+ [0600, 0600],
+ 'and the key attributes are readable by their owner only',
+ );
+
+ # Only the kernel's own groups are skipped: the ACL of another tool's
+ # subsystem named like one of them is read and protects the key.
+ my $other = "$root/subsystems/referrals";
+ mkdir($_) or die "mkdir $_: $!\n" for $other, "$other/allowed_hosts", "$other/namespaces";
+ for my $attr (keys $attrs{subsystem}->%*) {
+ open(my $fh, '>', "$other/$attr") or die "open: $!\n";
+ print {$fh} "$attrs{subsystem}->{$attr}\n";
+ close($fh);
+ }
+ symlink("$root/hosts/$HOSTS[0]", "$other/allowed_hosts/$HOSTS[0]") or die "symlink: $!\n";
+ mkdir($_)
+ or die "mkdir $_: $!\n"
+ for "$root/ports/1/referrals", "$root/ports/1/referrals/r";
+ my (undef, $raw) = sh(nv('_nvmet_render', $shim->([$find])));
+ ok(
+ $read->()->{subsystems}->{referrals}->{acl}->{ $HOSTS[0] },
+ 'a subsystem named referrals',
+ );
+ unlike($raw, qr{/ports/1/referrals}, 'while the referrals of a port are skipped');
+ unlink("$other/allowed_hosts/$HOSTS[0]", map { "$other/$_" } keys $attrs{subsystem}->%*);
+ rmdir($_)
+ or die "rmdir $_: $!\n"
+ for "$other/allowed_hosts", "$other/namespaces", $other, "$root/ports/1/referrals/r",
+ "$root/ports/1/referrals";
+
+ for my $unit (nv('_nvmet_plan_publish', $cfs, $NQN, $plan->{port_ids}, [@HOSTS])->@*) {
+ ($rc) = $run->($unit->{steps});
+ is($rc, 0, "the guarded link to port $unit->{port}");
+ }
+ is_deeply(nv('_nvmet_plan_publish', $read->(), $NQN, [1, 2], [@HOSTS]), [], 'publishes');
+
+ unlink("$root/ports/1/subsystems/$NQN", "$root/subsystems/$NQN/allowed_hosts/$HOSTS[1]");
+ my ($unit) = nv('_nvmet_plan_publish', $read->(), $NQN, [1], [@HOSTS])->@*;
+ ($rc) = $run->($unit->{steps});
+ ok($rc && !-l "$root/ports/1/subsystems/$NQN", 'the guard blocks a link without every ACL');
+ symlink("$root/hosts/$HOSTS[1]", "$root/subsystems/$NQN/allowed_hosts/$HOSTS[1]") or die;
+ ($rc) = $run->([write_step("$ROOT/subsystems/$NQN/attr_allow_any_host", 1)]);
+ ($rc) = $run->($unit->{steps});
+ ok($rc && !-l "$root/ports/1/subsystems/$NQN", 'and while any host may connect');
+ $run->([write_step("$ROOT/subsystems/$NQN/attr_allow_any_host", 0)]);
+ ($rc) = $run->($unit->{steps});
+ ok(!$rc && -l "$root/ports/1/subsystems/$NQN", 'but links once both hold');
+
+ my ($units) = nv('_nvmet_plan_unexport', $read->(), $NQN, $U{1});
+ ($rc) = $run->([map { $_->@* } $units->@*]);
+ ok(!$rc && !-e "$root/subsystems/$NQN/namespaces/1", 'the unexport chain removes it');
+ my ($teardown, $candidates) = nv(
+ '_nvmet_plan_delete_target', inventory(@POOL), $read->(), $NQN, [@HOSTS],
+ );
+ ($rc) = $run->([map { $_->@* } $teardown->@*]);
+ ok(!$rc && !-e "$root/subsystems/$NQN", 'the teardown chain removes the subsystem');
+ ($rc) = $run->([map { $_->@* } nv('_nvmet_plan_orphan_hosts', $read->(), $candidates)->@*]);
+ ok(!$rc && !-e "$root/hosts/$HOSTS[0]" && !-e "$root/hosts/$HOSTS[1]", 'and the hosts');
+ ok(-d "$root/ports/1" && -d "$root/ports/2", 'ports stay');
+};
+
+# Last: every command rendered during this test run, and every protocol
+# violation recorded (by a fake target: a key on a command line or in output,
+# a change without the lock; or a local command).
+subtest 'every rendered command is a POSIX command chain' => sub {
+ is_deeply(\@VIOLATIONS, [], 'no flow violated the target protocol');
+ my %shapes;
+ for my $command (keys %CORPUS) {
+ (my $shape = $command) =~ s/[0-9a-f]{8}(?:-[0-9a-f]{4}){3}-[0-9a-f]{12}/U/gi;
+ $shape =~ s/[0-9]+/N/g;
+ $shapes{$shape} //= $command;
+ }
+ ok(keys(%shapes) >= 40, 'the corpus covers ' . keys(%shapes) . ' command shapes');
+ my $meta = qr/[;&|<>`\$(){}\\*?\[\]~#!"]/;
+ my @bad;
+ for my $command (sort values %shapes) {
+ push @bad, "sh -n: $command" if system('/bin/sh', '-n', '-c', $command) != 0;
+ push @bad, "bash-ism: $command"
+ if $command =~ /\[\[|\$\(\(|<<<|\$'|pipefail|(?:\A|&& )(?:function|source|\.) /;
+ (my $bare = $command) =~ s/'[^']*'|\\'//g;
+ for my $segment (split / && /, $bare) {
+ next
+ if $segment =~ m{\Add\ ibs=4096\ obs=8192\ 2>/dev/null\ \|\ tee(?:\ /[^\s;&|<>]+)+
+ \ >/dev/null\z}x;
+ if ($segment =~ /\Aprintf +(\S*) +> (\S+)\z/) {
+ push @bad, "write: $segment" if "$1$2" =~ $meta;
+ next;
+ }
+ push @bad, "shell syntax: $segment" if $segment =~ $meta;
+ }
+ }
+ is_deeply(\@bad, [], 'every command is a plain POSIX command chain');
+};
+
+done_testing();
diff --git a/src/test/zfsnvme_test.pm b/src/test/zfsnvme_test.pm
new file mode 100644
index 00000000..3b311bfa
--- /dev/null
+++ b/src/test/zfsnvme_test.pm
@@ -0,0 +1,2353 @@
+package PVE::Storage::TestZFSNVMe;
+
+# Storage API and local NVMe host side of the zfsnvme plugin. The target side
+# is covered by zfsnvme_target_test.pm. Nothing here opens a connection or
+# changes the host outside a temporary directory: every probe of the local
+# host is mocked, and the children that the plugin forks for its connect and
+# delete steps run inline, except where the child itself is tested.
+
+use v5.36;
+
+use lib qw(..);
+
+use Compress::Zlib qw(crc32);
+use Config;
+use Errno qw(EACCES ECONNREFUSED EINVAL ENOENT);
+use Fcntl qw(O_NOCTTY O_RDWR);
+use File::Temp qw(tempdir);
+use MIME::Base64 qw(encode_base64);
+use POSIX qw(
+ ECHO ECHONL ICANON OPOST SIG_BLOCK SIGALRM SIGHUP SIGINT SIGQUIT SIGTERM TCSANOW WNOHANG
+ sigprocmask
+);
+use Test::MockModule;
+use Test::More;
+
+use PVE::Storage;
+use PVE::Storage::ZFSNVMePlugin;
+
+my $PLUGIN = 'PVE::Storage::ZFSNVMePlugin';
+my $NQN = 'nqn.2026-01.com.example:test';
+my $OTHER_NQN = 'nqn.2026-01.com.example:other';
+my $UUID = '12345678-1234-1234-1234-123456789abc';
+my $ROOT = '/sys/kernel/config/nvmet';
+my $hostnqn_a = 'nqn.2014-08.org.nvmexpress:uuid:00000000-0000-4000-8000-000000000001';
+my $hostnqn_b = 'nqn.2014-08.org.nvmexpress:uuid:00000000-0000-4000-8000-000000000002';
+
+sub make_key($hash, $length, $seed = 1) {
+ my $raw = join('', map { chr(($_ * 131 + $seed * 17) % 256) } 1 .. $length);
+ return "DHHC-1:$hash:" . encode_base64($raw . pack('V', crc32($raw)), '') . ':';
+}
+my $KEY = make_key('00', 32);
+my $KEY_B = make_key('00', 32, 2);
+
+sub scfg(%override) {
+ return {
+ type => 'zfsnvme',
+ server => '192.0.2.10',
+ pool => 'tank',
+ subsysnqn => $NQN,
+ 'nvme-portals' => '192.0.2.21,192.0.2.22',
+ 'nvme-host-ifaces' => 'ens19,ens20',
+ 'nvme-host-nqns' => "$hostnqn_a,$hostnqn_b",
+ %override,
+ };
+}
+
+# Calls a plugin package sub by name, so a mocked sub is the one called.
+sub nv($name, @args) {
+ my $code = $PLUGIN->can($name) // die "plugin has no sub '$name'\n";
+ return $code->(@args);
+}
+
+# ---------------------------------------------------------------------------
+# Mocks. Target calls are answered by $TARGET, the ssh command of the runner
+# by $COMMAND; sysfs and key files are read from %FILES and listed from %DIRS,
+# block devices are %BLOCK and symbolic links %LINKS. The children of the
+# bounded host steps run inline and record their bound in @FORKS. Waits
+# advance $NOW. The host side runs no external command: only the runner
+# tests, which set $COMMAND, may reach run_command.
+# ---------------------------------------------------------------------------
+
+our ($TARGET, $COMMAND, $NOW);
+our (%FILES, %DIRS, %CONFIG, %BLOCK, %LINKS, @TARGET_CALLS, @WARNINGS, @WRITES);
+our (@SYSFS_WRITES, @GLOBS, @CONNECTS, @ACTIVATIONS, @UNLINKED, @RESTRICTED, @DELETED, @FORKS);
+our (%REFUSE, %CONNECT_STATE); # address => the connect is refused, the new controller's state
+$NOW = 1_000_000;
+
+my $nvme_mock = Test::MockModule->new($PLUGIN);
+my $file_mock = Test::MockModule->new('PVE::File');
+my $sysfs_mock = Test::MockModule->new('PVE::SysFSTools');
+my $tools_mock = Test::MockModule->new('PVE::Tools');
+my $network_mock = Test::MockModule->new('PVE::Network');
+my $storage_mock = Test::MockModule->new('PVE::Storage');
+my $parent_mock = Test::MockModule->new('PVE::Storage::Plugin');
+
+$nvme_mock->redefine(
+ _nvmet_run => sub($scfg, $steps, %opts) {
+ push @TARGET_CALLS, { steps => $steps, %opts };
+ die "unexpected NVMe target call '$opts{op}'\n" if !$TARGET;
+ return $TARGET->($steps, %opts);
+ },
+);
+# A command is also recorded, in case the caller handles the error; the last
+# test checks that there was none.
+our @COMMANDS;
+my $no_command = sub($cmd, %opts) {
+ push @COMMANDS, $cmd;
+ die "test error: unexpected command '" . (ref($cmd) ? $cmd->[0] : $cmd) . "'\n";
+};
+$nvme_mock->redefine(
+ run_command => sub($cmd, %opts) {
+ return $COMMAND ? $COMMAND->($cmd, %opts) : $no_command->($cmd, %opts);
+ },
+);
+$tools_mock->redefine(run_command => $no_command);
+$nvme_mock->redefine(_now => sub () { return $NOW });
+$nvme_mock->redefine(_sleep => sub($seconds) { $NOW += $seconds; return });
+$nvme_mock->redefine(_block_device => sub($path) { return $BLOCK{$path} });
+$nvme_mock->redefine(_link_target => sub($path) { return $LINKS{$path} });
+$nvme_mock->redefine(_restrict_attr => sub($path) { push @RESTRICTED, $path; return });
+$nvme_mock->redefine(_unlink_file => sub($path) { push @UNLINKED, $path; return 1 });
+$nvme_mock->redefine(log_warn => sub($message) { push @WARNINGS, $message; return });
+$nvme_mock->redefine(file_read_firstline => sub($path) { return $FILES{$path} });
+$nvme_mock->redefine(
+ file_set_contents => sub($path, $data, $perm = undef, @rest) {
+ push @WRITES, { path => $path, data => $data, perm => $perm };
+ return;
+ },
+);
+$nvme_mock->redefine(make_path => sub(@args) { push @WRITES, { make_path => [@args] }; return });
+$nvme_mock->redefine(_local_iface_exists => sub($iface) { return 1 });
+
+# The connect child: creates a controller of the portal named in the options.
+my $connect_mock = sub($options, $names) {
+ my %option = map { split(/=/, $_, 2) } split(/,/, $options);
+ push @CONNECTS, "$option{traddr}\@$option{host_iface}";
+ die "connect failed: Connection refused\n" if $REFUSE{ $option{traddr} };
+ my $instance = 70 + scalar(@CONNECTS);
+ add_controller(
+ "nvme$instance",
+ $option{traddr},
+ $option{trsvcid},
+ $option{host_iface},
+ $CONNECT_STATE{ $option{traddr} } // 'live',
+ );
+ return $instance;
+};
+$nvme_mock->redefine(_fabrics_connect => $connect_mock);
+# The deletion child: the controllers disappear from sysfs.
+my $delete_mock = sub($controllers) {
+ for my $controller ($controllers->@*) {
+ push @DELETED, $controller->{name};
+ $DIRS{'/sys/class/nvme'} =
+ [grep { $_ ne $controller->{name} } ($DIRS{'/sys/class/nvme'} // [])->@*];
+ }
+ return scalar($controllers->@*);
+};
+$nvme_mock->redefine(_delete_controllers => $delete_mock);
+my $fork_inline = sub($timeout, $code, $opts = undef) {
+ push @FORKS, $timeout;
+ my $res = $code->();
+ return wantarray ? ($res, 0) : $res;
+};
+$tools_mock->redefine(run_fork_with_timeout => $fork_inline);
+$file_mock->redefine(
+ dir_glob_foreach => sub($dir, $regex, $func) {
+ push @GLOBS, [$dir, $regex];
+ for my $entry (($DIRS{$dir} // [])->@*) {
+ if (my @res = $entry =~ m/^($regex)$/) {
+ $func->(@res);
+ }
+ }
+ },
+);
+my $sysfs_write = sub($path, $data, $allow_existing = undef) {
+ push @SYSFS_WRITES, [$path, $data];
+ $FILES{$path} = $data =~ s/\n\z//r;
+ return 1;
+};
+$sysfs_mock->redefine(file_write => $sysfs_write);
+$network_mock->redefine(tcp_ping => sub($host, $port, $timeout = undef) { return 1 });
+$storage_mock->redefine(config => sub () { return { ids => {%CONFIG} } });
+
+# A local controller of $NQN, as nvme-core shows it in sysfs.
+sub add_controller($name, $address, $port, $iface, $state = 'live', $nqn = $NQN) {
+ my $base = "/sys/class/nvme/$name";
+ push $DIRS{'/sys/class/nvme'}->@*, $name;
+ $FILES{"$base/subsysnqn"} = $nqn;
+ $FILES{"$base/state"} = $state;
+ $FILES{"$base/address"} = "traddr=$address,trsvcid=$port,host_iface=$iface";
+ $FILES{"$base/reconnect_delay"} //= '2';
+ $FILES{"$base/ctrl_loss_tmo"} //= '600';
+ $FILES{"$base/fast_io_fail_tmo"} //= 'off';
+ return;
+}
+
+sub reset_host() {
+ %DIRS = ('/sys/class/nvme-subsystem' => ['nvme-subsys7']);
+ %FILES = (
+ '/sys/module/nvme_core/parameters/multipath' => 'Y',
+ '/etc/nvme/hostnqn' => $hostnqn_a,
+ '/etc/nvme/hostid' => '12345678-1234-1234-1234-123456789abc',
+ '/sys/class/nvme-subsystem/nvme-subsys7/subsysnqn' => $NQN,
+ );
+ %BLOCK = %LINKS = %CONNECT_STATE = ();
+ @CONNECTS = @ACTIVATIONS = @SYSFS_WRITES = @WRITES = @WARNINGS = @GLOBS = ();
+ @RESTRICTED = @UNLINKED = @DELETED = @FORKS = ();
+ $NOW += 61; # a new minute: no storage is in its slow-path backoff
+ return;
+}
+
+# ---------------------------------------------------------------------------
+# Volume names and features (same formats as the other ZFS plugins)
+# ---------------------------------------------------------------------------
+
+# [feature, volume, snapshot, running, expected]
+for my $case (
+ ['copy', 'base-100-disk-0', undef, 0, 1],
+ ['copy', 'vm-100-disk-0', undef, 0, 1],
+ ['copy', 'vm-100-disk-0', undef, 1, 1],
+ ['copy', 'base-100-disk-0/vm-101-disk-0', undef, 0, 1],
+ ['copy', 'vm-100-disk-0', 'snap', 0, undef],
+ ['clone', 'base-100-disk-0', undef, undef, 1],
+ ['clone', 'vm-100-disk-0', undef, undef, undef],
+ ['template', 'vm-100-disk-0', undef, undef, 1],
+ ['snapshot', 'vm-100-disk-0', undef, 0, 1],
+ ['snapshot', 'vm-100-disk-0', 'snap', 0, 1],
+) {
+ my ($feature, $volname, $snap, $running, $expected) = $case->@*;
+ my @args = defined($running) ? ($snap, $running, {}) : ();
+ is(
+ $PLUGIN->volume_has_feature({}, $feature, 'nvmetest', $volname, @args),
+ $expected,
+ "$feature of $volname"
+ . ($snap ? '@' . $snap : '')
+ . ($running ? ' while running' : ''),
+ );
+}
+ok(!$PLUGIN->storage_can_replicate({}, 'nvmetest'), 'no storage replication');
+
+for my $case (
+ ['vm-100-disk-0', ['images', 'vm-100-disk-0', 100, undef, undef, '', 'raw']],
+ ['base-100-disk-0', ['images', 'base-100-disk-0', 100, undef, undef, 1, 'raw']],
+ ['subvol-100-disk-0', ['images', 'subvol-100-disk-0', 100, undef, undef, '', 'subvol']],
+ ['basevol-100-disk-0', ['images', 'basevol-100-disk-0', 100, undef, undef, 1, 'subvol']],
+ [
+ 'base-100-disk-0/vm-200-disk-0',
+ ['images', 'vm-200-disk-0', 200, 'base-100-disk-0', 100, '', 'raw'],
+ ],
+ [
+ 'basevol-100-disk-0/subvol-200-disk-0',
+ ['images', 'subvol-200-disk-0', 200, 'basevol-100-disk-0', 100, '', 'subvol'],
+ ],
+) {
+ my ($name, $expected) = $case->@*;
+ is_deeply([$PLUGIN->parse_volname($name)], $expected, "parse_volname keeps the tuple of $name");
+}
+for my $name ('not-a-volume', 'vm-no-id-disk-0', 'vm-100-disk 0', 'vm-100-') {
+ eval { $PLUGIN->parse_volname($name) };
+ like($@, qr/unable to parse zfs volume name/, "rejects invalid volume name $name");
+}
+
+is_deeply(
+ \@PVE::Storage::ZFSNVMePlugin::ISA,
+ ['PVE::Storage::Plugin'],
+ 'the backend inherits directly from the storage plugin base',
+);
+for my $method (qw(
+ zfs_request zfs_get_lu_name zfs_create_lu zfs_resize_lu zfs_list_zvol zfs_parse_zvol_list
+ zfs_get_properties zfs_get_pool_stats zfs_create_zvol zfs_delete_zvol zfs_get_base
+))
+{
+ ok(!$PLUGIN->can($method), "the ZFS helper $method is private");
+}
+
+# ---------------------------------------------------------------------------
+# Configuration
+# ---------------------------------------------------------------------------
+
+is(nv('verify_nvme_nqn', $hostnqn_a), $hostnqn_a, 'accepts a standard NQN');
+ok(!nv('verify_nvme_nqn', 'not-an-nqn', 1), 'rejects an invalid NQN');
+ok(!nv('verify_nvme_nqn', 'nqn.x:' . ('a' x 220), 1), 'rejects a long NQN');
+is_deeply(
+ nv('parse_nvme_host_nqns', "$hostnqn_a, $hostnqn_b"),
+ [$hostnqn_a, $hostnqn_b],
+ 'parses the host allow-list',
+);
+ok(!nv('parse_nvme_host_nqns', "$hostnqn_a,$hostnqn_a", 1), 'rejects duplicate host NQNs');
+ok(!nv('parse_nvme_host_nqns', '', 1), 'requires a host NQN');
+
+is_deeply(
+ nv('parse_nvme_portals', '198.51.100.11:4420,[2001:db8::11]:4421,198.51.100.12'),
+ [
+ { address => '198.51.100.11', port => 4420, family => 'ipv4' },
+ { address => '2001:db8::11', port => 4421, family => 'ipv6' },
+ { address => '198.51.100.12', port => 4420, family => 'ipv4' },
+ ],
+ 'parses IPv4 and IPv6 portals',
+);
+is_deeply(
+ [
+ map { $_->{address} }
+ nv('parse_nvme_portals', '[2001:DB8:0:0::11],[2001:db8::0:12]:4421')->@*
+ ],
+ ['2001:db8::11', '2001:db8::12'],
+ 'IPv6 addresses are canonical',
+);
+eval { nv('parse_nvme_portals', '[2001:db8::11]:4420,[2001:db8:0::11]:4420'); };
+like($@, qr/duplicate NVMe\/TCP portal/, 'two spellings of one IPv6 portal are a duplicate');
+eval { nv('parse_nvme_portals', '198.51.100.11:4420,198.51.100.11:4420') };
+like($@, qr/duplicate NVMe\/TCP portal/, 'duplicate portal is rejected');
+eval { nv('parse_nvme_portals', '198.51.100.11,198.51.100.11:04420') };
+is($@, "duplicate NVMe/TCP portal '198.51.100.11:04420'\n", 'a port with a leading zero');
+eval { nv('parse_nvme_portals', '198.51.100.11:4421,[::ffff:198.51.100.11]:4421') };
+is(
+ $@,
+ "duplicate NVMe/TCP portal '[::ffff:198.51.100.11]:4421'\n",
+ 'an IPv4-mapped IPv6 address of an IPv4 portal',
+);
+ok(!nv('parse_nvme_portals', '198.51.100.11:70000', 1), 'rejects an invalid port');
+ok(!nv('parse_nvme_portals', '[198.51.100.11]', 1), 'rejects IPv4 in brackets');
+my $too_many = join(',', map { "198.51.100.$_:4420" } 1 .. 17);
+ok(!nv('parse_nvme_portals', $too_many, 1), 'limits the number of paths');
+
+is_deeply(
+ nv(
+ '_configured_portals',
+ {
+ 'nvme-portals' => '198.51.100.11:4420,198.51.100.12:4421',
+ 'nvme-host-ifaces' => 'ens20,ens21',
+ },
+ ),
+ [
+ { address => '198.51.100.11', port => 4420, family => 'ipv4', host_iface => 'ens20' },
+ { address => '198.51.100.12', port => 4421, family => 'ipv4', host_iface => 'ens21' },
+ ],
+ 'binds each portal to its local interface',
+);
+eval {
+ nv(
+ '_configured_portals',
+ {
+ 'nvme-portals' => '198.51.100.11,198.51.100.12',
+ 'nvme-host-ifaces' => 'ens20',
+ },
+ );
+};
+like($@, qr/one interface for each/, 'portal and interface counts must match');
+ok(!nv('parse_nvme_host_ifaces', 'ens20,not/an/interface', 1), 'rejects an invalid interface name');
+for my $name ('.', '..') {
+ ok(!nv('parse_nvme_host_ifaces', "ens20,$name", 1), "rejects the directory name '$name'");
+}
+for my $option ('nvme-host-ifaces', 'nvme-host-nqns') {
+ my $opts = $PLUGIN->options()->{$option};
+ ok(!$opts->{optional} && !$opts->{fixed}, "$option is required, and can be changed");
+ my $config = scfg(type => 'zfsnvme', content => 'images', blocksize => '16k');
+ delete $config->{$option};
+ eval { $PLUGIN->check_config('st', $config, 1, 1) };
+ like($@, qr/missing value for required option '\Q$option\E'/, "a new storage needs $option");
+}
+$nvme_mock->redefine(_local_iface_exists => sub($iface) { return $iface eq 'ens20' });
+eval { nv('_validate_local_ifaces', [{ host_iface => 'ens20' }, { host_iface => 'ens21' }]); };
+like($@, qr/host interface 'ens21' does not exist/, 'a missing local interface fails preflight');
+$nvme_mock->redefine(_local_iface_exists => sub($iface) { return 1 });
+
+my $listener_used = sub($address, $port, $storeid = 'existing') {
+ return "NVMe/TCP portal '$address' port $port is already used by storage '$storeid';"
+ . " give each storage its own address or port\n";
+};
+# [name, the existing storage, our overrides, error]
+for my $case (
+ [
+ 'the NQN of another storage',
+ { server => '192.0.2.11', pool => 'tank/existing', subsysnqn => $NQN },
+ {},
+ qr/NQN is already used by storage 'existing'/,
+ ],
+ [
+ 'the NQN and the listener of another storage',
+ { pool => 'data', subsysnqn => $NQN, 'nvme-portals' => '192.0.2.21' },
+ {},
+ "NVMe subsystem NQN is already used by storage 'existing'\n",
+ ],
+ [
+ 'the pool of another storage',
+ { pool => 'tank' },
+ {},
+ qr/ZFS pool 'tank' on '192\.0\.2\.10' is already used/,
+ ],
+ [
+ 'a pool inside ours',
+ { pool => 'tank/sub' },
+ {},
+ qr/ZFS pool 'tank' on '192\.0\.2\.10' overlaps the pool of storage 'existing'/,
+ ],
+ [
+ 'a pool above ours',
+ { pool => 'tank' },
+ { pool => 'tank/sub' },
+ qr/overlaps the pool of storage 'existing'/,
+ ],
+ [
+ 'a pool above ours on another server',
+ { pool => 'tank', server => '192.0.2.99' },
+ { pool => 'tank/sub' },
+ undef,
+ ],
+ ['a pool around ours', { pool => 'ta' }, {}, undef],
+ ['a parent pool', { pool => 'tan' }, {}, undef],
+ # one NVMe/TCP listener (family, address, port) per storage, on any server
+ [
+ 'the listener of another storage',
+ { pool => 'data', 'nvme-portals' => '192.0.2.31,192.0.2.22' },
+ {},
+ $listener_used->('192.0.2.22', 4420),
+ ],
+ [
+ 'the listener of another storage that names the default port',
+ { pool => 'data', 'nvme-portals' => '192.0.2.21:4420' },
+ {},
+ $listener_used->('192.0.2.21', 4420),
+ ],
+ [
+ 'the listener of a disabled storage',
+ { pool => 'data', disable => 1, 'nvme-portals' => '192.0.2.21' },
+ {},
+ $listener_used->('192.0.2.21', 4420),
+ ],
+ [
+ 'the listener of a storage that spells the server another way',
+ { server => 'nas.example', pool => 'data', 'nvme-portals' => '192.0.2.21' },
+ {},
+ $listener_used->('192.0.2.21', 4420),
+ ],
+ [
+ 'the listener of a storage on another server',
+ { server => '192.0.2.99', pool => 'tank', 'nvme-portals' => '192.0.2.21' },
+ {},
+ $listener_used->('192.0.2.21', 4420),
+ ],
+ [
+ 'an IPv6 listener written another way',
+ { pool => 'data', 'nvme-portals' => '[2001:db8:0::21]' },
+ { 'nvme-portals' => '[2001:DB8::0021]:4420' },
+ $listener_used->('2001:db8::21', 4420),
+ ],
+ [
+ 'an IPv4 listener reached through an IPv4-mapped IPv6 address',
+ { pool => 'data', 'nvme-portals' => '[::ffff:192.0.2.21]' },
+ {},
+ $listener_used->('::ffff:192.0.2.21', 4420),
+ ],
+ [
+ 'an IPv4-mapped IPv6 address of an IPv4 listener',
+ { pool => 'data', 'nvme-portals' => '192.0.2.21' },
+ { 'nvme-portals' => '[::FFFF:c000:215]' },
+ $listener_used->('192.0.2.21', 4420),
+ ],
+ [
+ 'the same addresses on other ports',
+ { pool => 'data', 'nvme-portals' => '192.0.2.21:4421,192.0.2.22:4421' },
+ {},
+ undef,
+ ],
+ [
+ 'the same address on another port through another spelling of the server',
+ {
+ server => 'nas.example',
+ pool => 'data',
+ 'nvme-portals' => '192.0.2.31,192.0.2.22:4421',
+ },
+ {},
+ "storage 'existing' reaches target address '192.0.2.22' through server 'nas.example';"
+ . " use the same server value for both storages\n",
+ ],
+ [
+ 'an IPv6 address written another way through another spelling of the server',
+ { server => 'nas.example', pool => 'data', 'nvme-portals' => '[2001:db8::0:21]:4421' },
+ { 'nvme-portals' => '[2001:db8::21]' },
+ qr/\Astorage 'existing' reaches target address '2001:db8::21' through server/,
+ ],
+ [
+ 'other addresses on another server',
+ { server => 'nas.example', pool => 'tank', 'nvme-portals' => '192.0.2.31,192.0.2.32' },
+ {},
+ undef,
+ ],
+ [
+ 'other addresses on the same server',
+ { pool => 'data', 'nvme-portals' => '192.0.2.31,[2001:db8::22]' },
+ {},
+ undef,
+ ],
+ ['a storage without portals', { pool => 'data' }, {}, undef],
+ # it cannot be activated, and is checked when its portals are fixed
+ [
+ 'a storage whose portals do not parse',
+ { pool => 'data', 'nvme-portals' => '192.0.2.21,192.0.2.300' },
+ {},
+ undef,
+ ],
+ [
+ 'the listener of another storage type',
+ { type => 'zfs', pool => 'tank', 'nvme-portals' => '192.0.2.21' },
+ {},
+ undef,
+ ],
+) {
+ my ($name, $other, $ours, $error) = $case->@*;
+ my $existing = {
+ type => 'zfsnvme',
+ server => '192.0.2.10',
+ subsysnqn => $OTHER_NQN,
+ pool => 'x',
+ $other->%*,
+ };
+ my $cfg = { ids => { existing => $existing } };
+ eval { nv('_assert_unique_target', 'new', scfg($ours->%*), $cfg) };
+ if (ref($error)) {
+ like($@, $error, "refuses $name");
+ } elsif (defined($error)) {
+ is($@, $error, "refuses $name");
+ } else {
+ is($@, '', "accepts $name");
+ }
+}
+
+{
+ my $states = {
+ '198.51.100.11:4420' => [{ state => 'live', host_iface => 'eth0' }],
+ '198.51.100.12:4420' => [
+ { state => 'connecting', host_iface => 'eth1' },
+ { state => 'live', host_iface => 'ens21' },
+ ],
+ '198.51.100.13:4420' => [{ state => 'dead', host_iface => 'ens22' }],
+ };
+ my $portals = [
+ { address => '198.51.100.11', port => 4420, host_iface => 'ens20' },
+ { address => '198.51.100.12', port => 4420, host_iface => 'ens21' },
+ { address => '198.51.100.13', port => 4420, host_iface => 'ens22' },
+ ];
+ is(nv('_live_portal_count', $states, $portals), 1, 'the wrong interface is not healthy');
+ is(nv('_live_portal_count', $states, $portals, 1), 2, 'but it carries I/O');
+}
+
+# DHHC-1:<hash>:<base64 of the key and its little-endian CRC-32>:
+for my $case (['00', 32], ['00', 48], ['00', 64], ['01', 32], ['02', 48], ['03', 64]) {
+ my $key = make_key($case->@*);
+ is(nv('_validate_secret', $key), $key, "accepts a $case->[1] byte key with hash $case->[0]");
+}
+my $bad_crc = do {
+ my $raw = 'k' x 32;
+ 'DHHC-1:00:' . encode_base64($raw . pack('V', crc32($raw) ^ 1), '') . ':';
+};
+for my $case (
+ ['a plaintext secret', 'plaintext'],
+ ['a key without the final colon', substr($KEY, 0, -1)],
+ ['an unknown hash', make_key('04', 32)],
+ ['a 48 byte key for SHA-256', make_key('01', 48)],
+ ['a 32 byte key for SHA-384', make_key('02', 32)],
+ ['a 32 byte key for SHA-512', make_key('03', 32)],
+ ['a 16 byte key', make_key('00', 16)],
+ ['a wrong CRC', $bad_crc],
+ ['truncated base64', substr($KEY, 0, -2) . ':'],
+ ['a key with a trailing newline', "$KEY\n"],
+) {
+ my ($name, $key) = $case->@*;
+ eval { nv('_validate_secret', $key) };
+ is($@, "invalid NVMe DH-HMAC-CHAP key representation\n", "rejects $name without echoing it");
+}
+eval { nv('_validate_secret', undef) };
+is($@, "missing NVMe DH-HMAC-CHAP key\n", 'a missing key');
+# check_config only validates. Defaults are applied where they are used.
+{
+ $parent_mock->redefine(
+ check_config => sub($class, $id, $config, $create, $skip = undef) { return $config },
+ );
+ my $config = scfg();
+ my $copy = { $config->%* };
+ is_deeply($PLUGIN->check_config('st', $config, 1, 1), $copy, 'no defaults on create');
+ is_deeply(
+ $PLUGIN->check_config('st', { 'nvme-host-ifaces' => 'ens20,ens21' }, 0),
+ { 'nvme-host-ifaces' => 'ens20,ens21' },
+ 'a partial update is accepted',
+ );
+ eval { $PLUGIN->check_config('st', scfg('nvme-fast-io-fail-tmo' => 700), 1, 1) };
+ like($@, qr/must not exceed nvme-ctrl-loss-tmo/, 'validated against the default');
+ eval { $PLUGIN->check_config('st', scfg('nvme-host-ifaces' => 'ens19'), 1, 1) };
+ like($@, qr/one interface for each/, 'cross-field validation stays');
+ # the pool is validated like every later use of it validates it
+ is(
+ $PLUGIN->check_config('st', scfg(pool => 'tank/nvme.1'), 1, 1)->{pool},
+ 'tank/nvme.1',
+ 'a nested pool is accepted',
+ );
+ for my $pool ('-rH', 'tank/vm 1', '') {
+ eval { $PLUGIN->check_config('st', scfg(pool => $pool), 1, 1) };
+ is($@, "invalid ZFS pool name\n", "the pool name '$pool' is refused on create");
+ eval { $PLUGIN->check_config('st', { pool => $pool }, 0) };
+ is($@, "invalid ZFS pool name\n", 'and on update');
+ }
+ $parent_mock->unmock('check_config');
+}
+eval {
+ nv(
+ '_validate_fail_fast_timeout',
+ { 'nvme-ctrl-loss-tmo' => 30, 'nvme-fast-io-fail-tmo' => 31 },
+ 600,
+ );
+};
+like($@, qr/must not exceed nvme-ctrl-loss-tmo/, 'fast I/O fail cannot outlive the loss timeout');
+eval {
+ nv(
+ '_validate_fail_fast_timeout',
+ { 'nvme-ctrl-loss-tmo' => -1, 'nvme-fast-io-fail-tmo' => 30 },
+ 600,
+ );
+};
+is($@, '', 'fast I/O fail stays valid with infinite reconnects');
+
+# ---------------------------------------------------------------------------
+# Storage hooks and secrets
+# ---------------------------------------------------------------------------
+
+{
+ my $secret = '/etc/pve/priv/storage/nvmetest.nvme-dhchap';
+ local %CONFIG = ();
+ local %FILES = ();
+ local @WRITES = ();
+ eval { $PLUGIN->on_add_hook('nvmetest', scfg()) };
+ is($@, "missing NVMe DH-HMAC-CHAP key\n", 'a new storage needs a key');
+ eval { $PLUGIN->on_add_hook('nvmetest', scfg(), 'dhchap-key' => $KEY) };
+ is($@, '', 'a new storage stores its key');
+ is_deeply(
+ [grep { $_->{path} } @WRITES],
+ [{ path => $secret, data => "$KEY\n", perm => 0600 }],
+ 'readable by root only',
+ );
+ is_deeply(
+ [map { $_->{make_path} } grep { $_->{make_path} } @WRITES],
+ [['/etc/pve/priv/storage', { mode => 0700 }]],
+ 'in a private directory',
+ );
+
+ local $FILES{$secret} = $KEY;
+ @WRITES = ();
+ eval {
+ $PLUGIN->on_update_hook_full(
+ 'nvmetest',
+ scfg(),
+ { 'nvme-host-ifaces' => 'ens21,ens22' },
+ );
+ };
+ is($@, '', 'a partial update validates against the current configuration');
+ eval { $PLUGIN->on_update_hook_full('nvmetest', scfg(), {}, undef, { 'dhchap-key' => $KEY }) };
+ is($@, '', 'the same key is accepted');
+ eval {
+ $PLUGIN->on_update_hook_full('nvmetest', scfg(), {}, undef, { 'dhchap-key' => $KEY_B });
+ };
+ like($@, qr/NVMe DH-HMAC-CHAP key rotation is not supported/, 'a different key is refused');
+ eval { $PLUGIN->on_update_hook_full('nvmetest', scfg(), { 'nvme-host-nqns' => $hostnqn_a }) };
+ like($@, qr/removing NVMe host NQN '\Q$hostnqn_b\E' is not supported/, 'no host removal');
+ eval { $PLUGIN->on_update_hook_full('nvmetest', scfg(), {}, ['nvme-host-nqns']) };
+ like($@, qr/at least one NVMe host NQN is required/, 'deleting the host list is refused');
+ @WRITES = ();
+ eval {
+ my $update = { 'nvme-host-nqns' => "$hostnqn_a,$hostnqn_b,nqn.x:c" };
+ $PLUGIN->on_update_hook_full('nvmetest', scfg(), $update);
+ };
+ is($@, '', 'adding a host is accepted');
+ is_deeply(\@WRITES, [], 'without storing the key again');
+
+ delete local $FILES{$secret};
+ @WRITES = ();
+ eval { $PLUGIN->on_update_hook_full('nvmetest', scfg(), {}, undef, { 'dhchap-key' => $KEY }) };
+ is($@, '', 'a storage without key file');
+ is_deeply(
+ [grep { $_->{path} } @WRITES],
+ [{ path => $secret, data => "$KEY\n", perm => 0600 }],
+ 'stores the new key',
+ );
+}
+
+# nvmet keeps one key per host NQN: storages on one target that share a host
+# must share the key.
+{
+ my $other = {
+ type => 'zfsnvme',
+ server => '192.0.2.10',
+ pool => 'data',
+ subsysnqn => $OTHER_NQN,
+ 'nvme-host-nqns' => $hostnqn_b,
+ };
+ my $cfg = { ids => { other => $other, nvmetest => scfg() } };
+ local %FILES = ('/etc/pve/priv/storage/other.nvme-dhchap' => $KEY_B);
+ eval { nv('_assert_shared_host_key', 'nvmetest', scfg(), $KEY, $cfg) };
+ is(
+ $@,
+ "storage 'other' uses a different DH-HMAC-CHAP key for the same NVMe host NQNs"
+ . " on '192.0.2.10'\n",
+ 'a different key for a shared host is refused',
+ );
+ for my $case (
+ ['the same key', $KEY_B, {}],
+ ['another server', $KEY, { server => '192.0.2.11' }],
+ ['no shared host', $KEY, { 'nvme-host-nqns' => 'nqn.2026-01.com.example:other-host' }],
+ ['another storage type', $KEY, { type => 'zfs' }],
+ ) {
+ my ($name, $key, $change) = $case->@*;
+ my $ids = { ids => { other => { $other->%*, $change->%* } } };
+ eval { nv('_assert_shared_host_key', 'nvmetest', scfg(), $key, $ids); };
+ is($@, '', "accepts $name");
+ }
+ {
+ local %FILES = ();
+ eval { nv('_assert_shared_host_key', 'nvmetest', scfg(), $KEY, $cfg); };
+ is($@, '', 'a storage without a key file is skipped');
+ }
+
+ local %CONFIG = (other => $other);
+ local @WRITES = ();
+ eval { $PLUGIN->on_add_hook('nvmetest', scfg(), 'dhchap-key' => $KEY) };
+ like($@, qr/storage 'other' uses a different DH-HMAC-CHAP key/, 'on_add_hook refuses it');
+ unlike($@, qr/DHHC-1/, 'without naming a key');
+ is_deeply([grep { $_->{path} } @WRITES], [], 'and stores no key');
+ eval { $PLUGIN->on_add_hook('nvmetest', scfg(), 'dhchap-key' => $KEY_B) };
+ is($@, '', 'the same key is accepted');
+
+ local $FILES{'/etc/pve/priv/storage/nvmetest.nvme-dhchap'} = $KEY;
+ my $current = scfg('nvme-host-nqns' => $hostnqn_a);
+ eval {
+ $PLUGIN->on_update_hook_full(
+ 'nvmetest',
+ $current,
+ { 'nvme-host-nqns' => "$hostnqn_a,$hostnqn_b" },
+ );
+ };
+ like($@, qr/storage 'other' uses a different DH-HMAC-CHAP key/, 'nor shared by an update');
+ eval { $PLUGIN->on_update_hook_full('nvmetest', $current, { 'nvme-iopolicy' => 'numa' }) };
+ is($@, '', 'other updates are accepted');
+}
+
+# Adding and updating a storage check its listeners; an update checks the new
+# portals, and the current definition of the storage itself does not count.
+{
+ my $secret = '/etc/pve/priv/storage/nvmetest.nvme-dhchap';
+ my $other = scfg(
+ subsysnqn => $OTHER_NQN,
+ pool => 'data',
+ 'nvme-portals' => '192.0.2.23',
+ 'nvme-host-ifaces' => 'ens21',
+ );
+ local %CONFIG = (other => $other);
+ local %FILES = ();
+ local @WRITES = ();
+ my $shared = scfg('nvme-portals' => '192.0.2.21,192.0.2.23');
+ eval { $PLUGIN->on_add_hook('nvmetest', $shared, 'dhchap-key' => $KEY) };
+ is($@, $listener_used->('192.0.2.23', 4420, 'other'), 'on_add_hook refuses a used listener');
+ is_deeply([grep { $_->{path} } @WRITES], [], 'and stores no key');
+
+ %CONFIG = (other => $other, nvmetest => scfg());
+ $FILES{$secret} = $KEY;
+ my $update = sub($current, $portals) {
+ eval {
+ $PLUGIN->on_update_hook_full('nvmetest', $current, { 'nvme-portals' => $portals });
+ };
+ return $@;
+ };
+ is($update->(scfg(), '192.0.2.22,192.0.2.21'), '', 'an update keeps its own listeners');
+ is(
+ $update->(scfg(), '192.0.2.21,192.0.2.23:4420'),
+ $listener_used->('192.0.2.23', 4420, 'other'),
+ 'on_update_hook_full refuses a used listener',
+ );
+ is($update->(scfg(), '192.0.2.21,192.0.2.23:4421'), '', 'but not its address on another port');
+ is($update->($shared, '192.0.2.21,192.0.2.22'), '', 'an update can leave a used listener');
+ is_deeply([grep { $_->{path} } @WRITES], [], 'without storing the key again');
+}
+
+# pvesm reads the key from a file, so it never appears on a command line.
+{
+ require PVE::CLI::pvesm;
+
+ my $dir = tempdir(CLEANUP => 1);
+ my $file = sub($name, $content) {
+ open(my $fh, '>', "$dir/$name") or die "open: $!\n";
+ print {$fh} $content;
+ close($fh);
+ return "$dir/$name";
+ };
+ my $padded = $file->('padded', " $KEY \nsecond line\n");
+ my $empty = $file->('empty', "\n");
+ my %map = map {
+ my $command = $_;
+ (
+ $command => (
+ grep { $_->{name} eq 'dhchap-key' }
+ PVE::CLI::pvesm::param_mapping($command)->@*
+ )[0],
+ )
+ } qw(create update);
+ ok($map{create} && $map{update}, 'pvesm create and update map dhchap-key');
+ my $read = $map{create}->{func};
+ eval { $read->($KEY) };
+ is(
+ $@,
+ "dhchap-key expects the path of a file containing the key\n",
+ 'a key in place of the file name is refused without repeating it',
+ );
+ is($read->($padded), $KEY, 'the first line of the file, trimmed');
+ eval { $read->($empty) };
+ like($@, qr/key file '.*' is empty/, 'an empty file is refused');
+ eval { $read->('/dev/null') };
+ like($@, qr/is empty/, 'a file that is not regular, like /dev/stdin, is read');
+ eval { $read->($dir) };
+ like($@, qr/expects the path of a file/, 'a directory is refused');
+}
+
+# ---------------------------------------------------------------------------
+# Volumes through the target runner
+# ---------------------------------------------------------------------------
+
+sub target(%answers) {
+ return sub($steps, %opts) {
+ my $answer = $answers{ $opts{op} } // die "unexpected NVMe target call '$opts{op}'\n";
+ return $answer->($steps, %opts) if ref($answer) eq 'CODE';
+ return { rc => 0, out => $answer, err => '' };
+ };
+}
+
+my @volume_list = (
+ "tank\t-\t-\tfilesystem",
+ "tank/vm-100-disk-0\t1048576\t-\tvolume",
+ "tank/vm-200-disk-0\t1048576\t-\tvolume",
+ "tank/vm-300-disk-0\t1048576\t-\tvolume",
+ "tank/vm-400-disk-0\t1048576\t-\tvolume",
+ "tank/base-90-disk-0\t1048576\t-\tvolume",
+ "tank/vm-101-disk-0\t1048576\ttank/base-90-disk-0\@__base__\tvolume",
+ "tank/vm-102-disk-0\t1048576\ttankXprod/base-90-disk-0\@__base__\tvolume",
+ "tank/subvol-103-disk-0\t-\t-\tfilesystem",
+ "tank/vm-104-disk-0\t-\t-\tfilesystem",
+);
+my @ownership = (
+ "tank\t-\t-",
+ "tank/vm-100-disk-0\t$NQN\tlocal",
+ "tank/vm-200-disk-0\t-\t-",
+ "tank/vm-300-disk-0\t$NQN\tinherited from tank",
+ "tank/vm-400-disk-0\t$NQN\treceived",
+ "tank/base-90-disk-0\t$NQN\tlocal",
+ "tank/vm-101-disk-0\t$NQN\tlocal",
+ "tank/vm-102-disk-0\t$NQN\tlocal",
+ "tank/subvol-103-disk-0\t$OTHER_NQN\tlocal",
+ "tank/vm-104-disk-0\t$NQN\tlocal",
+);
+{
+ local $TARGET = target('zfs list' => \@volume_list, 'zfs get' => \@ownership);
+ my $images = $PLUGIN->list_images('nvmetest', scfg());
+ my %by_volid = map { $_->{volid} => $_ } $images->@*;
+ my %by_name = map { $_->{name} => $_ } $images->@*;
+ is_deeply(
+ [sort keys %by_name],
+ ['base-90-disk-0', 'vm-100-disk-0', 'vm-101-disk-0', 'vm-102-disk-0', 'vm-400-disk-0'],
+ 'listing requires a zvol with a local or received owner',
+ );
+ is(
+ $by_name{'vm-102-disk-0'}->{parent},
+ 'tankXprod/base-90-disk-0@__base__',
+ 'a foreign pool origin is not mistaken for a local parent',
+ );
+ is_deeply(
+ $by_volid{'nvmetest:base-90-disk-0/vm-101-disk-0'},
+ {
+ name => 'vm-101-disk-0',
+ size => 1048576,
+ parent => 'base-90-disk-0@__base__',
+ format => 'raw',
+ vmid => 101,
+ volid => 'nvmetest:base-90-disk-0/vm-101-disk-0',
+ },
+ 'a linked clone keeps the volume ID format of the ZFS plugins',
+ );
+ is_deeply(
+ [map { $_->{volid} } $PLUGIN->list_images('nvmetest', scfg(), 400)->@*],
+ ['nvmetest:vm-400-disk-0'],
+ 'filtered by VM',
+ );
+ my $listed = $PLUGIN->list_images('nvmetest', scfg(), undef, ['nvmetest:vm-100-disk-0']);
+ is_deeply([map { $_->{volid} } $listed->@*], ['nvmetest:vm-100-disk-0'], 'by volume list');
+ is(
+ $PLUGIN->find_free_diskname('nvmetest', scfg(), 200),
+ 'vm-200-disk-1',
+ 'a name held by an unowned zvol is not handed out',
+ );
+ is($PLUGIN->find_free_diskname('nvmetest', scfg(), 90), 'vm-90-disk-1', 'nor one of a base');
+ is(
+ $PLUGIN->find_free_diskname('nvmetest', scfg(), 104),
+ 'vm-104-disk-1',
+ 'nor one held by a filesystem',
+ );
+ my @steps = map { $_->{steps}->@* } @TARGET_CALLS;
+ ok(!grep({ $_->[0] ne 'zfs' || $_->[1] !~ /\A(?:get|list)\z/ } @steps), 'listing only reads');
+}
+
+{
+ my @warnings;
+ local $SIG{__WARN__} = sub($warning) { push @warnings, $warning };
+ for my $case (
+ [['100', '200'], [300, 100, 200, 1]],
+ [['-', '200'], [0, 0, 0, 0]],
+ [['1'], [0, 0, 0, 0]],
+ ) {
+ local $TARGET = target('zfs get' => $case->[0]);
+ is_deeply([$PLUGIN->status('nvmetest', scfg())], $case->[1], "status for '$case->[0]->@*'");
+ }
+ local $TARGET =
+ target('zfs get' => sub(@args) { return { rc => 255, out => [], err => 'no route' } });
+ is_deeply([$PLUGIN->status('nvmetest', scfg())], [0, 0, 0, 0], 'unreachable is inactive');
+ is(scalar(@warnings), 3, 'and every inactive status warns');
+ like($warnings[0], qr/unexpected ZFS pool usage for storage 'nvmetest'/, 'about the usage');
+ is(
+ $warnings[-1],
+ "storage 'nvmetest': zfs error on '192.0.2.10': no route\n",
+ 'or naming the storage and the target',
+ );
+}
+
+{
+ local @TARGET_CALLS = ();
+ local $TARGET = target();
+ eval { $PLUGIN->volume_resize(scfg(), 'nvmetest', 'vm-100-disk-0', 2 * 1024**3, 1) };
+ like(
+ $@,
+ qr/online resize is not supported for NVMe\/TCP block devices; stop the VM first/,
+ 'online resize is refused',
+ );
+ eval { $PLUGIN->volume_resize(scfg(), 'nvmetest', 'vm-100-disk-0', 2 * 1024**3, 0, 'snap') };
+ like($@, qr/resizing a snapshot is not supported/, 'as is resizing a snapshot');
+ is_deeply(\@TARGET_CALLS, [], 'before any target call');
+ for my $method (qw(volume_export volume_import rename_volume rename_snapshot)) {
+ eval { $PLUGIN->$method(scfg(), 'nvmetest', 'vm-100-disk-0') };
+ like($@, qr/not supported/, "$method is not supported");
+ }
+ for my $direction (qw(export import)) {
+ my $formats = "volume_${direction}_formats";
+ is_deeply(
+ [$PLUGIN->$formats(scfg(), 'nvmetest', 'vm-100-disk-0')],
+ [],
+ "no $direction formats",
+ );
+ }
+ is($PLUGIN->deactivate_volume('nvmetest', scfg(), 'vm-100-disk-0'), 1, 'no-op deactivation');
+ eval { $PLUGIN->deactivate_volume('nvmetest', scfg(), 'vm-100-disk-0', 'snap') };
+ like($@, qr/unable to deactivate snapshot/, 'snapshots cannot be deactivated');
+}
+
+{
+ $nvme_mock->redefine(_nvmet_volume_uuid => sub($scfg, $name) { return $UUID });
+ is_deeply(
+ $PLUGIN->qemu_blockdev_options(scfg(), 'nvmetest', 'vm-100-disk-0'),
+ { driver => 'host_device', filename => "/dev/disk/by-id/nvme-uuid.$UUID" },
+ 'the QEMU host_device driver on the stable namespace UUID link',
+ );
+ eval {
+ my $options = { 'snapshot-name' => 's' };
+ $PLUGIN->qemu_blockdev_options(scfg(), 'nvmetest', 'vm-100-disk-0', undef, $options);
+ };
+ like($@, qr/direct access to snapshots not implemented/, 'snapshots have no block device');
+ $nvme_mock->unmock('_nvmet_volume_uuid');
+}
+
+# ---------------------------------------------------------------------------
+# The runner: one ssh argv, no Perl, no shell on the local side
+# ---------------------------------------------------------------------------
+
+{
+ my $run = $nvme_mock->original('_nvmet_run');
+ my @runs;
+ local $COMMAND = sub($cmd, %opts) {
+ push @runs, { cmd => $cmd, %opts };
+ return 0;
+ };
+ my $steps = [['zfs', 'get', '-H', 'type', 'tank'], ['printf', '%s\n', 'x y']];
+ my $res = $run->(scfg(), $steps, op => 'test');
+ is_deeply(
+ $runs[0]->{cmd},
+ [
+ '/usr/bin/ssh',
+ '-o',
+ 'BatchMode=yes',
+ '-o',
+ 'ConnectTimeout=10',
+ '-o',
+ 'LogLevel=ERROR',
+ '-i',
+ '/etc/pve/priv/zfs/192.0.2.10_id_rsa',
+ 'root@192.0.2.10',
+ q{zfs get -H type tank && printf '%s\n' 'x y'},
+ ],
+ 'one ssh argv with the rendered command, and no login banner in errors',
+ );
+ unlike(join(' ', $runs[0]->{cmd}->@*), qr/perl/, 'no Perl runs on the target');
+ is($runs[0]->{timeout}, 15, 'reads time out after 15 seconds by default');
+ ok(!$runs[0]->{noerr} && !exists($runs[0]->{input}),
+ 'failures are exceptions, stdin is closed');
+ is_deeply($res, { rc => 0, out => [], err => '' }, 'a successful call');
+
+ @runs = ();
+ $run->(
+ scfg(),
+ [{ key => ["$ROOT/hosts/$hostnqn_a/dhchap_key"] }],
+ op => 'key',
+ input => "$KEY\n",
+ timeout => 7,
+ );
+ is($runs[0]->{input}, "$KEY\n", 'the key goes to stdin');
+ unlike(join(' ', $runs[0]->{cmd}->@*), qr/DHHC-1/, 'and never into the command line');
+ is($runs[0]->{timeout}, 7, 'with the timeout of the caller');
+
+ for my $case (
+ [
+ 'output and the first error line',
+ sub(%o) {
+ $o{outfunc}->('a');
+ $o{outfunc}->('b');
+ $o{errfunc}->('');
+ $o{errfunc}->('first');
+ $o{errfunc}->('second');
+ die "command 'ssh' failed: exit code 3\n";
+ },
+ { rc => 3, out => ['a', 'b'], err => 'first' },
+ ],
+ [
+ 'an exit code without message',
+ sub(%o) { die "command 'ssh' failed: exit code 2\n" },
+ { rc => 2, out => [], err => 'exit code 2' },
+ ],
+ [
+ 'a timeout',
+ sub(%o) { die "command 'ssh' failed: got timeout\n" },
+ { rc => -1, out => [], err => 'timeout' },
+ ],
+ [
+ 'an ssh that could not connect',
+ sub(%o) { die "command 'ssh' failed: exit code 255\n" },
+ { rc => 255, out => [], err => 'exit code 255' },
+ ],
+ [
+ 'an ssh killed by a signal',
+ sub(%o) { die "command 'ssh' failed: got signal 9\n" },
+ { rc => -1, out => [], err => 'ssh failed' },
+ ],
+ ) {
+ my ($name, $behaviour, $expected) = $case->@*;
+ local $COMMAND = sub($cmd, %opts) { return $behaviour->(%opts) };
+ is_deeply(
+ $run->(scfg(), [['cat', '/proc/mounts']], op => 'test'),
+ $expected,
+ "returns $name",
+ );
+ }
+ for my $exception ("received interrupt\n", "command 'ssh $KEY' failed: received interrupt\n") {
+ local $COMMAND = sub($cmd, %opts) { $opts{outfunc}->('partial'); die $exception };
+ eval { $run->(scfg(), [['cat', '/proc/mounts']], op => 'test') };
+ is($@, "received interrupt\n", 'a stopped task dies with the task marker and nothing else');
+ }
+ @runs = ();
+ eval { $run->(scfg(), [['cat', '/proc/mounts']]) };
+ like($@, qr/NVMe target call without label/, 'every call has a label');
+ eval { $run->(scfg(), [['printf', '%s', 'x' x 65536]], op => 'test') };
+ like($@, qr/NVMe target command too long/, 'oversized calls are refused');
+ is_deeply(\@runs, [], 'before running anything');
+}
+
+# ---------------------------------------------------------------------------
+# The shared storage lock
+# ---------------------------------------------------------------------------
+
+subtest 'vdisk_alloc dispatches under the shared storage lock' => sub {
+ my $cluster_mock = Test::MockModule->new('PVE::Cluster');
+ my (@locks, @allocations);
+ $cluster_mock->redefine(
+ cfs_lock_storage => sub($storeid, $timeout, $func, @param) {
+ push @locks, [$storeid, $timeout];
+ return $func->(@param);
+ },
+ );
+ $parent_mock->redefine(lookup => sub($class, $type) { return $PLUGIN });
+ $storage_mock->redefine(activate_storage => sub($cfg, $storeid) { return 1 });
+ $nvme_mock->redefine(
+ alloc_image => sub($class, $storeid, $scfg, $vmid, $fmt, $name, $size) {
+ push @allocations, [$storeid, $vmid, $fmt, $name, $size];
+ return $name;
+ },
+ );
+ my $cfg = { ids => { nvmetest => scfg(shared => 1) } };
+ is(
+ PVE::Storage::vdisk_alloc($cfg, 'nvmetest', 105, 'raw', 'vm-105-disk-0', 131_072),
+ 'nvmetest:vm-105-disk-0',
+ 'core allocation returns the allocated volume ID',
+ );
+ is_deeply(\@locks, [['nvmetest', undef]], 'cfs_lock_storage gets the timeout of the core');
+ is_deeply(
+ \@allocations,
+ [['nvmetest', 105, 'raw', 'vm-105-disk-0', 131_072]],
+ 'alloc_image runs once inside the lock',
+ );
+ $nvme_mock->unmock('alloc_image');
+ $storage_mock->unmock('activate_storage');
+ $parent_mock->unmock('lookup');
+};
+
+# ---------------------------------------------------------------------------
+# Local NVMe host side
+# ---------------------------------------------------------------------------
+
+# The connect options of a controller: one write to the fabrics device.
+{
+ my $portal = { address => '192.0.2.22', port => 4420, host_iface => 'ens21' };
+ my $options = sub($scfg) {
+ return nv(
+ '_fabrics_options', $scfg, $portal, $hostnqn_a, $UUID, $KEY,
+ );
+ };
+ my ($string, $names) = $options->({ subsysnqn => $NQN });
+ is(
+ $string,
+ "transport=tcp,traddr=192.0.2.22,trsvcid=4420,host_iface=ens21,nqn=$NQN,"
+ . "hostnqn=$hostnqn_a,hostid=$UUID,dhchap_secret=$KEY,keep_alive_tmo=5,"
+ . 'reconnect_delay=2,ctrl_loss_tmo=600',
+ 'the portal, the identities, the key and the default timeouts',
+ );
+ is_deeply(
+ $names,
+ [
+ qw(transport traddr trsvcid host_iface nqn hostnqn hostid dhchap_secret),
+ qw(keep_alive_tmo reconnect_delay ctrl_loss_tmo),
+ ],
+ 'with the names of the options',
+ );
+ my $tuned = {
+ subsysnqn => $NQN,
+ 'nvme-keep-alive-tmo' => 10,
+ 'nvme-reconnect-delay' => 5,
+ 'nvme-ctrl-loss-tmo' => 60,
+ 'nvme-fast-io-fail-tmo' => 15,
+ 'nvme-nr-io-queues' => 4,
+ };
+ ($string, $names) = $options->($tuned);
+ my $tail = 'keep_alive_tmo=10,reconnect_delay=5,ctrl_loss_tmo=60,'
+ . 'fast_io_fail_tmo=15,nr_io_queues=4';
+ like($string, qr/,\Q$tail\E\z/, 'the configured timeouts and queues');
+ is_deeply([$names->@[-2, -1]], ['fast_io_fail_tmo', 'nr_io_queues'], 'are named too');
+ ($string) = $options->({ $tuned->%*, 'nvme-fast-io-fail-tmo' => 0 });
+ like($string, qr/,fast_io_fail_tmo=0,/, 'a fast I/O fail timeout of 0 is a real value');
+ ($string) = $options->({ $tuned->%*, 'nvme-ctrl-loss-tmo' => -1 });
+ like($string, qr/,ctrl_loss_tmo=-1,/, 'an infinite loss timeout');
+
+ for my $value ('6 0', '60,duplicate_connect', "60\x00", '') {
+ eval { $options->({ $tuned->%*, 'nvme-ctrl-loss-tmo' => $value }) };
+ is(
+ $@,
+ "internal error: invalid NVMe connect option 'ctrl_loss_tmo'\n",
+ 'a value that would end the option early is refused without echoing it',
+ );
+ }
+}
+
+# The connect runs in a child bounded by 10 seconds. Its result and errors
+# pass through; a stopped task and a timeout have their own message.
+{
+ my $portal = { address => '192.0.2.21', port => 4420, host_iface => 'ens19' };
+ my $connect = sub($child) {
+ local @FORKS = ();
+ $nvme_mock->redefine(_fabrics_connect => $child);
+ my $res = eval { nv('_connect_portal', scfg(), $portal, $hostnqn_a, $UUID, $KEY); };
+ my $error = $@;
+ $nvme_mock->redefine(_fabrics_connect => $connect_mock);
+ return ($res, $error, [@FORKS]);
+ };
+ my ($res, $error, $forks) = $connect->(sub($options, $names) { return 7 });
+ is($res, 7, 'returns the instance of the new controller');
+ is_deeply($forks, [10], 'from a child bounded by 10 seconds');
+ for my $exception ("received interrupt\n", "interrupted by unexpected signal\n") {
+ (undef, $error) = $connect->(sub(@args) {
+ die $exception;
+ });
+ is($error, "received interrupt\n", 'a stopped task dies with the task marker');
+ }
+ my @warnings;
+ local $SIG{__WARN__} = sub($warning) { push @warnings, $warning };
+ $tools_mock->redefine(
+ run_fork_with_timeout => sub($timeout, $code, $opts = undef) { return (undef, 1) },
+ );
+ (undef, $error) = $connect->(sub(@args) {
+ return 7;
+ });
+ is($error, "NVMe/TCP connect did not complete within 10 seconds\n", 'a timeout');
+ $tools_mock->redefine(
+ run_fork_with_timeout => sub($timeout, $code, $opts = undef) {
+ warn "malformed JSON string\n"; # what a child that died without a result leaves
+ return (undef, 0);
+ },
+ );
+ (undef, $error) = $connect->(sub(@args) {
+ return 7;
+ });
+ is($error, "NVMe/TCP connect left no result\n", 'a child without a result');
+ is_deeply(\@warnings, [], 'without the warning of run_fork_with_timeout in the task log');
+ $tools_mock->redefine(run_fork_with_timeout => $fork_inline);
+}
+
+# The signals a stopped task can bring that are blocked while a child runs.
+my %signal = (HUP => SIGHUP, INT => SIGINT, QUIT => SIGQUIT, TERM => SIGTERM, ALRM => SIGALRM);
+
+sub blocked_signals() {
+ my $mask = POSIX::SigSet->new();
+ sigprocmask(SIG_BLOCK, POSIX::SigSet->new(), $mask) or die "sigprocmask: $!\n";
+ return join(',', grep { $mask->ismember($signal{$_}) } sort keys %signal);
+}
+
+# The child itself: a real fork under the real run_fork_with_timeout.
+subtest 'the connect child' => sub {
+ $tools_mock->unmock('run_fork_with_timeout');
+ my $portal = { address => '192.0.2.21', port => 4420, host_iface => 'ens19' };
+ my $connect = sub() {
+ my $res = eval { nv('_connect_portal', scfg(), $portal, $hostnqn_a, $UUID, $KEY); };
+ return ($res, $@);
+ };
+ $nvme_mock->redefine(_fabrics_connect => sub(@args) { return 7 });
+ my ($res, $error) = $connect->();
+ is($res, 7, 'the result of the child reaches the parent');
+ is(blocked_signals(), '', 'and the signal mask is restored');
+ $nvme_mock->redefine(
+ _fabrics_connect => sub(@args) { die "connect failed: Connection refused\n" },
+ );
+ (undef, $error) = $connect->();
+ is($error, "connect failed: Connection refused\n", 'and so does its error');
+ is(blocked_signals(), '', 'with the signal mask restored');
+
+ # The warnings of the child go to the standard error of the task.
+ my $dir = tempdir(CLEANUP => 1);
+ $nvme_mock->redefine(_fabrics_connect => sub(@args) { warn "child warning\n"; return 7 });
+ open(my $saved, '>&', \*STDERR) or die "cannot save STDERR: $!\n";
+ open(STDERR, '>', "$dir/stderr") or die "cannot redirect STDERR: $!\n";
+ ($res, $error) = $connect->();
+ open(STDERR, '>&', $saved) or die "cannot restore STDERR: $!\n";
+ is($res, 7, 'a child that warns');
+ open(my $fh, '<', "$dir/stderr") or die "cannot read '$dir/stderr': $!\n";
+ is(join('', <$fh>), "child warning\n", 'is heard');
+ close($fh);
+
+ # A stopped task: the signal stays pending while the child runs and, once
+ # the child is reaped, the handler of the worker dies with the task marker.
+ for my $name (qw(TERM QUIT INT)) {
+ pipe(my $from_child, my $to_parent) or die "pipe: $!\n";
+ $nvme_mock->redefine(
+ _fabrics_connect => sub(@args) {
+ kill($name, getppid()) or die "kill: $!\n";
+ print {$to_parent} "$$ " . blocked_signals() . "\n";
+ close($to_parent);
+ return 7;
+ },
+ );
+ my $seen;
+ local $SIG{$name} = sub { # as in a PVE worker
+ chomp(my $report = <$from_child> // '');
+ my ($pid, $blocked) = split(/ /, $report, 2);
+ $seen = { blocked => $blocked, reaped => $pid && waitpid($pid, WNOHANG) == -1 };
+ die "received interrupt\n";
+ };
+ ($res, $error) = $connect->();
+ close($to_parent);
+ close($from_child);
+ is($error, "received interrupt\n", "$name stops the task with the task marker");
+ is(
+ $seen->{blocked},
+ 'HUP,INT,QUIT,TERM',
+ 'once the child, with the inherited mask, is done',
+ );
+ ok($seen->{reaped}, 'and reaped');
+ is(blocked_signals(), '', 'and the signal mask is restored');
+ }
+ $nvme_mock->redefine(_fabrics_connect => $connect_mock);
+ $tools_mock->redefine(run_fork_with_timeout => $fork_inline);
+};
+
+# The body of the connect child, inline. A pseudoterminal stands in for the
+# fabrics device: a character device that answers reads with the lines the
+# test writes to the master, and passes writes to the master.
+subtest 'the connect child body' => sub {
+ my $body = $nvme_mock->original('_fabrics_connect');
+ my $portal = { address => '192.0.2.21', port => 4420, host_iface => 'ens19' };
+ my ($options, $names) = nv('_fabrics_options', scfg(), $portal, $hostnqn_a, $UUID, $KEY);
+ my $tokens = sub(@names) {
+ return join(',', map { "$_=%s" } @names) . "\n";
+ };
+ # a body that waits for a line the test did not write must not hang the suite
+ local $SIG{ALRM} = sub { die "the child body did not finish\n" };
+ my $run = sub() {
+ @RESTRICTED = ();
+ alarm(30);
+ my $res = eval { $body->($options, $names) };
+ alarm(0);
+ return ($res, $@);
+ };
+ my $dir = tempdir(CLEANUP => 1);
+ my $file = "$dir/file";
+ open(my $fh, '>', $file) or die "cannot create '$file': $!\n";
+ close($fh);
+ for my $case (
+ [
+ 'a missing device',
+ "$dir/missing",
+ qr/\A'\Q$dir\E\/missing' does not exist; load the nvme-tcp kernel module\n\z/,
+ ],
+ ['a regular file', $file, qr/\A'\Q$file\E' is not a character device\n\z/],
+ [
+ 'a device without the option list',
+ '/dev/null',
+ qr/\Athe running kernel does not support the NVMe connect option 'transport'\n\z/,
+ ],
+ ) {
+ my ($name, $device, $expected) = $case->@*;
+ $nvme_mock->redefine(_fabrics_device => sub () { return $device });
+ my (undef, $error) = $run->();
+ like($error, $expected, "$name is refused");
+ unlike($error, qr/\Q$KEY\E/, 'without the key in the error');
+ }
+ is(-s $file, 0, 'and nothing is written before the device is checked');
+
+ my $pty = sub() {
+ # TIOCSPTLCK and TIOCGPTN are only known for these architectures
+ return if $Config{archname} !~ m/\A(?:x86_64|i[3-6]86|aarch64|arm|riscv64)-linux/;
+ sysopen(my $master, '/dev/ptmx', O_RDWR | O_NOCTTY) or return;
+ my $zero = pack('i', 0);
+ ioctl($master, 0x40045431, $zero) or return;
+ my $number = pack('i', 0);
+ ioctl($master, 0x80045430, $number) or return;
+ my $slave = '/dev/pts/' . unpack('i', $number);
+ # keeps the line discipline: canonical reads, no echo, no translation
+ sysopen(my $keep, $slave, O_RDWR | O_NOCTTY) or return;
+ my $termios = POSIX::Termios->new();
+ $termios->getattr(fileno($keep)) or die "tcgetattr: $!\n";
+ $termios->setlflag(($termios->getlflag() & ~(ECHO | ECHONL)) | ICANON);
+ $termios->setoflag($termios->getoflag() & ~OPOST);
+ $termios->setattr(fileno($keep), TCSANOW) or die "tcsetattr: $!\n";
+ return ($master, $slave, $keep);
+ };
+ my $received = sub($master) {
+ my $readable = '';
+ vec($readable, fileno($master), 1) = 1;
+ return '' if !select($readable, undef, undef, 0.1);
+ sysread($master, my $buffer, 65536) // die "cannot read the pseudoterminal: $!\n";
+ return $buffer;
+ };
+ my ($master, $slave, $keep) = $pty->();
+ SKIP: {
+ skip 'no pseudoterminal available', 25 if !$master;
+ $nvme_mock->redefine(_fabrics_device => sub () { return $slave });
+
+ my ($res, $error);
+ for my $instance (3, 12) {
+ syswrite($master, $tokens->($names->@*, 'tls'));
+ syswrite($master, "instance=$instance,cntlid=1\n");
+ ($res, $error) = $run->();
+ is($error, '', "a connect (instance $instance)");
+ is($res, $instance, 'returns the instance of the controller');
+ is($received->($master), $options, 'after exactly one write of the options');
+ my $attrs = "/sys/class/nvme/nvme$instance";
+ is_deeply(
+ \@RESTRICTED,
+ ["$attrs/dhchap_secret", "$attrs/dhchap_ctrl_secret"],
+ 'and restricts the secret attributes of that controller',
+ );
+ }
+
+ for my $missing (qw(host_iface nqn)) {
+ syswrite($master,
+ $tokens->(map { $_ eq $missing ? "${_}_suffix" : $_ } $names->@*));
+ ($res, $error) = $run->();
+ is(
+ $error,
+ "the running kernel does not support the NVMe connect option '$missing'\n",
+ "an option the kernel does not list ($missing)",
+ );
+ is($received->($master), '', 'is found before anything is written');
+ }
+
+ for my $result ("garbage\n", "instance=3,cntlid=1,garbage\n") {
+ syswrite($master, $tokens->($names->@*));
+ syswrite($master, $result);
+ ($res, $error) = $run->();
+ is($error, "unexpected connect result\n", 'an unexpected result is refused');
+ is($received->($master), $options, 'after the one write');
+ is_deeply(\@RESTRICTED, [], 'and nothing is restricted');
+ }
+
+ symlink($slave, "$dir/link") or die "symlink: $!\n";
+ $nvme_mock->redefine(_fabrics_device => sub () { return "$dir/link" });
+ ($res, $error) = $run->();
+ is(
+ $error,
+ "cannot open '$dir/link': Too many levels of symbolic links\n",
+ 'a symbolic link to the device is not followed',
+ );
+ $nvme_mock->redefine(_fabrics_device => sub () { return $slave });
+
+ # A failed write: the error has the errno text and never the options.
+ my $writes = 0;
+ $nvme_mock->redefine(
+ _fabrics_write => sub($fh, $data) { $writes++; $! = ECONNREFUSED; return undef },
+ );
+ syswrite($master, $tokens->($names->@*));
+ ($res, $error) = $run->();
+ is(
+ $error,
+ "connect failed: Connection refused\n",
+ 'a refused connect has the errno text',
+ );
+ is($writes, 1, 'after one write, which is not repeated');
+ is_deeply(\@RESTRICTED, [], 'and nothing is restricted');
+ $writes = 0;
+ $nvme_mock->redefine(
+ _fabrics_write => sub($fh, $data) { $writes++; return length($data) - 1 });
+ syswrite($master, $tokens->($names->@*));
+ syswrite($master, "instance=3,cntlid=1\n"); # never read: a short write ends the connect
+ ($res, $error) = $run->();
+ is($error, "connect failed: short write\n", 'a short write');
+ is($writes, 1, 'is not completed');
+ is_deeply(\@RESTRICTED, [], 'and restricts nothing');
+ $nvme_mock->unmock('_fabrics_write');
+ }
+ $nvme_mock->unmock('_fabrics_device');
+};
+
+# Deactivation deletes the controllers of the subsystem in one bounded child.
+{
+ reset_host();
+ add_controller('nvme1', '192.0.2.21', 4420, 'ens19');
+ add_controller('nvme2', '192.0.2.22', 4420, 'ens20');
+ add_controller('nvme3', '192.0.2.23', 4420, 'ens21', 'live', $OTHER_NQN);
+ $nvme_mock->redefine(
+ _namespace_openers =>
+ sub($nqn) { return ['qemu-system-x86_64 (PID 123, /dev/nvme0n1)'] },
+ );
+ eval { $PLUGIN->deactivate_storage('nvmetest', scfg()) };
+ like(
+ $@,
+ qr/refusing to disconnect.*namespace in use by qemu-system-x86_64/s,
+ 'deactivation refuses a namespace opened by a VM',
+ );
+ is_deeply(\@DELETED, [], 'and deletes nothing');
+ $nvme_mock->redefine(_namespace_openers => sub($nqn) { return [] });
+ is($PLUGIN->deactivate_storage('nvmetest', scfg()), 1, 'an unused storage is deactivated');
+ is_deeply(\@DELETED, ['nvme1', 'nvme2'], 'by deleting the controllers of its subsystem');
+ is_deeply(\@FORKS, [15], 'in one child bounded by 15 seconds');
+ is_deeply($DIRS{'/sys/class/nvme'}, ['nvme3'], 'the controllers of other subsystems stay');
+ is($PLUGIN->deactivate_storage('nvmetest', scfg()), 1, 'a storage without controllers');
+ is_deeply(\@FORKS, [15], 'needs no child');
+ $nvme_mock->unmock('_namespace_openers');
+
+ my $delete = $nvme_mock->original('_delete_controllers');
+ @SYSFS_WRITES = ();
+ is($delete->([{ name => 'nvme1' }, { name => 'nvme2' }]), 2, 'the child deletes each one');
+ is_deeply(
+ \@SYSFS_WRITES,
+ [map { ["/sys/class/nvme/$_/delete_controller", '1'] } qw(nvme1 nvme2)],
+ 'through sysfs',
+ );
+ # The stock file_write on files that fail: a failed write warns, with the
+ # errno text, and a failed open does not.
+ my $file_write = $sysfs_mock->original('file_write');
+ my $dir = tempdir(CLEANUP => 1);
+ my @warnings;
+ local $SIG{__WARN__} = sub($warning) { push @warnings, $warning };
+ for my $case (['/dev/full', 'No space left on device'], [$dir, 'Is a directory']) {
+ my ($path, $reason) = $case->@*;
+ $sysfs_mock->redefine(file_write => sub($file, @args) { $file_write->($path, @args) });
+ eval { $delete->([{ name => 'nvme1' }, { name => 'nvme2' }]) };
+ is(
+ $@,
+ "cannot delete NVMe controller 'nvme1': $reason\n",
+ "a failed delete names the controller and the errno text ($reason)",
+ );
+ }
+ $sysfs_mock->redefine(file_write => sub($file, @args) { $file_write->("$dir/a/b", @args) });
+ is(
+ $delete->([{ name => 'nvme1' }, { name => 'nvme2' }]),
+ 2,
+ 'a controller that went away since it was listed counts as deleted',
+ );
+ $sysfs_mock->redefine(
+ file_write => sub($path, $data, @rest) {
+ warn "error writing '$data' to '$path': Device or resource busy\n";
+ $! = 0;
+ return 0;
+ },
+ );
+ eval { $delete->([{ name => 'nvme1' }]) };
+ is(
+ $@,
+ "cannot delete NVMe controller 'nvme1': Device or resource busy\n",
+ 'the errno text of a failed write is the one of the warning',
+ );
+ is_deeply(\@warnings, [], 'which is not passed on');
+ $sysfs_mock->redefine(file_write => $sysfs_write);
+}
+
+# The users of the namespaces: file descriptors on a namespace or one of its
+# partitions, and kernel holders such as device-mapper.
+{
+ reset_host();
+ $DIRS{'/sys/class/nvme-subsystem/nvme-subsys7'} = ['nvme7n1', 'nvme7c1n1'];
+ $DIRS{'/sys/class/block/nvme7n1'} = ['nvme7n1p1', 'nvme7n1p2', 'holders'];
+ is_deeply(nv('_namespace_openers', $NQN), [], 'no device, no users');
+ my ($scan) = grep { $_->[0] eq '/sys/class/block/nvme7n1' } @GLOBS;
+ ok($scan, 'the partitions of the namespace are listed');
+ ok($scan && 'nvme7n1p1' =~ /^($scan->[1])$/ && 'nvme7n1' !~ /^($scan->[1])$/, 'only those');
+ ok(!grep({ $_->[0] eq '/sys/class/block/nvme7c1n1' } @GLOBS), 'not the path devices');
+
+ %BLOCK = map { ("/dev/$_" => 1) } qw(nvme7n1 nvme7n1p1 nvme7n1p2);
+ $DIRS{'/proc'} = [qw(100 200 300 self)];
+ $DIRS{'/proc/100/fd'} = [0, 1, 7];
+ $DIRS{'/proc/200/fd'} = [3];
+ $DIRS{'/proc/300/fd'} = [4];
+ %LINKS = (
+ '/proc/100/fd/0' => '/dev/null',
+ '/proc/100/fd/7' => '/dev/nvme7n1',
+ '/proc/200/fd/3' => '/dev/nvme7n1p1',
+ '/proc/300/fd/4' => '/dev/nvme8n1',
+ );
+ $FILES{'/proc/100/comm'} = 'qemu-system-x86';
+ $FILES{'/proc/200/comm'} = 'mkfs.ext4';
+ $DIRS{'/sys/class/block/nvme7n1p2/holders'} = ['dm-0'];
+ is_deeply(
+ nv('_namespace_openers', $NQN),
+ [
+ '/dev/nvme7n1p2 held by dm-0',
+ 'mkfs.ext4 (PID 200, /dev/nvme7n1p1)',
+ 'qemu-system-x86 (PID 100, /dev/nvme7n1)',
+ ],
+ 'open namespaces and partitions and their holders are users, other devices are not',
+ );
+}
+
+# The local device behind a namespace link must be that namespace.
+{
+ my $link = "/dev/disk/by-id/nvme-uuid.$UUID";
+ my $check = sub($target, %files) {
+ local %LINKS = (defined($target) ? ($link => $target) : ());
+ local %FILES = (
+ '/sys/block/nvme3n1/uuid' => uc($UUID),
+ '/sys/block/nvme3n1/nsid' => '7',
+ '/sys/block/nvme3n1/device/subsysnqn' => $NQN,
+ %files,
+ );
+ my $ok = nv('_nvmet_local_namespace_ok', $link, $NQN, 7, $UUID);
+ return $ok ? 1 : 0;
+ };
+ my $device = '../../nvme3n1';
+ is($check->($device), 1, 'accepts the namespace of the subsystem, whatever the UUID case');
+ is($check->($device, '/sys/block/nvme3n1/nsid' => '8'), 0, 'refuses another NSID');
+ is(
+ $check->($device, '/sys/block/nvme3n1/device/subsysnqn' => $OTHER_NQN),
+ 0,
+ 'another subsystem',
+ );
+ is(
+ $check->($device, '/sys/block/nvme3n1/uuid' => '22345678-1234-1234-1234-123456789abc'),
+ 0,
+ 'another UUID',
+ );
+ is($check->($device, '/sys/block/nvme3n1/uuid' => undef), 0, 'an unreadable device');
+ is($check->(undef), 0, 'a missing link');
+
+ for my $target ('../../nvme3c1n1', '../../nvme3n1p1', '../../sda') {
+ is($check->($target), 0, "a link to $target");
+ }
+
+ my $dir = tempdir(CLEANUP => 1);
+ symlink('../../nvme3n1', "$dir/link") or die "symlink: $!\n";
+ is($nvme_mock->original('_link_target')->("$dir/link"), '../../nvme3n1', 'read, not followed');
+}
+
+# Storage activation. Everything on the local host is mocked.
+my $activate = sub($storeid, $scfg, $cache = undef) {
+ local $FILES{"/etc/pve/priv/storage/$storeid.nvme-dhchap"} = $KEY;
+ local %CONFIG = (%CONFIG, $storeid => $scfg);
+ @ACTIVATIONS = ();
+ $nvme_mock->redefine(
+ _nvmet_activate_target => sub(@args) { push @ACTIVATIONS, [@args]; return },
+ );
+ my $res = eval { $PLUGIN->activate_storage($storeid, $scfg, $cache) };
+ my $error = $@;
+ $nvme_mock->unmock('_nvmet_activate_target');
+ return ($res, $error);
+};
+my $forced = sub($storeid) { return { 'zfsnvme-force-reconcile' => { $storeid => 1 } } };
+
+subtest 'activation preconditions' => sub {
+ my $file = sub($path, $value) {
+ return sub() { $FILES{$path} = $value }
+ };
+ my $missing_iface = sub() {
+ $nvme_mock->redefine(_local_iface_exists => sub($iface) { return $iface ne 'ens20' });
+ };
+ my $other_storage = sub(%override) {
+ return sub() { $CONFIG{other} = scfg(subsysnqn => $OTHER_NQN, %override) };
+ };
+ my $fabrics_injection = 'traddr=198.51.100.9';
+ for my $case (
+ [
+ 'NVMe kernel modules that are not loaded',
+ $file->('/sys/module/nvme_core/parameters/multipath', undef),
+ qr/the NVMe kernel modules are not loaded; load the nvme-tcp kernel module/,
+ ],
+ [
+ 'multipath disabled',
+ $file->('/sys/module/nvme_core/parameters/multipath', 'N'),
+ qr/native NVMe multipath is disabled/,
+ ],
+ [
+ 'no host NQN',
+ $file->('/etc/nvme/hostnqn', undef),
+ qr/missing \/etc\/nvme\/hostnqn; the nvme-cli package generates it/,
+ ],
+ [
+ 'a host that is not allowed',
+ $file->(
+ '/etc/nvme/hostnqn',
+ 'nqn.2014-08.org.nvmexpress:uuid:00000000-0000-4000-8000-00000000dead',
+ ),
+ qr/local NVMe host NQN .* is missing from nvme-host-nqns/,
+ ],
+ [
+ 'no host ID',
+ $file->('/etc/nvme/hostid', undef),
+ qr/missing \/etc\/nvme\/hostid; the nvme-cli package generates it/,
+ ],
+ [
+ 'the zero host ID',
+ $file->('/etc/nvme/hostid', '00000000-0000-0000-0000-000000000000'),
+ qr/invalid NVMe host ID/,
+ ],
+ [
+ 'a host ID with a connect option',
+ $file->('/etc/nvme/hostid', "$UUID,$fabrics_injection"),
+ qr/invalid NVMe host ID/,
+ ],
+ [
+ 'a host ID with a newline',
+ $file->('/etc/nvme/hostid', "$UUID\n"),
+ qr/invalid NVMe host ID/,
+ ],
+ [
+ 'a connection parameter that is not a number',
+ sub() { },
+ qr/invalid NVMe connection parameter 'reconnect-delay'/,
+ 'nvme-reconnect-delay' => "2,$fabrics_injection",
+ ],
+ [
+ 'a missing local interface',
+ $missing_iface,
+ qr/interface 'ens20' does not exist on this node/,
+ ],
+ [
+ 'a pool shared with another storage',
+ $other_storage->('nvme-portals' => '192.0.2.31,192.0.2.32'),
+ qr/ZFS pool 'tank' on '192\.0\.2\.10' is already used by storage 'other'/,
+ ],
+ [
+ 'a listener shared with another storage',
+ $other_storage->(pool => 'data', 'nvme-portals' => '192.0.2.31,192.0.2.22'),
+ qr/portal '192\.0\.2\.22' port 4420 is already used by storage 'other'/,
+ ],
+ ) {
+ my ($name, $setup, $error, %override) = $case->@*;
+ reset_host();
+ local %CONFIG = ();
+ $nvme_mock->redefine(_local_iface_exists => sub($iface) { return 1 });
+ $setup->();
+ my (undef, $got) = $activate->('zfsnvme-unit-s', scfg(%override));
+ like($got, $error, "refuses $name");
+ is_deeply([@ACTIVATIONS, @CONNECTS, @SYSFS_WRITES], [], 'and does nothing');
+ }
+ $nvme_mock->redefine(_local_iface_exists => sub($iface) { return 1 });
+};
+
+subtest 'a stopped task ends the activation at the connect' => sub {
+ for my $case (
+ ["received interrupt\n", 1],
+ ["interrupted by unexpected signal\n", 1],
+ ["connect failed: Connection refused\n", 0],
+ ["NVMe/TCP connect did not complete within 10 seconds\n", 0],
+ ) {
+ my ($exception, $cancelled) = $case->@*;
+ reset_host();
+ my $calls = 0;
+ $nvme_mock->redefine(
+ _fabrics_connect => sub($options, $names) {
+ my %option = map { split(/=/, $_, 2) } split(/,/, $options);
+ ++$calls;
+ # cancellation must escape even if the write created a live path
+ add_controller("nvme$calls", $option{traddr}, 4420, $option{host_iface})
+ if $cancelled || $calls > 1;
+ die $exception if $calls == 1;
+ return $calls;
+ },
+ );
+ my ($res, $err) = $activate->('zfsnvme-cancel', scfg());
+ is($err, $cancelled ? "received interrupt\n" : '', 'only a stopped task escapes');
+ is($res, $cancelled ? undef : 1, 'an ordinary failure still permits a degraded path');
+ is($calls, $cancelled ? 1 : 2, 'a stopped task never attempts the next portal');
+ unlike($err . join('', @WARNINGS), qr/\Q$KEY\E/, 'errors and warnings hide the key');
+ }
+ $nvme_mock->redefine(_fabrics_connect => $connect_mock);
+
+ reset_host();
+ no warnings 'redefine';
+ local *PVE::Storage::ZFSNVMePlugin::_restrict_attr =
+ sub($path) { die "received interrupt\n" };
+ my ($res, $err) = $activate->('zfsnvme-cancel-scan', scfg());
+ is($err, "received interrupt\n",
+ 'cancellation in the post-connect permission scan escapes');
+ is(scalar(@CONNECTS), 1, 'no next portal after a cancelled permission scan');
+};
+
+# Deleting a controller runs in a bounded child: 5 seconds for a dead path
+# during an activation, which goes on, and 15 seconds for a deactivation,
+# which fails. Only a stopped task escapes an activation.
+subtest 'deleting a controller is bounded and keeps a stopped task' => sub {
+ my $interrupt = "received interrupt\n";
+ my $failed = "cannot delete NVMe controller 'nvme1'\n";
+ # [label, what the child does, activation error, deactivation error]
+ for my $case (
+ ['a stopped task', $interrupt, $interrupt, $interrupt],
+ ['a signalled child', "interrupted by unexpected signal\n", $interrupt, $interrupt],
+ ['a failed deletion', $failed, '', $failed],
+ [
+ 'a timeout',
+ undef,
+ '',
+ "disconnecting NVMe subsystem '$NQN' did not complete within 15 seconds\n",
+ ],
+ ) {
+ my ($label, $exception, $activation_error, $deactivation_error) = $case->@*;
+ $tools_mock->redefine(
+ run_fork_with_timeout => sub($seconds, $code, $opts = undef) {
+ push @FORKS, $seconds;
+ # without an exception, the deletion times out
+ return (undef, 1) if !defined($exception) && $seconds != 10;
+ return ($code->(), 0);
+ },
+ );
+ $nvme_mock->redefine(_delete_controllers => sub($controllers) { die $exception })
+ if defined($exception);
+
+ reset_host();
+ add_controller('nvme1', '192.0.2.21', 4420, 'ens19', 'dead');
+ my ($res, $err) = $activate->('zfsnvme-delete', scfg());
+ is($err, $activation_error, "$label while a dead path is deleted");
+ is($FORKS[0], 5, 'in a child bounded by 5 seconds');
+ is(scalar(@CONNECTS), $activation_error ? 0 : 2, 'only a failure goes on to connect');
+ my $warned = $exception
+ // "deleting NVMe controller 'nvme1' did not complete within 5 seconds\n";
+ like(join("\n", @WARNINGS), qr/\Q$warned\E/, 'and is warned about')
+ if !$activation_error;
+
+ reset_host();
+ add_controller('nvme1', '192.0.2.21', 4420, 'ens19');
+ eval { $PLUGIN->deactivate_storage('zfsnvme-delete', scfg()) };
+ is($@, $deactivation_error, "$label fails the deactivation");
+ is_deeply(\@FORKS, [15], 'in a child bounded by 15 seconds');
+ $nvme_mock->redefine(_delete_controllers => $delete_mock);
+ }
+ $tools_mock->redefine(run_fork_with_timeout => $fork_inline);
+};
+
+subtest 'storage deletion preserves task cancellation at each catch boundary' => sub {
+ for my $case (
+ ['deactivate', "received interrupt\n", 1, 0],
+ ['deactivate', "ordinary disconnect failure\n", 0, 1],
+ [
+ 'deactivate',
+ "disconnecting NVMe subsystem '$NQN' did not complete within 15 seconds\n",
+ 0,
+ 1,
+ ],
+ ['target', "received interrupt\n", 1, 1],
+ ['target', "ordinary target failure\n", 0, 1],
+ ) {
+ my ($phase, $exception, $cancelled, $expected_target_calls) = $case->@*;
+ reset_host();
+ my $target_calls = 0;
+ no warnings 'redefine';
+ local *PVE::Storage::ZFSNVMePlugin::deactivate_storage = sub(@args) {
+ die $exception if $phase eq 'deactivate';
+ return 1;
+ };
+ local *PVE::Storage::ZFSNVMePlugin::_nvmet_delete_target = sub(@args) {
+ ++$target_calls;
+ die $exception if $phase eq 'target';
+ return;
+ };
+ eval { $PLUGIN->on_delete_hook('zfsnvme-delete-cancel', scfg()) };
+ is(
+ $@,
+ $cancelled ? "received interrupt\n" : '',
+ "$phase preserves only task cancellation",
+ );
+ is($target_calls, $expected_target_calls, 'no target deletion after a cancellation');
+ is_deeply(
+ \@UNLINKED,
+ ['/etc/pve/priv/storage/zfsnvme-delete-cancel.nvme-dhchap'],
+ 'the key file is removed first',
+ );
+ chomp(my $message = $exception);
+ like(
+ join("\n", @WARNINGS),
+ qr/not disconnecting NVMe storage 'zfsnvme-delete-cancel': \Q$message\E/,
+ 'an ordinary deactivation failure only warns',
+ ) if $phase eq 'deactivate' && !$cancelled;
+ }
+};
+
+subtest 'slow-path backoff' => sub {
+ reset_host();
+ local %REFUSE = ('192.0.2.22' => 1);
+ add_controller('nvme1', '192.0.2.21', 4420, 'ens19');
+ my $scfg = scfg();
+ my ($res, $error) = $activate->('zfsnvme-unit-a', $scfg);
+ is($error, '', 'a degraded storage activates');
+ is(scalar(@ACTIVATIONS), 1, 'through the target');
+ is_deeply(
+ [$ACTIVATIONS[0]->@[0, 2 .. 4]],
+ [
+ 'zfsnvme-unit-a',
+ [
+ map { {
+ family => 'ipv4',
+ address => "192.0.2.2$_",
+ port => 4420,
+ host_iface => 'ens' . (18 + $_),
+ } } 1,
+ 2,
+ ],
+ [$hostnqn_a, $hostnqn_b],
+ $KEY,
+ ],
+ 'with the storage, its parsed portals, hosts and key',
+ );
+ like(
+ join("\n", @WARNINGS),
+ qr/storage 'zfsnvme-unit-a' is degraded: 1\/2 paths live/,
+ 'and warns',
+ );
+ ($res, $error) = $activate->('zfsnvme-unit-a', $scfg);
+ is($error, '', 'within a minute, a live path takes the fast path');
+ is(scalar(@ACTIVATIONS), 0, 'without a target call');
+ {
+ local $NOW = $NOW + 61;
+ ($res, $error) = $activate->('zfsnvme-unit-a', $scfg);
+ is(scalar(@ACTIVATIONS), 1, 'a minute later the target is tried again');
+ }
+ my $cache = $forced->('zfsnvme-unit-a');
+ ($res, $error) = $activate->('zfsnvme-unit-a', $scfg, $cache);
+ is(scalar(@ACTIVATIONS), 1, 'a forced activation always reaches the target');
+ ok(!$cache->{'zfsnvme-force-reconcile'}->{'zfsnvme-unit-a'}, 'and consumes the request');
+ $FILES{'/sys/class/nvme/nvme1/state'} = 'dead';
+ ($res, $error) = $activate->('zfsnvme-unit-a', $scfg);
+ is(
+ $error,
+ "no live NVMe/TCP path for storage 'zfsnvme-unit-a'\n",
+ 'within a minute, without a live path the activation fails at once',
+ );
+ is_deeply([@ACTIVATIONS, @DELETED], [], 'without a target call or a reconnect');
+ {
+ local $NOW = $NOW + 61;
+ ($res, $error) = $activate->('zfsnvme-unit-a', $scfg);
+ is(scalar(@ACTIVATIONS), 1, 'a minute later the target is tried');
+ is_deeply(\@DELETED, ['nvme1'], 'and the dead path is reconnected');
+ }
+
+ reset_host();
+ local %REFUSE = ('192.0.2.22' => 1);
+ add_controller('nvme1', '192.0.2.21', 4420, 'ens19');
+ local $NOW = 30; # the monotonic clock shortly after boot
+ ($res, $error) = $activate->('zfsnvme-unit-b', $scfg);
+ is(scalar(@ACTIVATIONS), 1, 'a storage that was never activated is not in its backoff');
+
+ # A storage without any path waits for one once per backoff period, not
+ # on every status update.
+ reset_host();
+ local %REFUSE = ('192.0.2.21' => 1, '192.0.2.22' => 1);
+ ($res, $error) = $activate->('zfsnvme-unit-d', $scfg);
+ is($error, "no live NVMe/TCP path for storage 'zfsnvme-unit-d'\n", 'no path, no storage');
+ is_deeply(
+ [scalar(@ACTIVATIONS), scalar(@CONNECTS)],
+ [1, 2],
+ 'after the target and each portal',
+ );
+ @CONNECTS = ();
+ ($res, $error) = $activate->('zfsnvme-unit-d', $scfg);
+ is(
+ $error,
+ "no live NVMe/TCP path for storage 'zfsnvme-unit-d'\n",
+ 'within a minute it fails',
+ );
+ is_deeply([@ACTIVATIONS, @CONNECTS], [], 'at once, without a target call or a connect');
+ ($res, $error) = $activate->('zfsnvme-unit-d', $scfg, $forced->('zfsnvme-unit-d'));
+ is_deeply(
+ [scalar(@ACTIVATIONS), scalar(@CONNECTS)],
+ [1, 2],
+ 'a forced activation tries again',
+ );
+ {
+ local $NOW = $NOW + 61;
+ @CONNECTS = ();
+ ($res, $error) = $activate->('zfsnvme-unit-d', $scfg);
+ is_deeply(
+ [scalar(@ACTIVATIONS), scalar(@CONNECTS)],
+ [1, 2],
+ 'so does one a minute later',
+ );
+ }
+
+ # A slow path whose target call fails starts the backoff as well.
+ reset_host();
+ add_controller('nvme1', '192.0.2.21', 4420, 'ens19');
+ local $FILES{'/etc/pve/priv/storage/zfsnvme-unit-t.nvme-dhchap'} = $KEY;
+ local %CONFIG = ('zfsnvme-unit-t' => $scfg);
+ my $calls = 0;
+ $nvme_mock->redefine(
+ _nvmet_activate_target => sub(@args) {
+ $calls++;
+ die "NVMe target '192.0.2.10' is unreachable: timeout\n";
+ },
+ );
+ eval { $PLUGIN->activate_storage('zfsnvme-unit-t', $scfg) };
+ like($@, qr/is unreachable/, 'an unreachable target fails the slow path');
+ eval { $PLUGIN->activate_storage('zfsnvme-unit-t', $scfg) };
+ is($@, '', 'within a minute the live path takes the fast path');
+ is($calls, 1, 'without another target call');
+ $nvme_mock->unmock('_nvmet_activate_target');
+};
+
+subtest 'connecting the paths' => sub {
+ reset_host();
+ add_controller('nvme1', '192.0.2.21', 4420, 'ens19');
+ $network_mock->redefine(
+ tcp_ping => sub($host, $port, $timeout = undef) { return $host ne '192.0.2.22' },
+ );
+ my (undef, $error) = $activate->('zfsnvme-unit-v', scfg(), $forced->('zfsnvme-unit-v'));
+ is_deeply(\@CONNECTS, [], 'an unreachable portal is not connected');
+ like(join("\n", @WARNINGS), qr/'192\.0\.2\.22:4420' is unreachable/, 'but warned about');
+ $network_mock->redefine(tcp_ping => sub($host, $port, $timeout = undef) { return 1 });
+
+ # After a target restart, the kernel reconnects its controllers itself.
+ reset_host();
+ add_controller('nvme1', '192.0.2.21', 4420, 'ens19', 'connecting');
+ local %REFUSE = ('192.0.2.22' => 1);
+ my $sleeps = 0;
+ $nvme_mock->redefine(
+ _sleep => sub($seconds) {
+ $NOW += $seconds;
+ $FILES{'/sys/class/nvme/nvme1/state'} = 'live' if ++$sleeps == 3;
+ return;
+ },
+ );
+ (undef, $error) = $activate->('zfsnvme-unit-w', scfg());
+ is($error, '', 'an activation waits for a controller that reconnects');
+ is($sleeps, 3, 'until it is live');
+ $nvme_mock->redefine(_sleep => sub($seconds) { $NOW += $seconds; return });
+};
+
+subtest 'moving a path to its configured interface' => sub {
+ my $single = scfg('nvme-portals' => '192.0.2.21', 'nvme-host-ifaces' => 'ens19');
+
+ reset_host();
+ add_controller('nvme1', '192.0.2.21', 4420, 'eth9');
+ my ($res, $error) = $activate->('zfsnvme-unit-c', scfg());
+ is($error, '', 'a storage whose only live path uses another interface');
+ is_deeply(\@CONNECTS, ['192.0.2.21@ens19', '192.0.2.22@ens20'], 'connects both paths');
+ is_deeply(\@DELETED, ['nvme1'], 'and then disconnects the old one');
+ is_deeply(\@WARNINGS, [], 'the storage is healthy');
+
+ reset_host();
+ add_controller('nvme1', '192.0.2.21', 4420, 'eth9');
+ local %CONNECT_STATE = ('192.0.2.21' => 'connecting');
+ my $start = $NOW;
+ ($res, $error) = $activate->('zfsnvme-unit-g', $single);
+ is($error, '', 'a new path that does not become live');
+ is_deeply(\@DELETED, [], 'keeps the old path');
+ like(
+ join("\n", @WARNINGS),
+ qr/'192\.0\.2\.21:4420' is not live on 'ens19' yet; keeping its path on another interface/,
+ 'and warns',
+ );
+ like(join("\n", @WARNINGS), qr/is degraded: 0\/1 paths live/, 'the storage is degraded');
+ is($NOW - $start, 10, 'after waiting 10 seconds for the new path');
+ ($res, $error) = $activate->('zfsnvme-unit-g', $single);
+ ok(!$error && !@ACTIVATIONS, 'within a minute, the old path takes the fast path');
+
+ reset_host();
+ add_controller('nvme1', '192.0.2.21', 4420, 'eth9');
+ local %REFUSE = ('192.0.2.21' => 1);
+ ($res, $error) = $activate->('zfsnvme-unit-h', $single);
+ is($error, '', 'a new path that cannot connect');
+ is_deeply(\@DELETED, [], 'keeps the old path too');
+ like(
+ join("\n", @WARNINGS),
+ qr/192\.0\.2\.21:4420: connect failed: Connection refused/,
+ 'and warns',
+ );
+
+ reset_host();
+ add_controller('nvme1', '192.0.2.21', 4420, 'eth9');
+ my $connect = $PLUGIN->can('_connect_portal');
+ {
+ no warnings 'redefine';
+ local %REFUSE = ();
+ local *PVE::Storage::ZFSNVMePlugin::_connect_portal = sub($scfg, $portal, @ids) {
+ $connect->($scfg, $portal, @ids);
+ die "connection failed after creation\n";
+ };
+ ($res, $error) = $activate->('zfsnvme-unit-h-partial', $single);
+ }
+ ok(!$error && $res == 1, 'a partially successful connect leaves the storage usable');
+ is($FILES{'/sys/class/nvme/nvme71/state'}, 'live', 'a live replacement was created');
+ is_deeply(\@DELETED, [], 'the connect error still preserves the old interface path');
+ like(
+ join("\n", @WARNINGS),
+ qr/192\.0\.2\.21:4420: connection failed after creation/,
+ 'and preserves the original connect warning',
+ );
+};
+
+subtest 'reconnect timeouts reach connected controllers' => sub {
+ my $sysfs = sub () {
+ return [
+ map { [$_->[0] =~ s{\A/sys/class/nvme/}{}r, $_->[1] =~ s/\n\z//r] }
+ grep { $_->[0] =~ m{\A/sys/class/nvme/} } @SYSFS_WRITES
+ ];
+ };
+ my $case = sub($storeid, $scfg, %current) {
+ reset_host();
+ add_controller('nvme1', '192.0.2.21', 4420, 'ens19');
+ add_controller('nvme2', '192.0.2.22', 4420, 'ens20');
+ for my $name (qw(nvme1 nvme2)) {
+ $FILES{"/sys/class/nvme/$name/$_"} = $current{$_} for keys %current;
+ }
+ my ($res, $error) = $activate->($storeid, $scfg, $forced->($storeid));
+ is($error, '', "activation of $storeid");
+ return $sysfs->();
+ };
+ my $timeouts = scfg(
+ 'nvme-reconnect-delay' => 5,
+ 'nvme-ctrl-loss-tmo' => 60,
+ 'nvme-fast-io-fail-tmo' => 10,
+ );
+ my $written = [
+ map { (
+ ["$_/reconnect_delay", 5], ["$_/ctrl_loss_tmo", 60], ["$_/fast_io_fail_tmo", 10],
+ ) } qw(nvme1 nvme2)
+ ];
+ is_deeply($case->('zfsnvme-unit-i', $timeouts), $written, 'changes reach every controller');
+ is_deeply(
+ [grep { $_->[0] =~ /iopolicy/ } @SYSFS_WRITES],
+ [['/sys/class/nvme-subsystem/nvme-subsys7/iopolicy', "round-robin\n"]],
+ 'with the multipath policy',
+ );
+ my %current = (reconnect_delay => '5', ctrl_loss_tmo => '60', fast_io_fail_tmo => '10');
+ is_deeply(
+ $case->('zfsnvme-unit-j', $timeouts, %current),
+ [],
+ 'unchanged ones are not written',
+ );
+ is_deeply(
+ $case->('zfsnvme-unit-k', scfg('nvme-ctrl-loss-tmo' => -1)),
+ [map { ["$_/ctrl_loss_tmo", -1] } qw(nvme1 nvme2)],
+ 'an infinite loss timeout',
+ );
+ is_deeply(
+ $case->('zfsnvme-unit-l', scfg('nvme-ctrl-loss-tmo' => 7), ctrl_loss_tmo => '8'),
+ [],
+ 'the loss timeout the kernel rounds up to the reconnect delay',
+ );
+ is_deeply(
+ $case->('zfsnvme-unit-m', scfg(), fast_io_fail_tmo => '5'),
+ [map { ["$_/fast_io_fail_tmo", -1] } qw(nvme1 nvme2)],
+ 'an unset fast I/O fail timeout is turned off',
+ );
+
+ reset_host();
+ add_controller('nvme1', '192.0.2.21', 4420, 'ens19');
+ add_controller('nvme2', '192.0.2.22', 4420, 'ens20');
+ my ($res, $error) = $activate->('zfsnvme-unit-n', $timeouts);
+ ok(!$error && $res == 1, 'a healthy storage activates on the fast path');
+ is(scalar(@ACTIVATIONS) + scalar(@CONNECTS), 0, 'without a target call or connect');
+ is_deeply($sysfs->(), $written, 'which also writes changed timeouts');
+ is_deeply(
+ [grep { $_->[0] =~ /iopolicy/ } @SYSFS_WRITES],
+ [['/sys/class/nvme-subsystem/nvme-subsys7/iopolicy', "round-robin\n"]],
+ 'and the multipath policy',
+ );
+ delete $FILES{'/etc/pve/priv/storage/zfsnvme-unit-n.nvme-dhchap'};
+ $res = eval {
+ local %CONFIG = ('zfsnvme-unit-n' => $timeouts);
+ $PLUGIN->activate_storage('zfsnvme-unit-n', $timeouts);
+ };
+ is($@, "missing NVMe DH-HMAC-CHAP key\n", 'the fast path needs the key file');
+
+ # file_write warns about a failed write, and is silent about a failed open,
+ # leaving its errno in $!
+ my @perl_warnings;
+ local $SIG{__WARN__} = sub($warning) { push @perl_warnings, $warning };
+ my $failing = sub($pattern, $errno, $warn) {
+ $sysfs_mock->redefine(
+ file_write => sub($path, $data, @rest) {
+ return $sysfs_write->($path, $data) if $path !~ $pattern;
+ $! = $errno;
+ return undef if !$warn;
+ warn "error writing '$data' to '$path': $!\n";
+ $! = ENOENT; # not the reason: file_write closed the file since
+ return 0;
+ },
+ );
+ };
+ my $reason = sub($errno) { local $! = $errno; return "$!" };
+ reset_host();
+ add_controller('nvme1', '192.0.2.21', 4420, 'ens19');
+ add_controller('nvme2', '192.0.2.22', 4420, 'ens20');
+ $failing->(qr{/nvme1/ctrl_loss_tmo\z}, EINVAL, 1);
+ ($res, $error) = $activate->('zfsnvme-unit-e', $timeouts);
+ is($error, '', 'a timeout that cannot be written does not fail the activation');
+ my $einval = $reason->(EINVAL);
+ is_deeply(
+ [@WARNINGS, @perl_warnings],
+ ["cannot set ctrl_loss_tmo of NVMe controller 'nvme1': $einval"],
+ 'and is reported once, as a task warning with the reason',
+ );
+ reset_host();
+ @perl_warnings = ();
+ add_controller('nvme1', '192.0.2.21', 4420, 'ens19');
+ add_controller('nvme2', '192.0.2.22', 4420, 'ens20');
+ $failing->(qr{/nvme2/ctrl_loss_tmo\z}, EACCES, 0);
+ ($res, $error) = $activate->('zfsnvme-unit-tmo-open', $timeouts);
+ is($error, '', 'a timeout that cannot be opened does not fail the activation');
+ is_deeply(
+ [@WARNINGS, @perl_warnings],
+ ["cannot set ctrl_loss_tmo of NVMe controller 'nvme2': " . $reason->(EACCES)],
+ 'and is reported once, with the reason',
+ );
+ reset_host();
+ @perl_warnings = ();
+ add_controller('nvme1', '192.0.2.21', 4420, 'ens19');
+ add_controller('nvme2', '192.0.2.22', 4420, 'ens20');
+ $failing->(qr{/iopolicy\z}, EACCES, 0);
+ ($res, $error) = $activate->('zfsnvme-unit-f', $timeouts);
+ is(
+ $error,
+ 'unable to set NVMe multipath policy: ' . $reason->(EACCES) . "\n",
+ 'a policy that cannot be set fails the activation',
+ );
+ is_deeply([@WARNINGS, @perl_warnings], [], 'with the reason in the error only');
+ reset_host();
+ add_controller('nvme1', '192.0.2.21', 4420, 'ens19');
+ add_controller('nvme2', '192.0.2.22', 4420, 'ens20');
+ $failing->(qr{/iopolicy\z}, EINVAL, 1);
+ ($res, $error) = $activate->('zfsnvme-unit-policy-write', $timeouts);
+ is(
+ $error,
+ "unable to set NVMe multipath policy: $einval\n",
+ 'a policy that cannot be written fails the activation',
+ );
+ is_deeply([@WARNINGS, @perl_warnings], [],
+ 'and the warning of file_write is not passed on');
+
+ # The stock file_write on an attribute that does not exist: unlike a
+ # deletion or a rescan, setting a timeout or the policy treats the failed
+ # open (ENOENT) as a failure like any other.
+ my $file_write = $sysfs_mock->original('file_write');
+ my $missing = tempdir(CLEANUP => 1) . '/a/b';
+ my $absent = sub($pattern) {
+ $sysfs_mock->redefine(
+ file_write => sub($path, $data, @rest) {
+ return $sysfs_write->($path, $data) if $path !~ $pattern;
+ return $file_write->($missing, $data, @rest);
+ },
+ );
+ };
+ my $enoent = $reason->(ENOENT);
+ reset_host();
+ @perl_warnings = ();
+ add_controller('nvme1', '192.0.2.21', 4420, 'ens19');
+ add_controller('nvme2', '192.0.2.22', 4420, 'ens20');
+ $absent->(qr{/nvme1/reconnect_delay\z});
+ ($res, $error) = $activate->('zfsnvme-unit-tmo-missing', $timeouts);
+ is($error, '', 'a missing timeout attribute does not fail the activation');
+ is_deeply(
+ [@WARNINGS, @perl_warnings],
+ ["cannot set reconnect_delay of NVMe controller 'nvme1': $enoent"],
+ 'but is reported once, as a task warning with the reason',
+ );
+ reset_host();
+ @perl_warnings = ();
+ add_controller('nvme1', '192.0.2.21', 4420, 'ens19');
+ add_controller('nvme2', '192.0.2.22', 4420, 'ens20');
+ $absent->(qr{/iopolicy\z});
+ ($res, $error) = $activate->('zfsnvme-unit-policy-missing', $timeouts);
+ is(
+ $error,
+ "unable to set NVMe multipath policy: $enoent\n",
+ 'a missing policy attribute fails the activation',
+ );
+ is_deeply([@WARNINGS, @perl_warnings], [], 'with the reason in the error only');
+ $sysfs_mock->redefine(file_write => $sysfs_write);
+};
+
+subtest 'the key attributes of the controllers are private' => sub {
+ my $attrs = sub(@names) {
+ my @attrs = qw(dhchap_secret dhchap_ctrl_secret);
+ return [
+ map {
+ my $name = $_;
+ map { "/sys/class/nvme/$name/$_" } @attrs
+ } @names
+ ];
+ };
+ reset_host();
+ add_controller('nvme1', '192.0.2.21', 4420, 'ens19');
+ add_controller('nvme2', '192.0.2.22', 4420, 'ens20');
+ add_controller('nvme3', '192.0.2.23', 4420, 'ens21', 'live', $OTHER_NQN);
+ my ($res, $error) = $activate->('zfsnvme-unit-o', scfg());
+ is($error, '', 'a healthy storage');
+ is_deeply(
+ \@RESTRICTED,
+ $attrs->('nvme1', 'nvme2'),
+ 'restricts its controllers on the fast path',
+ );
+
+ reset_host();
+ add_controller('nvme1', '192.0.2.21', 4420, 'ens19');
+ ($res, $error) = $activate->('zfsnvme-unit-p', scfg(), $forced->('zfsnvme-unit-p'));
+ is($error, '', 'a forced activation');
+ is_deeply(
+ \@RESTRICTED,
+ $attrs->('nvme1', 'nvme1', 'nvme71'),
+ 'restricts those of the existing controllers first, then those of new ones',
+ );
+
+ reset_host();
+ local %CONNECT_STATE = map { ($_ => 'connecting') } '192.0.2.21', '192.0.2.22';
+ ($res, $error) = $activate->('zfsnvme-unit-q', scfg());
+ is(
+ $error,
+ "no live NVMe/TCP path for storage 'zfsnvme-unit-q'\n",
+ 'an activation that fails',
+ );
+ is_deeply(
+ \@RESTRICTED,
+ $attrs->('nvme71', 'nvme71', 'nvme72'),
+ 'each one after its connect',
+ );
+
+ reset_host();
+ add_controller('nvme1', '192.0.2.21', 4420, 'ens19', 'connecting');
+ $nvme_mock->redefine(_nvmet_activate_target => sub(@args) { die "unreachable\n" });
+ $res = eval {
+ local $FILES{'/etc/pve/priv/storage/zfsnvme-unit-x.nvme-dhchap'} = $KEY;
+ local %CONFIG = ('zfsnvme-unit-x' => scfg());
+ $PLUGIN->activate_storage('zfsnvme-unit-x', scfg());
+ };
+ is($@, "unreachable\n", 'a slow path that fails at the target');
+ is_deeply(\@RESTRICTED, $attrs->('nvme1'), 'still restricts the existing controllers');
+ $nvme_mock->unmock('_nvmet_activate_target');
+
+ # The first connect fails after it created a live controller; the scan of
+ # the permissions after it fails too in the second case.
+ my $connect = $PLUGIN->can('_connect_portal');
+ my $connect_error = "connection failed after creation\n";
+ for my $scan_error (0, 1) {
+ reset_host();
+ my $scan_failed = 0;
+ {
+ no warnings 'redefine';
+ local *PVE::Storage::ZFSNVMePlugin::_connect_portal = sub($scfg, $portal, @ids) {
+ is_deeply(
+ \@RESTRICTED,
+ $attrs->('nvme71'),
+ 'a failed connect restricts both attributes before the next connect',
+ ) if @CONNECTS == 1 && !$scan_error;
+ $connect->($scfg, $portal, @ids);
+ die $connect_error if @CONNECTS == 1;
+ };
+ local *PVE::Storage::ZFSNVMePlugin::_restrict_attr = sub($path) {
+ push @RESTRICTED, $path;
+ eval { 1 }; # A permission helper must not erase the saved connect error.
+ die "permission scan failed\n" if $scan_error && !$scan_failed++;
+ return;
+ };
+ ($res, $error) = $activate->("zfsnvme-unit-secret-failure-$scan_error", scfg());
+ }
+ is($error, '', 'a connect error does not discard its live controller');
+ ok(
+ scalar(grep { $_ eq "192.0.2.21:4420: $connect_error" } @WARNINGS),
+ 'and warns about it',
+ );
+ if ($scan_error) {
+ ok(
+ scalar(
+ grep { $_ eq "cannot restrict NVMe controller attributes for '$NQN'" }
+ @WARNINGS
+ ),
+ 'the scan failure has its own warning',
+ );
+ unlike(join("\n", @WARNINGS), qr/permission scan failed/, 'without its exception');
+ } else {
+ is_deeply(
+ \@RESTRICTED,
+ $attrs->('nvme71', 'nvme71', 'nvme72'),
+ 'each attempt restricts the controllers, including a failed last connect',
+ );
+ }
+ }
+
+ my $dir = tempdir(CLEANUP => 1);
+ my $restrict = $nvme_mock->original('_restrict_attr');
+ for my $mode (0644, 0640, 0600, 0400) {
+ my $file = sprintf('%s/attr-%o', $dir, $mode);
+ open(my $fh, '>', $file) or die "open: $!\n";
+ close($fh);
+ chmod($mode, $file);
+ $restrict->($file);
+ is(
+ (stat($file))[2] & 07777,
+ $mode & 077 ? 0600 : $mode,
+ sprintf('an attribute with mode %o', $mode),
+ );
+ }
+ @WARNINGS = ();
+ $restrict->("$dir/missing");
+ is_deeply(\@WARNINGS, [], 'a missing attribute is skipped');
+ my $not_directory = "$dir/attr-600/child";
+ $restrict->($not_directory);
+ like(join("\n", @WARNINGS), qr/\Q$not_directory\E/, 'another stat failure warns');
+};
+
+subtest 'controllers match IPv6 portals in any spelling' => sub {
+ reset_host();
+ add_controller('nvme1', '2001:db8:0:0::11', 4420, 'ens19');
+ add_controller('nvme2', '2001:DB8::12', 4420, 'ens20');
+ my $scfg = scfg('nvme-portals' => '[2001:db8::11],[2001:db8::0:12]');
+ my ($res, $error) = $activate->('zfsnvme-unit-r', $scfg);
+ is($error, '', 'both paths are live');
+ is(scalar(@ACTIVATIONS) + scalar(@CONNECTS), 0, 'without a target call or connect');
+};
+
+is_deeply(\@COMMANDS, [], 'the host side ran no command');
+
+done_testing();
+
+1;
^ permalink raw reply related [flat|nested] 7+ messages in thread
* [PATCH storage v3 3/4] zfsnvme: fence target commands of abandoned transactions
2026-10-05 0:26 [PATCH storage v3 0/4] add ZFS over NVMe/TCP storage plugin Joaquin Varela
2026-10-05 0:26 ` [PATCH storage v3 1/4] zfsnvme: " Joaquin Varela
2026-10-05 0:26 ` [PATCH storage v3 2/4] test: add zfsnvme plugin tests Joaquin Varela
@ 2026-10-05 0:26 ` Joaquin Varela
2026-10-05 0:26 ` [PATCH storage v3 4/4] zfsnvme: wait up to 30 seconds for the shared storage lock Joaquin Varela
` (2 subsequent siblings)
5 siblings, 0 replies; 7+ messages in thread
From: Joaquin Varela @ 2026-10-05 0:26 UTC (permalink / raw)
To: pve-devel
ssh does not stop a command on the target when the client gives up on
it. A zfs command that outlives its connection, a chain whose connection
broke in the middle, or a call of a task that was stopped keeps running
on the target after the plugin reported the outcome as unknown. The
pmxcfs domain lock cannot see it: cfs_lock aborts the callback after 60
seconds and pmxcfs lets another node take the lock after 120 seconds,
but neither stops the command on the target. The next transaction can
then interleave its own commands with the rest of the abandoned one,
and a late command of an old plan could change a zvol that a newer plan
already replaced under the same name.
Wrap every call of a transaction in a fixed POSIX sh guard that uses
flock(1) on /run/pve-storage-nvmet/lock and an owner file next to it.
Each transaction first claims the target: it reads the sha256 digest of
the current owner file without the flock, then, under the flock, writes
its own random token only if the owner still has that digest
(compare-and-set), so an abandoned claim that reaches the target late
cannot replace an owner that took over in between. Every later call of
the transaction, reads included, takes the same flock and checks the
token before its commands run. A command still running under the flock
thus drains before a new owner can claim, and a command queued by a
superseded owner exits with code 73 instead of running, as does a claim
whose compare-and-set fails, which is reported as a target that another
transaction took over. Every guarded call first checks the directory:
if /run/pve-storage-nvmet is a symlink or not a root-owned directory
with mode 0700, or its lock or owner entry is a symlink or exists as
anything but a regular file, the call exits with code 74 before the
flock and the error names that directory problem; only a call that
carried a token is reported as superseded. A claim that cannot get the
flock within 30 seconds exits with code 200 and is reported as "target
busy", so a queue of transactions does not look like an unknown
mutation outcome. Calls outside a transaction are sent unchanged and
never wait for a long mutation.
The guard is the only shell control flow on the target: a directory
check, an if/else in the owner read, a command substitution for stat and
one for the owner digest (kept in a variable, so a failing sha256sum is
not taken for a match), and 'flock ... /bin/sh -c' around the unchanged
command chains. It needs flock (util-linux), stat and sha256sum
(coreutils) on the target. Chunking now sizes the wrapped call. A fenced
call aborts the operation ("superseded by a newer one; the target state
is unknown") and is never compensated; the next activation repairs from
the target state as before.
This is control-plane exclusion only: it fences stale commands of this
plugin on the target, not hosts or guest I/O, which remain the job of
HA and of the NVMe controller timeouts.
The emulated target gains the owner, claim and token bookkeeping. New
tests cover a callback fenced after its pmxcfs lock expired, a claim
that timed out, a claim that finds another owner, a busy claim during
alloc_image, unusable owner observations, an unusable lock directory,
exit code 73 outside a transaction, the exit code classification of the
runner, and the guard itself with /bin/sh in a temporary directory,
which now also needs flock, stat and timeout. In the existing create
test, the delayed create is now fenced instead of landing after the
second allocation.
Signed-off-by: Joaquin Varela <joaquinvarela@neatech.ar>
---
src/PVE/Storage/ZFSNVMePlugin.pm | 159 ++++++++++-
src/test/zfsnvme_target_test.pm | 476 ++++++++++++++++++++++++++++++-
src/test/zfsnvme_test.pm | 43 +++
3 files changed, 652 insertions(+), 26 deletions(-)
diff --git a/src/PVE/Storage/ZFSNVMePlugin.pm b/src/PVE/Storage/ZFSNVMePlugin.pm
index 127680fb..dfd43d43 100644
--- a/src/PVE/Storage/ZFSNVMePlugin.pm
+++ b/src/PVE/Storage/ZFSNVMePlugin.pm
@@ -36,6 +36,13 @@ use base qw(PVE::Storage::Plugin);
# commands, sometimes joined with `&&`, to read or change ZFS/configfs state.
# A pmxcfs domain lock per target serializes the changes within the cluster,
# held like the storage lock of the other shared storage types.
+#
+# A fixed POSIX sh transport guard also uses a target flock and owner token:
+# a transaction claims the target by comparing the observed predecessor
+# before publishing its token, and its later calls, reads included, check
+# that token under the flock. Observations outside transactions and the
+# pre-claim owner observation bypass the flock. The guard excludes stale
+# control-plane commands, not guest I/O or hosts.
# ---------------------------------------------------------------------------
# Constants and regular expressions
@@ -60,6 +67,10 @@ my $nvmet_max_command = 65536;
# Seconds to wait for the target lock while another node changes the target;
# a destroy or rollback can hold it for several seconds.
my $nvmet_lock_wait = 30;
+my $nvmet_lock_dir = '/run/pve-storage-nvmet';
+my $nvmet_fenced = 73; # a superseded token, or a claim whose compare-and-set failed
+my $nvmet_unsafe_dir = 74; # $nvmet_lock_dir cannot be used as the lock directory
+my $nvmet_busy = 200; # reserved for lock acquisition during a claim, never a storage command
# On the pool dataset: the highest NSID handed out for a volume of the pool.
my $nvmet_last_nsid = 'proxmox:nvme-last-nsid';
@@ -77,8 +88,9 @@ my @nvmet_cfs_attrs = qw(
attr_model attr_serial attr_allow_any_host
addr_trtype addr_adrfam addr_traddr addr_trsvcid
);
-# Every command the plugin runs on the target, besides the `dd | tee` pipe of
-# a key write. The configfs read runs find through env to set its locale.
+# Allowlist for ordinary steps. The configfs read runs find through env to set
+# its locale; key writes use a dedicated `dd | tee` pipe. The transport guard
+# additionally uses POSIX sh, flock, stat, sha256sum and shell builtins.
my %nvmet_commands = map { $_ => 1 } qw(
cat chmod env grep ln mkdir modprobe mount printf rm rmdir test zfs
);
@@ -143,6 +155,8 @@ my $RE_BASE_SNAPSHOT = qr{^ (?<base>\S+) \@__base__ $}nxx;
my $RE_UNSIGNED_INTEGER = qr{^ (?<value>\d+) $}nxx;
my $RE_NVMET_UUID = qr{\A [0-9a-fA-F]{8} (?: - [0-9a-fA-F]{4}){3} - [0-9a-fA-F]{12} \z}nxx;
+my $RE_NVMET_OWNER_DIGEST = qr{\A [0-9a-f]{64} \z}nxx;
+my $RE_NVMET_OWNER_SNAPSHOT = qr{\A (?<digest>[0-9a-f]{64}) \x20{2} - \z}nxx;
my $RE_NVMET_NSID = qr{\A [1-9] [0-9]* \z}nxx;
my $RE_NVMET_POOL = qr{\A [A-Za-z0-9] [A-Za-z0-9_.:/-]* \z}nxx;
my $RE_NVMET_DATASET_NAME = qr{\A [A-Za-z0-9] [A-Za-z0-9_.:-]* \z}nxx;
@@ -469,7 +483,8 @@ my sub nvmet_quote($word) {
# Renders steps into one POSIX shell command line: simple commands joined with
# `&&`, plus `>` for configfs writes and one `dd | tee` pipe for a key read from
-# stdin. There are no loops, variables, conditionals or substitutions. Every
+# stdin. This renderer has no loops, variables, conditionals or substitutions
+# in its output; _nvmet_render_call adds the separate transport guard. Every
# word is quoted; all operands are built from validated configuration.
sub _nvmet_render($steps) {
my $invalid = "internal error: invalid NVMe target step\n";
@@ -514,16 +529,19 @@ sub _nvmet_render($steps) {
return join(' && ', @commands);
}
-# Packs whole units into calls of at most $max bytes of rendered command, the
-# single argument sshd runs, below its limit. A rendered command is ASCII
-# (every operand is validated), so its length is its size in bytes.
+# Packs whole units into calls of at most $max bytes, including the transport
+# guard and the outer shell quoting, below sshd's single-argument limit. A
+# rendered command is ASCII (every operand is validated), so its length is
+# its size in bytes.
sub _nvmet_chunk($units, $max = $nvmet_max_command) {
my (@chunks, @current);
+ my $sizing_token = '00000000-0000-4000-8000-000000000000';
for my $unit ($units->@*) {
my @next = (@current, $unit->@*);
- if (length(_nvmet_render(\@next)) > $max) {
+ if (length(_nvmet_render_call(\@next, token => $sizing_token)) > $max) {
die "internal error: NVMe target command too long\n"
- if !@current || length(_nvmet_render($unit)) > $max;
+ if !@current
+ || length(_nvmet_render_call($unit, token => $sizing_token)) > $max;
push @chunks, [@current];
@next = $unit->@*;
}
@@ -533,6 +551,78 @@ sub _nvmet_chunk($units, $max = $nvmet_max_command) {
return \@chunks;
}
+# Transport fencing, not target-side storage logic. The fixed target lock is
+# inherited by sh and its children, so a command outliving SSH still excludes
+# the next owner. Token checks reject commands queued by a superseded owner.
+# Claims compare their observed predecessor, so an abandoned claim cannot
+# replace an owner which took over after that observation.
+sub _nvmet_render_call($steps, %opts) {
+ # Observations outside a transaction must not wait for a long-running
+ # mutation. Transactional reads still carry a token and take the lock.
+ return _nvmet_render($steps)
+ if !$opts{read_owner} && !defined($opts{claim}) && !defined($opts{token});
+
+ my $dir = nvmet_quote($nvmet_lock_dir);
+ my $lock = nvmet_quote("$nvmet_lock_dir/lock");
+ my $owner = nvmet_quote("$nvmet_lock_dir/owner");
+ my $token = $opts{claim} // $opts{token};
+ die "internal error: invalid NVMe target token\n"
+ if defined($token) && $token !~ $RE_NVMET_UUID;
+ # /run is root-owned. Never follow a pre-existing symlink or reuse an
+ # incorrectly owned/mode directory; never unlink the lock inode. A
+ # directory that cannot be used is a target problem, not a fencing verdict.
+ my $prepare =
+ "umask 077; (mkdir -m 0700 $dir 2>/dev/null || test -d $dir)"
+ . " && test ! -L $dir && test -d $dir"
+ . " && test \"\$(stat -c '%u:%a' $dir)\" = '0:700'"
+ . " && test ! -L $lock && (test ! -e $lock || test -f $lock)"
+ . " && test ! -L $owner && (test ! -e $owner || test -f $owner)"
+ . " || exit $nvmet_unsafe_dir; ";
+ if ($opts{read_owner}) {
+ die "internal error: NVMe owner observation has commands or a token\n"
+ if $steps->@* || defined($token);
+ # A failed hash/read is not an absent owner. The digest also represents
+ # empty files and arbitrary non-token content without interpreting
+ # either as a valid transaction token.
+ return $prepare
+ . "if test -e $owner; then sha256sum < $owner; else printf '%s\\n' missing; fi";
+ }
+ my $body;
+ if (defined($opts{claim})) {
+ die "internal error: NVMe target claim has commands\n" if $steps->@*;
+ my $expected = $opts{expected_owner};
+ die "internal error: invalid NVMe target predecessor\n"
+ if !defined($expected)
+ || ($expected ne 'missing' && $expected !~ $RE_NVMET_OWNER_DIGEST);
+ my $compare =
+ $expected eq 'missing'
+ ? "test ! -e $owner"
+ : "current=\$(sha256sum < $owner) && test \"\$current\" = "
+ . nvmet_quote("$expected -");
+ # The claim has no storage commands. Normalize its only operation's
+ # failure so it cannot be mistaken for flock's conflict exit code.
+ $body =
+ "$compare || exit $nvmet_fenced; printf '%s\\n' "
+ . nvmet_quote($token)
+ . " > $owner || exit 1";
+ } else {
+ $body = _nvmet_render($steps);
+ $body = 'grep -Fqx -- ' . nvmet_quote($token) . " $owner || exit $nvmet_fenced; $body"
+ if defined($token);
+ }
+ my $conflict = defined($opts{claim}) ? "-E $nvmet_busy " : '';
+ return
+ $prepare
+ . "flock -x -w $nvmet_lock_wait $conflict$lock /bin/sh -c "
+ . nvmet_quote($body);
+}
+
+sub _nvmet_new_token() {
+ my $token = lc(file_read_firstline('/proc/sys/kernel/random/uuid') // '');
+ die "cannot generate NVMe target transaction token\n" if $token !~ $RE_NVMET_UUID;
+ return $token;
+}
+
# Never propagate the context run_command adds to the task marker: it quotes
# the command line.
my sub rethrow_task_interrupt($error) {
@@ -543,10 +633,16 @@ my sub rethrow_task_interrupt($error) {
# { rc => exit code, or -1 when ssh did not run to its end, out => [lines],
# err => text } and never dies on a failing command; a stopped task dies with
# the task marker. %opts: op (label, required), timeout and input (stdin,
-# used for keys).
+# used for keys); the transport guard options of _nvmet_render_call:
+# read_owner, claim with expected_owner, and token. A guarded call that
+# finds the lock directory unusable (exit $nvmet_unsafe_dir) gets the
+# directory problem as its error text; a claim that could not acquire the
+# flock (exit $nvmet_busy) is reported as target busy, and one whose
+# compare-and-set failed (exit $nvmet_fenced) as a target that another
+# transaction took over, which a repeated operation claims anew.
sub _nvmet_run($scfg, $steps, %opts) {
die "internal error: NVMe target call without label\n" if !defined($opts{op});
- my $command = _nvmet_render($steps);
+ my $command = _nvmet_render_call($steps, %opts);
die "internal error: NVMe target command too long\n" if length($command) > $nvmet_max_command;
my $cmd = [@ssh_cmd, '-i', nvmet_ssh_key($scfg), 'root@' . nvmet_server($scfg), $command];
my (@out, $err);
@@ -568,6 +664,15 @@ sub _nvmet_run($scfg, $steps, %opts) {
return { rc => -1, out => [], err => 'ssh failed' } if $error !~ $RE_COMMAND_EXIT;
$rc = $+{code};
}
+ $err =
+ "unsafe lock directory $nvmet_lock_dir on the target: it must be a root-owned"
+ . " directory with mode 0700 whose lock and owner entries are regular files"
+ if ($opts{read_owner} || defined($opts{claim}) || defined($opts{token}))
+ && $rc == $nvmet_unsafe_dir;
+ $err = 'target busy: could not acquire transaction lock'
+ if defined($opts{claim}) && $rc == $nvmet_busy;
+ $err = 'another transaction took over the target'
+ if defined($opts{claim}) && $rc == $nvmet_fenced;
return { rc => $rc, out => \@out, err => $err // ($rc ? "exit code $rc" : '') };
}
@@ -1096,10 +1201,12 @@ sub _nvmet_plan_orphan_hosts($cfs, $candidates) {
# ---------------------------------------------------------------------------
my %nvmet_lock_owner; # lock id => pid of the process holding the domain lock
+my %nvmet_lock_token; # lock id => target-side transaction token
# Runs $code under the pmxcfs domain lock of the target. Like the pmxcfs
# lock itself, it is not re-entrant, and a child forked inside is not the
-# owner.
+# owner. pmxcfs can break a lock after 120 seconds: the target flock drains
+# an executing command before a new token fences all calls of the old owner.
my sub nvmet_locked($scfg, $code) {
my $id = nvmet_lock_id($scfg);
die "cluster not quorate - refusing NVMe target changes\n"
@@ -1108,7 +1215,29 @@ my sub nvmet_locked($scfg, $code) {
$id,
$nvmet_lock_wait,
sub {
+ my $transaction = _nvmet_new_token();
+ my $snapshot = _nvmet_run(
+ $scfg, [],
+ read_owner => 1,
+ op => 'read target transaction owner',
+ );
+ die "cannot observe NVMe target transaction owner: $snapshot->{err}\n"
+ if $snapshot->{rc};
+ my $line = $snapshot->{out}->[0] // '';
+ die "invalid NVMe target transaction owner observation\n"
+ if $snapshot->{out}->@* != 1
+ || ($line ne 'missing' && $line !~ $RE_NVMET_OWNER_SNAPSHOT);
+ my $expected = $line eq 'missing' ? $line : $+{digest};
+ my $claim = _nvmet_run(
+ $scfg, [],
+ claim => $transaction,
+ expected_owner => $expected,
+ op => 'claim target transaction',
+ timeout => $nvmet_lock_wait + 15,
+ );
+ die "cannot claim NVMe target transaction: $claim->{err}\n" if $claim->{rc};
local $nvmet_lock_owner{$id} = $$;
+ local $nvmet_lock_token{$id} = $transaction;
return $code->();
},
);
@@ -1124,8 +1253,9 @@ my sub nvmet_locked($scfg, $code) {
# did not run to its end (-1: its timeout, or ssh was killed), and a call
# during which the task was stopped (it dies with the task marker) may still
# be running there. Its outcome is unknown, so it is never compensated from a
-# read that could come before the rest of it. The next activation repairs
-# from the target state.
+# read that could come before the rest of it. The next activation, under a
+# new token, repairs from the target state. A same-token read cannot settle a
+# lost reply: it could overtake the abandoned command before flock.
my sub nvmet_exec($scfg, $steps, %opts) {
my $locked = ($nvmet_lock_owner{ nvmet_lock_id($scfg) } // 0) == $$;
die "internal error: NVMe target change without target lock\n"
@@ -1134,8 +1264,11 @@ my sub nvmet_exec($scfg, $steps, %opts) {
$scfg, $steps,
op => $opts{op},
timeout => $opts{timeout} // 15,
+ ($locked ? (token => $nvmet_lock_token{ nvmet_lock_id($scfg) }) : ()),
(defined($opts{input}) ? (input => $opts{input}) : ()),
);
+ die "NVMe target transaction was superseded by a newer one; the target state is unknown\n"
+ if $locked && $res->{rc} == $nvmet_fenced;
die "NVMe target operation '$opts{op}' did not complete ($res->{err});"
. " the target state is unknown\n"
if ($locked && $res->{rc} == -1)
diff --git a/src/test/zfsnvme_target_test.pm b/src/test/zfsnvme_target_test.pm
index 991d2084..1b109bfd 100644
--- a/src/test/zfsnvme_target_test.pm
+++ b/src/test/zfsnvme_target_test.pm
@@ -9,6 +9,7 @@ use lib qw(..);
use Compress::Zlib qw(crc32);
use Digest::SHA qw(sha256_hex);
+use Fcntl qw(F_GETFD F_SETFD FD_CLOEXEC LOCK_EX LOCK_NB);
use File::Temp qw(tempdir);
use FindBin;
use IPC::Open3;
@@ -112,6 +113,7 @@ our (%LOCK_HELD, @LOCKS, @NESTED_LOCKS, @QUORUM, @WARNINGS, @SYSFS_WRITES, %FILE
our (%CORPUS, @VIOLATIONS, @WRITES, @UNLINKED, %BLOCK, @UUIDS);
my $uuid_seq = 0;
+my $token_seq = 0;
sub lock_id($scfg) {
return 'zfsnvme-' . ($scfg->{server} =~ s/[^A-Za-z0-9.-]/_/gr);
@@ -145,11 +147,18 @@ my $tools_mock = Test::MockModule->new('PVE::Tools');
$plugin_mock->redefine(
_nvmet_run => sub($scfg, $steps, %opts) {
die "test error: no fake target\n" if !$FAKE;
+ return $FAKE->read_owner() if $opts{read_owner};
+ return $FAKE->claim($opts{claim}, %opts) if defined($opts{claim});
my $rendered = nv('_nvmet_render', $steps);
$CORPUS{$rendered} //= $opts{op} // '';
return $FAKE->run($scfg, $steps, $rendered, %opts);
},
);
+$plugin_mock->redefine(
+ _nvmet_new_token => sub () {
+ return sprintf('ffffffff-ffff-4000-8000-%012x', ++$token_seq);
+ },
+);
# The target is the only place that runs commands, and only through _nvmet_run.
# A command is also a violation, in case the caller handles the error.
my $no_command = sub($cmd, %opts) {
@@ -292,6 +301,11 @@ package FakeTarget {
violations => [],
faults => {}, # call index => before | after | late | cut:<step> | lost:<step>
late => [], # calls that ssh gave up on, applied by land()
+ claims => [],
+ owner => undef,
+ owner_read => undef, # the answer of the owner observation, if not the owner
+ unsafe_dir => 0, # the lock directory of the target is unusable
+ fenced => 0, # every command exits with the fencing code
# { after => call index, mode => before | unreachable | failed }
read_fault => undef,
step_fault => undef, # sub ($step, $call) returning an error text
@@ -398,6 +412,61 @@ package FakeTarget {
# --- execution --------------------------------------------------------
+ sub owner_digest($self) {
+ return 'missing' if !defined($self->{owner});
+ return sha256_hex("$self->{owner}\n");
+ }
+
+ # What _nvmet_run returns for a guarded call that found the lock directory
+ # unusable.
+ sub unsafe_dir_answer($self) {
+ return {
+ rc => 74,
+ out => [],
+ err => 'unsafe lock directory /run/pve-storage-nvmet on the target: it must be'
+ . ' a root-owned directory with mode 0700 whose lock and owner entries are'
+ . ' regular files',
+ };
+ }
+
+ sub read_owner($self) {
+ return $self->unsafe_dir_answer if $self->{unsafe_dir};
+ return $self->{owner_read} if $self->{owner_read};
+ my $owner = $self->owner_digest;
+ return { rc => 0, out => [$owner eq 'missing' ? $owner : "$owner -"], err => '' };
+ }
+
+ sub claim($self, $token, %opts) {
+ push $self->{claims}->@*, $token;
+ my $expected = $opts{expected_owner};
+ $self->violation('claim without expected owner') if !defined($expected);
+ $self->violation("invalid expected owner '$expected'")
+ if $expected ne 'missing' && $expected !~ /\A[0-9a-f]{64}\z/;
+ return $self->unsafe_dir_answer if $self->{unsafe_dir};
+ return {
+ rc => 200,
+ out => [],
+ err => 'target busy: could not acquire transaction lock',
+ }
+ if $self->{claim_busy};
+ if ($self->{claim_fault}) {
+ $self->{late_claim} = { token => $token, expected_owner => $expected };
+ return { rc => -1, out => [], err => 'timeout' };
+ }
+ # A command already inside flock finishes before another owner can
+ # claim. Commands delayed before their guard instead see the new token.
+ $self->land(1);
+ return { rc => 73, out => [], err => 'another transaction took over the target' }
+ if $self->owner_digest ne $expected;
+ $self->{owner} = $token;
+ return { rc => 0, out => [], err => '' };
+ }
+
+ sub replay_late_claim($self) {
+ my $late = $self->{late_claim} // $self->violation('no late claim');
+ return $self->claim($late->{token}, expected_owner => $late->{expected_owner});
+ }
+
sub run($self, $scfg, $steps, $rendered, %opts) {
my $call = {
index => scalar($self->{calls}->@*),
@@ -410,6 +479,7 @@ package FakeTarget {
mutating => (grep { main::mutating_step($_) } $steps->@*) ? 1 : 0,
changing => main::changing($steps),
now => $main::NOW,
+ token => $opts{token},
};
push $self->{calls}->@*, $call;
$self->audit($call, $opts{input});
@@ -452,6 +522,8 @@ package FakeTarget {
if defined($input) && !grep { $input eq "$_\n" } $self->{secrets}->@*;
push @bad, "target change without the target lock in '$op'"
if $call->{mutating} && !$call->{locked};
+ push @bad, "transaction call without a fencing token in '$op'"
+ if $call->{locked} && !defined($call->{token});
$self->flag(@bad);
}
@@ -465,12 +537,22 @@ package FakeTarget {
# Applies the calls that ssh gave up on, as the target finally runs them:
# each chain stops at its first failing step.
- sub land($self) {
+ sub land($self, $active_only = 0) {
+ my @pending;
for my $late (splice($self->{late}->@*)) {
+ if ($active_only && !$late->{guarded}) {
+ push @pending, $late;
+ next;
+ }
+ next
+ if !$late->{guarded}
+ && defined($late->{token})
+ && ($self->{owner} // '') ne $late->{token};
for my $step ($late->{steps}->@*) {
last if !eval { $self->step($step, $late->{input}); 1 };
}
}
+ push $self->{late}->@*, @pending;
return $self;
}
@@ -480,6 +562,7 @@ package FakeTarget {
# still runs, later), or ssh gives up on it while it runs to its end later
# (late: -1).
sub execute($self, $call, $steps, $input) {
+ $self->land(1) if defined($call->{token}); # only transactions drain active commands
my $closed = 'Connection to 192.0.2.10 closed by remote host.';
my $broken = { rc => 255, out => [], err => $closed };
my $no_route = 'ssh: connect to host 192.0.2.10: No route to host';
@@ -497,9 +580,18 @@ package FakeTarget {
$call->{fault} = $mode if $mode;
return { rc => 1, out => [], err => 'injected failure' } if $mode eq 'before';
if ($mode eq 'late') {
- push $self->{late}->@*, { steps => dclone($steps), input => $input };
+ push $self->{late}->@*,
+ {
+ steps => dclone($steps),
+ input => $input,
+ token => $call->{token},
+ };
return { rc => -1, out => [], err => 'timeout' };
}
+ return $self->unsafe_dir_answer if $self->{unsafe_dir} && defined($call->{token});
+ return { rc => 73, out => [], err => 'transaction fenced' }
+ if $self->{fenced}
+ || (defined($call->{token}) && ($self->{owner} // '') ne $call->{token});
my ($cut) = $mode =~ /\Acut:([0-9]+)\z/;
my ($lost) = $mode =~ /\Alost:([0-9]+)\z/;
my @out;
@@ -507,7 +599,13 @@ package FakeTarget {
return $broken if defined($cut) && $index == $cut;
if (defined($lost) && $index == $lost) {
my @rest = $steps->@[$index .. $steps->$#*];
- push $self->{late}->@*, { steps => dclone(\@rest), input => $input };
+ push $self->{late}->@*,
+ {
+ steps => dclone(\@rest),
+ input => $input,
+ token => $call->{token},
+ guarded => 1,
+ };
return $broken;
}
my $step = $steps->[$index];
@@ -2544,7 +2642,7 @@ subtest 'template name and chunking' => sub {
my $chunks = nv('_nvmet_chunk', \@units);
ok($chunks->@* > 1, 'large plans are split');
ok(
- !grep({ length(nv('_nvmet_render', $_)) > 65536 } $chunks->@*),
+ !grep({ length(nv('_nvmet_render_call', $_, token => $U{1})) > 65536 } $chunks->@*),
'no call exceeds 64 KiB',
);
ok(!grep({ $_->[0]->[0] ne 'test' || $_->@* % 6 } $chunks->@*), 'chunks hold whole units');
@@ -2553,7 +2651,7 @@ subtest 'template name and chunking' => sub {
my $small = nv(
'_nvmet_chunk',
[@units[0 .. 3]],
- length(nv('_nvmet_render', [map { $_->@* } @units[0 .. 1]])),
+ length(nv('_nvmet_render_call', [map { $_->@* } @units[0 .. 1]], token => $U{1})),
);
is(scalar($small->@*), 2, 'the limit is configurable');
eval { nv('_nvmet_chunk', [$units[0]], 10) };
@@ -3206,14 +3304,12 @@ subtest 'create' => sub {
$fake->land;
is_deeply(
[map { owned_volumes($fake->{m})->{"tank/vm-$_-disk-0"}->[0] } 200, 201],
- [4, 5],
- 'never gets the NSID of the create that completes afterwards',
+ [undef, 5],
+ 'fences the delayed create and never reuses its reserved NSID',
);
$res = flow($fake, $ACT{activate});
- ok(
- !$res->{error} && ns_of($fake, 4) && ns_of($fake, 5),
- 'and the next activation exports both',
- );
+ ok(!$res->{error} && !ns_of($fake, 4) && ns_of($fake, 5),
+ 'only the second one is exported');
($fake, $res) = run_on(
'alloc',
@@ -3664,6 +3760,66 @@ subtest 'resize' => sub {
is($fake->{m}->{ds}->{'tank/vm-100-disk-0'}->{volsize}, 2 * 1024**3, 'keeps the new size');
};
+subtest 'a new owner fences a still-running cluster-lock callback' => sub {
+ my $fake = lifecycle_fake();
+ my $superseded = 0;
+ $fake->{before_call} = sub($f, $call) {
+ return if $call->{op} ne 'resize zvol' || $superseded++;
+ $NOW += 121;
+ local %LOCK_HELD; # pmxcfs permits the expired lease to be replaced
+ my $new =
+ flow($f, sub { $PLUGIN->volume_resize(scfg(), 'st', 'vm-100-disk-0', 3 * 1024**3) });
+ is($new->{error}, '', 'a second owner completes its fresh plan');
+ };
+ my $res = flow($fake, $ACT{resize});
+ like(
+ $res->{error},
+ qr/transaction was superseded by a newer one; the target state is unknown\n\z/,
+ 'the old callback cannot resume',
+ );
+ is($fake->{m}->{ds}->{'tank/vm-100-disk-0'}->{volsize}, 3 * 1024**3, 'the new size stays');
+ is(scalar($fake->{claims}->@*), 2, 'each callback claims only once');
+};
+
+subtest 'an uncertain claim never starts a transaction' => sub {
+ my $fake = lifecycle_fake(claim_fault => 1);
+ my $before = dclone($fake->{m});
+ my $res = flow($fake, $ACT{resize});
+ like($res->{error}, qr/cannot claim NVMe target transaction/, 'claim timeout aborts');
+ is(
+ scalar($fake->{calls}->@*),
+ 0,
+ 'no planning read or mutation follows the uncertain claim',
+ );
+ is_deeply($fake->{m}, $before, 'storage state is unchanged');
+ $fake->{claim_fault} = 0;
+ is($fake->replay_late_claim->{rc}, 0, 'the abandoned claim can still reach the target');
+ is($fake->{owner}, $fake->{late_claim}->{token}, 'the late claim publishes its token');
+ is_deeply($fake->{m}, $before, 'the accepted late claim changes no storage state');
+ my $next =
+ flow($fake, sub { $PLUGIN->volume_resize(scfg(), 'st', 'vm-100-disk-0', 3 * 1024**3) });
+ is($next->{error}, '', 'a fresh transaction claims the late owner by its digest');
+ isnt($fake->{owner}, $fake->{late_claim}->{token}, 'the fresh transaction replaces it');
+ is($fake->{m}->{ds}->{'tank/vm-100-disk-0'}->{volsize}, 3 * 1024**3, 'with its own resize');
+};
+
+subtest 'a claim that finds another owner' => sub {
+ # the owner changed between the observation and the claim
+ my ($fake, $res) = run_on(
+ 'resize',
+ owner => $U{2},
+ owner_read => { rc => 0, out => [sha256_hex("$U{1}\n") . ' -'], err => '' },
+ );
+ is(
+ $res->{error},
+ "cannot claim NVMe target transaction: another transaction took over the target\n",
+ 'reports the takeover',
+ );
+ is_deeply([scalar($fake->{claims}->@*), $res->{calls}], [1, []], 'after the one claim');
+ is($fake->{owner}, $U{2}, 'the owner that took over stays');
+ is($fake->{m}->{ds}->{'tank/vm-100-disk-0'}->{volsize}, 1024**3, 'and nothing changed');
+};
+
sub removal_fake(%opts) {
my $fake = target_fake(%opts);
$fake->add_zvol('other/foreign-disk', identity => [$FOREIGN_NQN, 1, $U{8}]);
@@ -3876,6 +4032,90 @@ subtest 'snapshots and the generic ZFS calls' => sub {
}
};
+subtest 'a busy claim prevents public allocation' => sub {
+ my $fake = lifecycle_fake(claim_busy => 1, owner => $U{1});
+ my $before = dclone($fake->{m});
+ my $res = flow(
+ $fake, sub { $PLUGIN->alloc_image('st', scfg(), 200, 'raw', 'vm-200-disk-0', 1024) },
+ );
+ is(
+ $res->{error},
+ "cannot claim NVMe target transaction: target busy: could not acquire transaction lock\n",
+ 'public allocation reports busy, not an unknown mutation outcome',
+ );
+ is(scalar($fake->{claims}->@*), 1, 'the claim is attempted once');
+ is_deeply($res->{calls}, [],
+ 'no planning, mutation or compensation follows the busy claim');
+ is($fake->{owner}, $U{1}, 'the existing owner is unchanged');
+ is_deeply($fake->{m}, $before, 'storage state is unchanged');
+};
+
+subtest 'an unusable owner snapshot never starts a claim' => sub {
+ for my $case (
+ [
+ 'a timeout',
+ { rc => -1, out => [], err => 'timeout' },
+ qr/cannot observe NVMe target transaction owner: timeout/,
+ ],
+ [
+ 'a failed hash with matching stdout',
+ { rc => 1, out => [sha256_hex("owner\n") . ' -'], err => 'Input/output error' },
+ qr/cannot observe NVMe target transaction owner: Input\/output error/,
+ ],
+ [
+ 'a malformed snapshot',
+ { rc => 0, out => ['not an owner snapshot'], err => '' },
+ qr/invalid NVMe target transaction owner observation/,
+ ],
+ ) {
+ my ($name, $snapshot, $error) = $case->@*;
+ my ($fake, $res) = run_on('resize', owner_read => $snapshot);
+ like($res->{error}, $error, "$name aborts the transaction");
+ is_deeply([$fake->{claims}->@*, $fake->{calls}->@*], [], 'without a claim or a call');
+ }
+};
+
+subtest 'an unusable lock directory on the target' => sub {
+ my ($fake, $res) = run_on('resize', unsafe_dir => 1);
+ is(
+ $res->{error},
+ 'cannot observe NVMe target transaction owner: '
+ . $fake->unsafe_dir_answer->{err} . "\n",
+ 'the owner observation names the directory problem',
+ );
+ is_deeply([$fake->{claims}->@*, $fake->{calls}->@*], [], 'without a claim or a call');
+
+ # the directory becomes unusable while a transaction runs
+ $fake = lifecycle_fake(
+ before_call => sub($f, $call) { $f->{unsafe_dir} = 1 if $call->{op} eq 'resize zvol' },
+ );
+ $res = flow($fake, $ACT{resize});
+ is(
+ $res->{error},
+ "NVMe target operation 'resize zvol' failed: " . $fake->unsafe_dir_answer->{err} . "\n",
+ 'a later call reports the directory problem, not a superseded transaction',
+ );
+ is($fake->{m}->{ds}->{'tank/vm-100-disk-0'}->{volsize}, 1024**3, 'and changed nothing');
+};
+
+subtest 'exit code 73 is a fencing verdict only for a call with a token' => sub {
+ my $fake = lifecycle_fake(fenced => 1);
+ my $res = flow($fake, sub { [$PLUGIN->path(scfg(), 'vm-100-disk-0', 'st')] });
+ is(
+ $res->{error},
+ "cannot read NVMe target state: transaction fenced\n",
+ 'a read outside a transaction reports a failed command',
+ );
+ is(scalar($res->{locks}->@*), 0, 'it ran without the lock');
+ $res = flow($fake, $ACT{resize});
+ is(
+ $res->{error},
+ "NVMe target transaction was superseded by a newer one; the target state is unknown\n",
+ 'the same exit code fences the read of a transaction',
+ );
+ is_deeply(ops($res), ['read target state'], 'which ends at that read');
+};
+
subtest 'lock ownership' => sub {
my $fake = lifecycle_fake();
pipe(my $reader, my $writer) or die "pipe: $!\n";
@@ -4595,7 +4835,7 @@ sub sh($command, $input = undef) {
# The subtests that run rendered commands need these tools. A Debian build
# has them, so a missing one is an error, never a silently skipped test.
{
- my @tools = qw(find grep sha256sum dd tee);
+ my @tools = qw(find grep sha256sum dd tee flock stat timeout);
BAIL_OUT('needs /bin/sh with ' . join(', ', @tools))
if !-x '/bin/sh' || (sh('command -v ' . join(' ', @tools)))[0] != 0;
}
@@ -4813,9 +5053,219 @@ subtest 'the rendered chains on a configfs stand-in' => sub {
ok(-d "$root/ports/1" && -d "$root/ports/2", 'ports stay');
};
+subtest 'transport guard executes with POSIX sh in an isolated directory' => sub {
+ my $dir = tempdir(CLEANUP => 1);
+ my $runtime = "$dir/runtime";
+ my $output = "$dir/output";
+ my $digest = sub($bytes) { return sha256_hex($bytes) };
+ my $ownership = $> . ':700';
+ my ($signal_rc) = sh('kill -TERM $$');
+ is($signal_rc, 143, 'a signaled shell is not reported as a successful command');
+ my $run = sub($steps, %opts) {
+ my $short_wait = delete $opts{test_short_wait};
+ my $printf_failure = delete $opts{test_printf_failure};
+ my $sha256_failure = delete $opts{test_sha256_failure};
+ my $command = nv('_nvmet_render_call', $steps, %opts);
+ is(system('/bin/sh', '-n', '-c', $command), 0, 'the transport command is POSIX syntax');
+ # Keep production guards, apart from the temporary path/uid and the
+ # explicitly requested short wait or failing command below.
+ # No test command may access the production runtime directory.
+ $command =~ s{\Q/run/pve-storage-nvmet\E}{$runtime}g;
+ $command =~ s/'0:700'/'$ownership'/g;
+ die "test command escaped temporary runtime\n" if $command =~ m{/run/pve-storage-nvmet};
+ if ($short_wait) {
+ $command =~ s/\bflock -x -w 30 /flock -x -w 0.1 /
+ or die "test cannot shorten the target lock wait\n";
+ }
+ if ($printf_failure) {
+ my $prefix = PVE::Tools::shellquote('printf() { return 200; }; ');
+ $command =~ s{(/bin/sh -c )}{$1$prefix}
+ or die "test cannot inject the failing printf\n";
+ }
+ if ($sha256_failure) {
+ my $failure = 'sha256sum() { command sha256sum "$@"; return 200; }; ';
+ if ($opts{read_owner}) {
+ $command = $failure . $command;
+ } else {
+ my $prefix = PVE::Tools::shellquote($failure);
+ $command =~ s{(/bin/sh -c )}{$1$prefix}
+ or die "test cannot inject the failing sha256sum\n";
+ }
+ }
+ # Kill the whole finite command group if a regression blocks on flock.
+ return sh('timeout -k 1 2 /bin/sh -c ' . PVE::Tools::shellquote($command));
+ };
+ for my $option (qw(claim token)) {
+ for my $invalid ('', 0) {
+ eval { nv('_nvmet_render_call', [['cat', $output]], $option => $invalid) };
+ is(
+ $@,
+ "internal error: invalid NVMe target token\n",
+ "$option '$invalid' is refused",
+ );
+ }
+ }
+ for my $case (
+ [['cat', $output], read_owner => 1],
+ [[], read_owner => 1, token => $U{1}],
+ [[], claim => $U{1}],
+ [[], claim => $U{1}, expected_owner => 'not-a-digest'],
+ ) {
+ my ($steps, %opts) = $case->@*;
+ eval { nv('_nvmet_render_call', $steps, %opts) };
+ like(
+ $@,
+ qr/internal error: (?:NVMe owner observation has commands or a token|invalid .*)/,
+ 'owner metadata is validated before rendering',
+ );
+ }
+ my ($owner_rc, $owner_out) = $run->([], read_owner => 1);
+ is($owner_rc, 0, 'an absent owner is observed without the target flock');
+ is($owner_out, "missing\n", 'the initial owner snapshot is missing');
+ my ($rc) = $run->([], claim => $U{1}, expected_owner => 'missing');
+ is($rc, 0, 'the first owner claims');
+ my $value = q{literal 'quotes' and $variables};
+ ($rc) = $run->([write_step($output, $value)], token => $U{1});
+ is($rc, 0, 'the current owner executes a quoted command');
+ is(slurp($output), "$value\n", 'outer quoting preserves literal operands');
+ {
+ open(my $held, '+<', "$runtime/lock") or die "open temporary lock: $!\n";
+ my $flags = fcntl($held, F_GETFD, 0) // die "get lock FD flags: $!\n";
+ fcntl($held, F_SETFD, $flags | FD_CLOEXEC) or die "set lock FD flags: $!\n";
+ flock($held, LOCK_EX | LOCK_NB) or die "hold temporary lock: $!\n";
+ for my $options ({}, { token => undef, claim => undef }) {
+ my ($read_rc, $read_out) = $run->([['cat', $output]], $options->%*);
+ is($read_rc, 0, 'a read completes while another command holds the target flock');
+ is($read_out, "$value\n", 'the read returns the actual file content');
+ }
+ ($owner_rc, $owner_out) = $run->([], read_owner => 1);
+ is($owner_rc, 0, 'an owner snapshot does not wait for a held flock');
+ is($owner_out, $digest->("$U{1}\n") . " -\n", 'the snapshot hashes the owner bytes');
+ ($rc) = $run->(
+ [],
+ claim => $U{2},
+ expected_owner => $digest->("$U{1}\n"),
+ test_short_wait => 1,
+ );
+ is($rc, 200, 'a conflicting claim returns the dedicated busy exit code');
+ ($rc) = $run->([['cat', $output]], token => $U{1}, test_short_wait => 1);
+ is($rc, 1, 'a tokenized call retains the default flock conflict exit code');
+ ($rc) = $run->([write_step($output, 'blocked')], token => $U{1});
+ is($rc, 124, 'a tokenized call still waits for the exclusive target flock');
+ is(slurp("$runtime/owner"), "$U{1}\n", 'the blocked claim did not replace the owner');
+ is(slurp($output), "$value\n", 'the blocked change did not run');
+ close($held) or die "close temporary lock: $!\n";
+ }
+ my ($read_rc, $read_out) = $run->([['cat', $output]], token => $U{1});
+ is($read_rc, 0, 'a tokenized read runs once the exclusive lock is released');
+ is($read_out, "$value\n", 'the tokenized read returns the same file content');
+ ($rc) = $run->([], claim => $U{2}, expected_owner => $digest->("$U{1}\n"));
+ is($rc, 0, 'a new owner claims');
+ ($rc) = $run->([write_step($output, 'stale')], token => $U{1});
+ is($rc, 73, 'an old owner is fenced');
+ is(slurp($output), "$value\n", 'a fenced write does not run');
+ ($rc) = $run->([['printf', '%s', 'unused']], token => $U{2}, test_printf_failure => 1);
+ is($rc, 200, 'a tokenized child exit code 200 is passed through');
+ is(slurp("$runtime/owner"), "$U{2}\n", 'a failed child did not replace the owner');
+ ($rc) = $run->(
+ [],
+ claim => $U{3},
+ expected_owner => $digest->("$U{2}\n"),
+ test_printf_failure => 1,
+ );
+ is($rc, 1, 'a failed claim printf is normalized, never reported as busy');
+ is(slurp("$runtime/owner"), '', 'the failed printf did not publish a new owner');
+ is(slurp($output), "$value\n", 'the failing claim write did no storage work');
+ open(my $owner, '>', "$runtime/owner") or die "open restored owner: $!\n";
+ print {$owner} "$U{2}\n";
+ close($owner) or die "close restored owner: $!\n";
+
+ ($rc) = $run->(
+ [],
+ claim => $U{3},
+ expected_owner => $digest->("$U{2}\n"),
+ test_sha256_failure => 1,
+ );
+ is($rc, 73, 'a failed CAS hash is fenced, never reported as busy');
+ is(slurp("$runtime/owner"), "$U{2}\n", 'a failed CAS hash leaves the owner unchanged');
+ is(slurp($output), "$value\n", 'a failed CAS hash starts no storage work');
+ ($owner_rc, $owner_out) = $run->([], read_owner => 1, test_sha256_failure => 1);
+ is($owner_rc, 200, 'a failed owner snapshot hash preserves its failure status');
+ isnt($owner_out, "missing\n", 'a failed hash is not an absent-owner snapshot');
+ is(slurp("$runtime/owner"), "$U{2}\n", 'a failed snapshot leaves the owner unchanged');
+
+ my $predecessor = $digest->("$U{2}\n");
+ my $late_a = nv('_nvmet_render_call', [], claim => $U{4}, expected_owner => $predecessor);
+ is(system('/bin/sh', '-n', '-c', $late_a), 0, 'the delayed A claim is POSIX syntax');
+ $late_a =~ s{\Q/run/pve-storage-nvmet\E}{$runtime}g;
+ $late_a =~ s/'0:700'/'$ownership'/g;
+ die "test command escaped temporary runtime\n" if $late_a =~ m{/run/pve-storage-nvmet};
+ ($rc) = $run->([], claim => $U{5}, expected_owner => $predecessor);
+ is($rc, 0, 'B claims after the same existing-owner snapshot');
+ my ($late_rc) = sh('timeout -k 1 2 /bin/sh -c ' . PVE::Tools::shellquote($late_a));
+ is($late_rc, 73, 'A cannot overwrite B after its stale existing-owner snapshot');
+ ($rc) = $run->([], claim => $U{4}, expected_owner => 'missing');
+ is($rc, 73, 'a stale missing-owner snapshot cannot overwrite B either');
+ is(slurp("$runtime/owner"), "$U{5}\n", 'both stale snapshots preserve the actual owner');
+ ($rc) = $run->([write_step($output, 'B')], token => $U{5});
+ is($rc, 0, 'B writes after rejecting delayed A');
+ is(slurp($output), "B\n", 'B remains the transaction owner');
+
+ # Any existing owner content, empty included, is observed by its digest
+ # and can only be claimed with that exact digest.
+ for my $content ('', "maintenance:1700000000\n") {
+ open(my $owner, '>', "$runtime/owner") or die "open owner: $!\n";
+ print {$owner} $content;
+ close($owner) or die "close owner: $!\n";
+ ($owner_rc, $owner_out) = $run->([], read_owner => 1);
+ is($owner_rc . $owner_out, '0' . $digest->($content) . " -\n",
+ 'the owner is observed');
+ ($rc) = $run->([], claim => $U{8}, expected_owner => $digest->($content));
+ is($rc, 0, 'and can be recovered by its exact digest');
+ }
+
+ unlink("$runtime/owner") or die "unlink temporary owner: $!\n";
+ ($rc) = $run->([], claim => $U{3}, expected_owner => $digest->("$U{8}\n"));
+ is($rc, 73, 'a stale digest cannot claim after its owner disappears');
+ ok(!-e "$runtime/owner", 'the stale digest does not recreate the missing owner');
+ ($rc) = $run->([write_step($output, 'missing')], token => $U{2});
+ is($rc, 73, 'a missing marker fails closed');
+
+ # An unusable directory is refused with its own exit code before the
+ # flock, by every guarded call.
+ chmod(0755, $runtime) or die "chmod temporary runtime: $!\n";
+ ($rc) = $run->([], claim => $U{3}, expected_owner => 'missing');
+ is($rc, 74, 'an unsafe runtime mode refuses a claim');
+ ($owner_rc, $owner_out) = $run->([], read_owner => 1);
+ is($owner_rc . $owner_out, '74', 'and an owner observation, without a snapshot');
+ ($rc) = $run->([write_step($output, 'unsafe')], token => $U{2});
+ is($rc, 74, 'and a tokenized call, before its token is checked');
+ chmod(0700, $runtime) or die "chmod temporary runtime: $!\n";
+ rename("$runtime/lock", "$runtime/lock.held") or die "rename temporary lock: $!\n";
+ mkdir("$runtime/lock") or die "mkdir temporary lock: $!\n";
+ ($rc) = $run->([], claim => $U{3}, expected_owner => 'missing');
+ is($rc, 74, 'a lock entry that is not a regular file is refused');
+ rmdir("$runtime/lock") or die "rmdir temporary lock: $!\n";
+ symlink("$runtime/lock.held", "$runtime/lock") or die "symlink temporary lock: $!\n";
+ ($rc) = $run->([], claim => $U{3}, expected_owner => 'missing');
+ is($rc, 74, 'so is a symlink lock');
+ unlink("$runtime/lock") or die "unlink temporary lock: $!\n";
+ rename("$runtime/lock.held", "$runtime/lock") or die "rename temporary lock: $!\n";
+ mkdir("$runtime/owner") or die "mkdir temporary owner: $!\n";
+ ($rc) = $run->([], claim => $U{3}, expected_owner => 'missing');
+ is($rc, 74, 'so is an owner entry that is not a regular file');
+ rmdir("$runtime/owner") or die "rmdir temporary owner: $!\n";
+ rename($runtime, "$dir/actual") or die "rename temporary runtime: $!\n";
+ symlink("$dir/actual", $runtime) or die "symlink temporary runtime: $!\n";
+ ($rc) = $run->([], claim => $U{3}, expected_owner => 'missing');
+ is($rc, 74, 'and a symlink runtime');
+ is(slurp($output), "B\n", 'all refusal paths preserve B output');
+};
+
# Last: every command rendered during this test run, and every protocol
# violation recorded (by a fake target: a key on a command line or in output,
-# a change without the lock; or a local command).
+# a change without the lock, a transaction call without a token; or a local
+# command).
subtest 'every rendered command is a POSIX command chain' => sub {
is_deeply(\@VIOLATIONS, [], 'no flow violated the target protocol');
my %shapes;
diff --git a/src/test/zfsnvme_test.pm b/src/test/zfsnvme_test.pm
index 3b311bfa..8a963bfe 100644
--- a/src/test/zfsnvme_test.pm
+++ b/src/test/zfsnvme_test.pm
@@ -1053,6 +1053,49 @@ my @ownership = (
eval { $run->(scfg(), [['cat', '/proc/mounts']], op => 'test') };
is($@, "received interrupt\n", 'a stopped task dies with the task marker and nothing else');
}
+ # The exit codes of the transport guard: 74 (an unusable lock directory)
+ # names the directory problem for every guarded call, 73 (a failed
+ # compare-and-set) and 200 (busy) are explained for a claim only, and a
+ # call without the guard keeps them all as they are.
+ my $unsafe =
+ 'unsafe lock directory /run/pve-storage-nvmet on the target: it must be a root-owned'
+ . ' directory with mode 0700 whose lock and owner entries are regular files';
+ for my $option (qw(read_owner claim token)) {
+ my $steps = $option eq 'token' ? [['cat', '/proc/mounts']] : [];
+ my %call = (op => 'test', $option => ($option eq 'read_owner' ? 1 : $UUID));
+ $call{expected_owner} = 'missing' if $option eq 'claim';
+ for my $rc (0, 1, 73, 74, 200, 255) {
+ local $COMMAND = sub($cmd, %opts) {
+ die "command 'ssh' failed: exit code $rc\n" if $rc;
+ return 0;
+ };
+ my $err = $rc ? "exit code $rc" : '';
+ $err = $unsafe if $rc == 74;
+ $err = 'another transaction took over the target' if $option eq 'claim' && $rc == 73;
+ $err = 'target busy: could not acquire transaction lock'
+ if $option eq 'claim' && $rc == 200;
+ is_deeply(
+ $run->(scfg(), $steps, %call),
+ { rc => $rc, out => [], err => $err },
+ "$option preserves exit $rc, names the directory problem of 74, and only a"
+ . ' claim classifies 73 as taken over and 200 as busy',
+ );
+ }
+ local $COMMAND = sub($cmd, %opts) { die "command 'ssh' failed: got timeout\n" };
+ is_deeply(
+ $run->(scfg(), $steps, %call),
+ { rc => -1, out => [], err => 'timeout' },
+ "$option does not confuse an SSH timeout with target busy",
+ );
+ }
+ for my $rc (73, 74, 200) {
+ local $COMMAND = sub($cmd, %opts) { die "command 'ssh' failed: exit code $rc\n" };
+ is_deeply(
+ $run->(scfg(), [['cat', '/proc/mounts']], op => 'test'),
+ { rc => $rc, out => [], err => "exit code $rc" },
+ "a call without the guard keeps exit $rc as it is",
+ );
+ }
@runs = ();
eval { $run->(scfg(), [['cat', '/proc/mounts']]) };
like($@, qr/NVMe target call without label/, 'every call has a label');
^ permalink raw reply related [flat|nested] 7+ messages in thread
* [PATCH storage v3 4/4] zfsnvme: wait up to 30 seconds for the shared storage lock
2026-10-05 0:26 [PATCH storage v3 0/4] add ZFS over NVMe/TCP storage plugin Joaquin Varela
` (2 preceding siblings ...)
2026-10-05 0:26 ` [PATCH storage v3 3/4] zfsnvme: fence target commands of abandoned transactions Joaquin Varela
@ 2026-10-05 0:26 ` Joaquin Varela
2026-10-05 0:26 ` [PATCH docs v3] storage: document ZFS over NVMe/TCP Joaquin Varela
2026-10-05 0:26 ` [PATCH manager v3] ui: storage: add ZFS over NVMe/TCP editor Joaquin Varela
5 siblings, 0 replies; 7+ messages in thread
From: Joaquin Varela @ 2026-10-05 0:26 UTC (permalink / raw)
To: pve-devel
The core takes the shared storage lock without a timeout around
vdisk_alloc, vdisk_free, vdisk_clone, vdisk_create_base, rename_volume
and rename_snapshot, so cfs_lock applies its default of 10 seconds for
acquiring it. A zfsnvme allocation holds that lock for several SSH
round trips (read the target state, reserve the NSID, create the zvol,
wait for its device, export and verify the namespace), and keeps it
while it waits for the target lock, which a destroy or rollback on the
same target can hold for several seconds. When two guests get disks on
one storage at the same time, the second one can then fail with "got
lock request timeout" although nothing went wrong.
Override cluster_lock_storage() to wait up to 30 seconds for the shared
lock when the caller passes no timeout, the same acquisition budget as
for the target lock. An explicit timeout is passed on unchanged (for 0,
cfs_lock still applies its own default of 10 seconds), and so are the
execution timeout, error handling and exclusion of the parent method.
Signed-off-by: Joaquin Varela <joaquinvarela@neatech.ar>
---
src/PVE/Storage/ZFSNVMePlugin.pm | 9 ++++++++
src/test/zfsnvme_test.pm | 35 +++++++++++++++++++++++++++++++-
2 files changed, 43 insertions(+), 1 deletion(-)
diff --git a/src/PVE/Storage/ZFSNVMePlugin.pm b/src/PVE/Storage/ZFSNVMePlugin.pm
index dfd43d43..67595350 100644
--- a/src/PVE/Storage/ZFSNVMePlugin.pm
+++ b/src/PVE/Storage/ZFSNVMePlugin.pm
@@ -2452,6 +2452,15 @@ sub clone_image($class, $scfg, $storeid, $volname, $vmid, $snap = undef) {
return "$basename/$name";
}
+sub cluster_lock_storage($class, $storeid, $shared, $timeout, $func, @param) {
+ # Remote lifecycle operations take several SSH round trips while the core
+ # holds this lock. Allow a bounded queue of such operations, using the same
+ # acquisition budget as the target lock. Keep caller-supplied deadlines and
+ # the parent's execution timeout, error handling and exclusion unchanged.
+ $timeout = $nvmet_lock_wait if $shared && !defined($timeout);
+ return $class->SUPER::cluster_lock_storage($storeid, $shared, $timeout, $func, @param);
+}
+
sub alloc_image($class, $storeid, $scfg, $vmid, $fmt, $name, $size) {
die "unsupported format '$fmt'" if $fmt ne 'raw';
die "illegal name '$name' - should be 'vm-$vmid-*'\n"
diff --git a/src/test/zfsnvme_test.pm b/src/test/zfsnvme_test.pm
index 8a963bfe..7a218bf5 100644
--- a/src/test/zfsnvme_test.pm
+++ b/src/test/zfsnvme_test.pm
@@ -1108,6 +1108,39 @@ my @ownership = (
# The shared storage lock
# ---------------------------------------------------------------------------
+subtest 'cluster lock timeout defaults and passthrough' => sub {
+ my @base_calls;
+ $parent_mock->redefine(
+ cluster_lock_storage => sub($class, $storeid, $shared, $timeout, $func, @param) {
+ push @base_calls, [$class, $storeid, $shared, $timeout, @param];
+ return $func->(@param);
+ },
+ );
+ for my $case (
+ ['shared default', 1, undef, 30],
+ ['shared explicit', 1, 12, 12],
+ ['shared zero (the parent applies its own default)', 1, 0, 0],
+ ['local default', 0, undef, undef],
+ ) {
+ my ($label, $shared, $timeout, $expected) = $case->@*;
+ my $result = $PLUGIN->cluster_lock_storage(
+ 'nvmetest',
+ $shared,
+ $timeout,
+ sub(@args) { return \@args },
+ 'left',
+ 'right',
+ );
+ is_deeply(
+ $base_calls[-1],
+ [$PLUGIN, 'nvmetest', $shared, $expected, 'left', 'right'],
+ "$label: the parent gets its arguments",
+ );
+ is_deeply($result, ['left', 'right'], "$label: the callback runs with its parameters");
+ }
+ $parent_mock->unmock('cluster_lock_storage');
+};
+
subtest 'vdisk_alloc dispatches under the shared storage lock' => sub {
my $cluster_mock = Test::MockModule->new('PVE::Cluster');
my (@locks, @allocations);
@@ -1131,7 +1164,7 @@ subtest 'vdisk_alloc dispatches under the shared storage lock' => sub {
'nvmetest:vm-105-disk-0',
'core allocation returns the allocated volume ID',
);
- is_deeply(\@locks, [['nvmetest', undef]], 'cfs_lock_storage gets the timeout of the core');
+ is_deeply(\@locks, [['nvmetest', 30]], 'cfs_lock_storage waits the bounded default');
is_deeply(
\@allocations,
[['nvmetest', 105, 'raw', 'vm-105-disk-0', 131_072]],
^ permalink raw reply related [flat|nested] 7+ messages in thread
* [PATCH docs v3] storage: document ZFS over NVMe/TCP
2026-10-05 0:26 [PATCH storage v3 0/4] add ZFS over NVMe/TCP storage plugin Joaquin Varela
` (3 preceding siblings ...)
2026-10-05 0:26 ` [PATCH storage v3 4/4] zfsnvme: wait up to 30 seconds for the shared storage lock Joaquin Varela
@ 2026-10-05 0:26 ` Joaquin Varela
2026-10-05 0:26 ` [PATCH manager v3] ui: storage: add ZFS over NVMe/TCP editor Joaquin Varela
5 siblings, 0 replies; 7+ messages in thread
From: Joaquin Varela @ 2026-10-05 0:26 UTC (permalink / raw)
To: pve-devel
Document the zfsnvme storage backend: target and node requirements,
its properties with a configuration example, native multipath and
DH-HMAC-CHAP authentication, security and availability limits, and the
lifecycle of the target configuration. Include the chapter in pvesm and
link its wiki page.
Every storage needs its own NVMe/TCP listeners on the target, and they
must not publish any other nvmet subsystem: nvmet rejects a connection
to a subsystem that is not published on the listener yet with DNR, and
the node deletes the controller regardless of the controller loss
timeout. The chapter says which portals and server values the backend
refuses for that reason, and that adding, updating or activating a
storage fails on them.
Nodes connect each path with a single write to /dev/nvme-fabrics and
delete controllers through sysfs, so nvme-cli is only needed for the
host identity files and as the administration tool. The chapter
states which connect options the running kernel must support, that
the key is never put on a command line or in a temporary file, and
which secret attributes are restricted to mode 0600 on the nodes and
on the target. Revoking a node also requires removing it from the
cluster and replacing the SSH key of every target, since every cluster
member can read the storage keys and the SSH keys give root access to
the targets.
It also explains what happens when every path stays down longer than
the controller loss timeout: the kernel removes the controllers and
the multipath block devices, and running guests have to be stopped and
started once the paths are back. A controller loss timeout of -1
needs a finite fast I/O fail timeout to keep guest I/O from blocking
for the whole outage.
Signed-off-by: Joaquin Varela <joaquinvarela@neatech.ar>
---
v3, accompanying "[PATCH storage v3 0/4] add ZFS over NVMe/TCP
storage plugin":
- the five v2 patches are squashed into one; the structure of the
chapter is unchanged
- changes that follow the storage v3 implementation:
- target requirements: the plugin only runs standard utilities over
SSH (listed), and the highest namespace ID handed out is kept in
the proxmox:nvme-last-nsid property of the pool
- nodes: paths are created through /dev/nvme-fabrics and deleted
through sysfs, so nvme-cli is only needed for the host identity
files (hostid must be a nonzero UUID) and as the administration
tool; the running kernel must support every connect option used,
and a failed connect is logged per portal
- every storage needs its own listeners (address and port), which
must not publish any other nvmet subsystem, because nvmet rejects
a subsystem that is not published on a listener yet with DNR and
the node then deletes the controller regardless of the controller
loss timeout; the plugin refuses a listener used by another
zfsnvme storage and a target address reached through a different
server value, and the server value is the key of the cluster-wide
lock per target
- target changes, including the restore after a target restart, take
that lock and therefore need quorum; an uncertain command result
aborts the operation and the next one reads the target again;
queries outside a transaction, such as capacity, do not take it
- one cluster per target, pools of one target must not overlap,
storages whose host NQNs overlap share the key
- nvme-host-ifaces and nvme-host-nqns are required in the schema,
an interface change is make-before-break, host NQNs can only be
added, so revoking a node has its own procedure
- sparse is no longer set by default, like in the web interface
- pvesm takes --dhchap-key as the path of a key file
- how the key reaches the kernel and the target, and the mode 0600
restriction of the secret attributes on the nodes and the target
- the connection tunables are applied to connected controllers on
the next activation
- no pvesm export/import
- removing the storage never deletes volumes; what deactivation and
the best-effort target cleanup do
- the optional transaction guard of storage 3/4: one sentence in the
target requirements and one paragraph in "Security and
Availability"; both go away if that patch is dropped
- corrections of v2 statements that were wrong or misleading:
- DH-HMAC-CHAP only authenticates the hosts to the target, with one
key shared by all nodes (v2: "authenticates the endpoints")
- the target kernel needs NVMe in-band authentication, which v2
already required but did not list
- the plugin itself loads nvmet_tcp and mounts configfs on the
target; loading the transport at boot is recommended, not required
- all-path loss: after the controller loss timeout the kernel
deletes the controllers and the multipath devices, and running
guests need a stop and start; a rejection deletes the controller
at once; -1 needs a finite fast I/O fail timeout. v2 only said
that I/O stays queued, and "600 seconds in the example" was wrong
because the example sets a fast I/O fail timeout of 30 seconds
- the configuration example uses documentation addresses and an
example NQN, and drops "shared 1", which is not a property of
this type
- the pvesm wiki link uses the page name of the chapter title
- nothing else was added, apart from wording ("controller loss
timeout" without hyphens) and the anchor for the cross-references to
"Security and Availability"; happy to trim further
v2: https://lore.proxmox.com/pve-devel/cover.1785636981.git.joaquinvarela@neatech.ar/
pve-storage-zfsnvme.adoc | 373 +++++++++++++++++++++++++++++++++++++++
pvesm.adoc | 4 +
2 files changed, 377 insertions(+)
create mode 100644 pve-storage-zfsnvme.adoc
diff --git a/pve-storage-zfsnvme.adoc b/pve-storage-zfsnvme.adoc
new file mode 100644
index 0000000..c590636
--- /dev/null
+++ b/pve-storage-zfsnvme.adoc
@@ -0,0 +1,373 @@
+[[storage_zfsnvme]]
+ZFS over NVMe/TCP Backend
+-------------------------
+ifdef::wiki[]
+:pve-toplevel:
+:title: Storage: ZFS over NVMe/TCP
+endif::wiki[]
+
+Storage pool type: `zfsnvme`
+
+This backend accesses a remote Linux machine with a ZFS pool and the kernel
+NVMe target through `ssh`. For each guest disk it creates a ZVOL, exports it as
+an NVMe namespace, and connects the {pve} nodes through native Linux NVMe/TCP
+multipath.
+
+The backend supports thin provisioning, snapshots, rollback, templates, linked
+clones, offline resize, and shared-storage live migration. Namespace UUIDs are
+stored as ZFS user properties, and the stable `nvme-uuid` device link is used
+for guest disks.
+
+Configuration
+~~~~~~~~~~~~~
+
+The target needs OpenZFS, configfs, and the `nvmet` and `nvmet-tcp` kernel
+modules with NVMe in-band authentication support. The backend runs target
+utilities over SSH and does all parsing and planning on the {pve} node, so
+neither Perl nor Bash is required on the target. Root's login shell must be a
+POSIX-compatible shell, and the target needs these standard utilities: `find`
+(with `-prune`, `-path`, and `-exec ... {} +`), `grep`, `env`, `printf`, `test`,
+`cat`, `mkdir`, `rmdir`, `rm`, `ln`, `chmod`, `dd`, `tee`, `sha256sum`,
+`modprobe`, and `mount`. The transaction guard described in
+xref:storage_zfsnvme_security[Security and Availability] also needs `/bin/sh` to
+be a POSIX-compatible shell, util-linux `flock` with `-w` and `-E` support, and
+`stat` with `-c '%u:%a'` support. Configure root SSH access like for the ZFS
+over iSCSI backend. The key for the server address is stored at
+`/etc/pve/priv/zfs/<server>_id_rsa`.
+
+The NVMe target modules must be installed on the target. The backend loads
+`nvmet_tcp` and, if necessary, mounts configfs before it restores the target
+configuration after a reboot, but loading the transport at boot is recommended.
+On a Linux target using systemd, load the transport at boot and verify the
+configfs mount:
+
+----
+# echo nvmet_tcp >/etc/modules-load.d/pve-nvmet.conf
+# modprobe nvmet_tcp
+# mountpoint /sys/kernel/config
+----
+
+The target configuration below configfs is derived state. It does not need a
+separate persistence service: after the module and ZFS pool are available, the
+backend reconstructs namespaces and UUIDs from ZFS user properties, and
+reconstructs the subsystem, host ACLs, and ports from the shared storage
+configuration during activation. The backend also records the highest
+namespace ID it handed out in the `proxmox:nvme-last-nsid` property of the
+pool dataset, so that a namespace ID is not reused for a new volume.
+
+Do not leave cloud-init device discovery enabled on a dedicated Linux target.
+A guest cloud-init ZVOL contains a `cidata` filesystem and can otherwise be
+mistaken for the target host's own NoCloud datasource during boot. After the
+target host is provisioned, remove cloud-init or disable it according to the
+distribution's documentation. For distributions supporting the standard
+disable marker, use:
+
+----
+# touch /etc/cloud/cloud-init.disabled
+----
+
+Each {pve} node needs the `nvme-tcp` module, native NVMe multipath, a unique
+`/etc/nvme/hostnqn`, and a nonzero UUID in `/etc/nvme/hostid`. The `nvme-cli`
+package creates both files and provides the NVMe administration tools.
+Interface names listed in the storage configuration must exist on every node
+where the storage is enabled.
+
+The backend does not use `nvme-cli` to connect or disconnect. It creates each
+missing path with a single write to the kernel's `/dev/nvme-fabrics` interface
+from a short-lived child process, supplying both host identities explicitly,
+and deletes controllers through sysfs. The running kernel must support every
+connect option the backend uses, including DH-HMAC-CHAP and `host_iface`;
+otherwise the connection fails with an error naming the option. A failed
+connection is logged as a warning for its portal, with the error text reported
+by the kernel.
+
+The following properties are specific to ZFS over NVMe/TCP:
+
+server::
+
+IP address or DNS name used for the SSH control connection. Connected guest I/O
+uses the independent NVMe/TCP data paths, but capacity reporting and lifecycle
+operations require this endpoint. Use a redundant management DNS name or VIP
+where the storage appliance provides one. Use the same value for a target in
+all storage definitions. A cluster-wide lock serializes transactions using this
+value. The backend can only detect another spelling of the same target, such as
+its IP address and its DNS name, when the storage definitions share a target
+address in `nvme-portals`.
+
+Use each target with only one {pve} cluster. Storage definitions on the same
+target need their own pool, NQN, and listener, see `nvme-portals`. They share
+the cluster-wide lock, which does not coordinate with other clusters.
+
+pool::
+
+ZFS pool or child dataset used exclusively by this storage definition. It must
+not contain, or be contained in, the pool of another storage definition on the
+same target. Do not share an NQN or pool with another cluster.
+
+subsysnqn::
+
+NVMe qualified name of the target subsystem. It must be unique across the
+cluster's storage definitions.
+
+nvme-portals::
+
+Comma-separated NVMe/TCP target addresses. The default service is `4420`.
+Specify IPv6 addresses in brackets, for example `[2001:db8::10]:4420`.
++
+Each portal is a listener on the target, that is, an address and port. Every
+storage definition needs its own listeners, and they must not publish any other
+nvmet subsystem, including target configuration not made by {pve}. nvmet
+rejects a connection to a subsystem that is not published on that listener yet
+with the Do Not Retry (DNR) status, for example while the target configuration
+is restored after a target restart, and the node then deletes the controller
+regardless of the controller loss timeout. The backend compares the portals
+with those of every other ZFS over NVMe/TCP storage definition, including
+disabled ones. Adding a storage, any update of its configuration, and its
+activation on a node fail if one of its portals uses an address and port that
+another storage definition already uses, whatever their `server` values, or an
+address, on any port, that another storage definition reaches through a
+different `server` value. IPv6 addresses are compared in canonical form, and an
+IPv4-mapped IPv6 address counts as the IPv4 address. Use different addresses or
+ports, and the same `server` value, for storage definitions on the same target.
+
+nvme-host-ifaces::
+
+Required. Comma-separated local interfaces, matched to `nvme-portals` by
+position. An explicit interface prevents a failed data path from reconnecting
+over the management network. Activation fails before changing target state if
+any configured interface is missing on the local node. When an interface
+changes, the backend connects the path on the new interface before it
+disconnects the old one.
+
+nvme-host-nqns::
+
+Required. Comma-separated contents of `/etc/nvme/hostnqn` from every cluster
+node that may activate the storage. The complete allow-list lets any one node
+restore all host ACLs before publishing the subsystem after a target reboot.
+Activation fails if the local node's Host NQN is absent. Update this property
+before enabling the storage on a newly added cluster node. Host NQNs can only
+be added. To revoke a host, follow the procedure in
+xref:storage_zfsnvme_security[Security and Availability], which requires
+removing the node from the cluster, a new SSH key for the target, and a new key.
+
+dhchap-key::
+
+NVMe DH-HMAC-CHAP secret in `DHHC-1` representation. The value is handled as a
+sensitive property and stored below `/etc/pve/priv/storage/` with mode `0600`.
+The target stores the key per Host NQN, so storage definitions on the same
+target whose `nvme-host-nqns` overlap must use the same key. Key rotation is
+not supported, see
+xref:storage_zfsnvme_security[Security and Availability].
+
+nvme-iopolicy::
+
+Native multipath policy: `round-robin`, `queue-depth`, or `numa`.
+
+nvme-keep-alive-tmo::
+
+Keep-alive timeout in seconds.
+
+nvme-reconnect-delay::
+
+Delay between controller reconnect attempts in seconds.
+
+nvme-ctrl-loss-tmo::
+
+Time in seconds during which the kernel keeps reconnecting a lost controller
+before it deletes the controller. `-1` retries indefinitely. See
+xref:storage_zfsnvme_security[Security and Availability] for the consequences.
+
+nvme-fast-io-fail-tmo::
+
+Optional time in seconds before outstanding I/O fails while a controller is
+reconnecting. If this property is absent, the kernel queues I/O until
+`nvme-ctrl-loss-tmo` expires. A configured value must not exceed a finite
+controller loss timeout. A value of `0` selects immediate fail-fast behavior.
+
+nvme-nr-io-queues::
+
+Optional number of I/O queues per controller.
+
+blocksize::
+
+ZFS volume block size.
+
+sparse::
+
+Use ZFS thin provisioning instead of reserving the full virtual size. If this
+property is not set, volumes are thick-provisioned, which is also the default
+in the web interface.
+
+.Configuration Example (`/etc/pve/storage.cfg`)
+----
+zfsnvme: nvme-shared
+ server 203.0.113.10
+ pool tank/pve-nvme
+ subsysnqn nqn.2026-01.com.example:pve-nvme
+ nvme-portals 192.0.2.10:4420,198.51.100.10:4420
+ nvme-host-ifaces ens1f0,ens1f1
+ nvme-host-nqns nqn.2014-08.org.nvmexpress:uuid:11111111-1111-1111-1111-111111111111,nqn.2014-08.org.nvmexpress:uuid:22222222-2222-2222-2222-222222222222
+ blocksize 16k
+ sparse 1
+ nvme-iopolicy round-robin
+ nvme-keep-alive-tmo 5
+ nvme-reconnect-delay 2
+ nvme-ctrl-loss-tmo 600
+ nvme-fast-io-fail-tmo 30
+ content images
+----
+
+The DH-HMAC-CHAP key is intentionally not shown in `storage.cfg`. Set it with
+the storage creation API or web interface. With `pvesm`, pass the path of a file
+that contains the key, for example `--dhchap-key /root/nvme-dhchap.key`. Create
+the file with mode `0600`, for example after `umask 077`, and remove it once the
+storage is created.
+
+[[storage_zfsnvme_security]]
+Security and Availability
+~~~~~~~~~~~~~~~~~~~~~~~~~
+
+The target uses an ACL for the unique Host NQN of every node and never enables
+`allow_any_host`. All nodes share the DH-HMAC-CHAP key of the storage, and it
+only authenticates the hosts to the target. The target is not authenticated to
+the hosts, no Diffie-Hellman group is configured, and data is not encrypted.
+Host NQNs are not secret, so any holder of the key can connect as any allowed
+host. This backend does not configure NVMe/TCP TLS, so use isolated storage
+networks or equivalent protection.
+
+Apart from its key file, the backend keeps the key in memory. It passes the key
+to the local kernel in the connect write to `/dev/nvme-fabrics`, and to the
+target on the standard input of an SSH command. The key never appears on a
+command line, and the backend writes no temporary files.
+
+The kernel exposes the key with mode `0644` in the `dhchap_secret` sysfs
+attributes of the NVMe controllers on the nodes and in the host entries below
+`/sys/kernel/config/nvmet/hosts/` on the target. The backend attempts to
+restrict the controller attributes on the {pve} nodes to mode `0600` right after
+it creates a controller, during storage activation, and after each connection
+attempt. This is best effort: permission changes can fail, and controller
+attributes can appear or reappear after the permission check. Until then, any
+process on the node that can read sysfs, including processes in containers, can
+read the key. There is no guaranteed upper bound on the exposure interval. Check
+the actual permissions and activation warnings; do not rely on this mitigation
+to protect keys from untrusted local users. On the target, the backend restricts
+the `dhchap_key` and `dhchap_ctrl_key` attributes of a host entry to mode `0600`
+before it writes the key. A process on the target that opened one of these
+attributes before the change can still read the key through that descriptor.
+Host entries that already hold the key, for example ones created by hand, keep
+their mode. Limit shell access on both the target and the {pve} nodes to trusted
+administrators.
+
+To rotate the key, or to revoke a node:
+
+. To revoke a node, first remove it from the cluster. Every cluster member can
+ read the keys below `/etc/pve/priv/`: the DH-HMAC-CHAP key of every ZFS over
+ NVMe/TCP storage and the SSH key of every target, which gives root access to
+ the target, including its keys and volumes. Then replace the SSH key of each
+ target: create a new key pair at `/etc/pve/priv/zfs/<server>_id_rsa` (for
+ every `server` value used for the target), replace the old public key with
+ the new one in `/root/.ssh/authorized_keys` on the target, and check that
+ file for other unknown keys.
+. Move or remove every volume of the storage, including unused disks,
+ templates, VM state volumes, and disks with snapshots, which cannot be moved
+ while their snapshots exist. `pvesm list <storage>` must show no volume.
+. Remove this storage definition. The backend removes the subsystem from the
+ target once it owns no volumes. If the removal task warns that it kept the
+ NVMe target configuration, remove the remaining volumes, or the subsystem on
+ the target, before you continue.
+. On every other node that had the storage active, disconnect it with
+ `nvme disconnect --nqn <subsysnqn>` or reboot the node.
+. Create a new storage with a new key and a new `subsysnqn`, without the Host
+ NQN of a revoked node.
+
+To revoke a node, repeat this for every ZFS over NVMe/TCP storage of the
+cluster, since the node could read all of their keys. Other storage definitions
+on the same target that list any of the same Host NQNs share the key, because
+the target stores it per Host NQN. Rotate their keys together: remove all of
+them before you create the new ones.
+
+Shared access relies on {pve} cluster locking and fencing. Test quorum and
+fencing before placing production guests on the storage. Use independent
+failure domains for the configured paths and make sure the ZFS target itself
+is not a single point of failure.
+
+Changes to the target configuration, including restoring it after a target
+restart, are serialized by a cluster-wide lock per target and therefore require
+quorum. An uncertain command result aborts the operation; a subsequent
+operation reads the target again. Queries outside a transaction, such as
+capacity and namespace identity queries, do not take this lock. They can see
+intermediate state, and the underlying ZFS queries can still be delayed by
+storage activity.
+
+In addition, the transaction guard of the backend serializes the commands of
+each transaction with a lock on the target, in the runtime directory
+`/run/pve-storage-nvmet`. Before claiming a transaction, the backend observes
+the current owner without taking the target lock. The claim compares that
+observation under the lock before publishing a new token, so an abandoned claim
+cannot replace an owner that has since changed. Subsequent command groups,
+including reads, acquire the same lock and validate the token. This prevents a
+command delayed by an SSH failure or an expired cluster lock from modifying a
+volume after a newer operation has replaced it. The lock and token are shared by
+all pools and storage definitions on the target host, so they also reject
+commands from transactions that used another spelling of the server address.
+Queries outside a transaction do not wait for the target lock. The runtime
+directory must be owned by root with mode `0700`. Do not delete it or replace
+its `lock` or `owner` files while commands can still be running. A stuck command
+can block further changes; the backend does not break this lock to restore
+availability.
+
+Choose the all-path outage policy for the workload. The default configuration
+leaves `nvme-fast-io-fail-tmo` unset: transient outages can recover without
+guest block errors, but an I/O request can stall until the controller is lost.
+The default controller loss timeout is 600 seconds; the example instead
+explicitly sets a fast I/O fail timeout of 30 seconds. Set a shorter fast I/O
+fail timeout when the application prefers a prompt block error over a long
+stall. This is a service-level decision, not a universally safer default.
+
+When every path of a node stays down past the controller loss timeout, the
+kernel deletes the controllers and the multipath block devices of the storage on
+that node, and queued and new I/O fails. A running guest keeps the removed
+device open, so once the paths are back, stop and start the guest, for example
+with `qm stop <vmid>` and `qm start <vmid>`; a reboot inside the guest is not
+enough. A target that rejects the connection, for example because of an unknown
+Host NQN, a wrong key, or a listener that does not publish the subsystem yet
+(see `nvme-portals`), makes the kernel delete the controller regardless of the
+timeout. A controller loss timeout of `-1` keeps the devices, but combine it
+with a finite fast I/O fail timeout. Otherwise guest I/O, and stopping or
+migrating the guest, can block for the whole outage.
+
+Changes to `nvme-iopolicy`, `nvme-reconnect-delay`, `nvme-ctrl-loss-tmo`, and
+`nvme-fast-io-fail-tmo` are applied to connected controllers on the next
+storage activation. They do not affect an outage of all paths that is already
+in progress. `nvme-keep-alive-tmo` and `nvme-nr-io-queues` only apply to new
+controller connections, for example after a node reboot.
+
+Online resize of a running guest disk is not supported. Stop the VM before
+resizing; the new size is detected when the block device is reopened.
+
+The backend does not support volume transfer streams through `pvesm export`
+or `pvesm import`, including `raw+size`. Offline or remote migration workflows
+that require these streams cannot transfer disks to or from this backend.
+Shared-storage live migration remains supported: it does not transfer the disk
+contents.
+
+Removing the storage definition never deletes volumes. The node that removes it
+deactivates the storage, which deletes its controllers. Deactivation fails if
+one of the namespaces is still in use locally, or if the controllers cannot be
+deleted within 15 seconds (a deletion that is stuck in the kernel keeps the
+task waiting until the kernel returns); the removal then logs a warning and the
+node keeps its connections. Other nodes keep their connections. On each of
+them, run `nvme disconnect --nqn <subsysnqn>` or reboot the node. If no volumes
+of the storage remain, the backend also tries to remove the subsystem and
+unused host entries from the target. This cleanup is best effort: if it fails,
+or if volumes remain, the target configuration is kept and a warning is logged.
+
+Storage Features
+~~~~~~~~~~~~~~~~
+
+.Storage features for backend `zfsnvme`
+[width="100%",cols="m,m,3*d",options="header"]
+|==============================================================================
+|Content types |Image formats |Shared |Snapshots |Clones
+|images |raw |yes |yes |yes
+|==============================================================================
diff --git a/pvesm.adoc b/pvesm.adoc
index 5bd24b2..4ada3f0 100644
--- a/pvesm.adoc
+++ b/pvesm.adoc
@@ -439,6 +439,8 @@ See Also
* link:/wiki/Storage:_ZFS_over_ISCSI[Storage: ZFS over ISCSI]
+* link:/wiki/Storage:_ZFS_over_NVMe/TCP[Storage: ZFS over NVMe/TCP]
+
endif::wiki[]
ifndef::wiki[]
@@ -471,6 +473,8 @@ include::pve-storage-btrfs.adoc[]
include::pve-storage-zfs.adoc[]
+include::pve-storage-zfsnvme.adoc[]
+
ifdef::manvolnum[]
include::pve-copyright.adoc[]
^ permalink raw reply related [flat|nested] 7+ messages in thread
* [PATCH manager v3] ui: storage: add ZFS over NVMe/TCP editor
2026-10-05 0:26 [PATCH storage v3 0/4] add ZFS over NVMe/TCP storage plugin Joaquin Varela
` (4 preceding siblings ...)
2026-10-05 0:26 ` [PATCH docs v3] storage: document ZFS over NVMe/TCP Joaquin Varela
@ 2026-10-05 0:26 ` Joaquin Varela
5 siblings, 0 replies; 7+ messages in thread
From: Joaquin Varela @ 2026-10-05 0:26 UTC (permalink / raw)
To: pve-devel
Add an input panel for the zfsnvme storage type, so that ZFS over
NVMe/TCP storages can be added and edited in the web interface, and
register it in the storage type list and the JavaScript bundle.
The general tab covers the SSH server, ZFS pool, subsystem NQN, block
size, thin provisioning, NVMe/TCP portals, host interfaces, allowed
host NQNs, the DH-HMAC-CHAP key and the multipath I/O policy. The
connection timeouts and the number of I/O queues are advanced options.
Ranges and defaults are those of the storage plugin schema.
Apply the backend's restrictions in the form, so that they show up
before submitting: properties that are fixed after creation are
read-only when editing, host NQNs can only be added, and a fast I/O
fail timeout must not exceed a finite controller loss timeout. The key
is only requested on creation. Since the backend does not support key
rotation, editing neither shows nor resends it.
The help button links to the storage_zfsnvme section of the admin
guide, so generating OnlineHelpInfo.js requires a pve-doc-generator
that includes it.
Signed-off-by: Joaquin Varela <joaquinvarela@neatech.ar>
---
v3, accompanying "[PATCH storage v3 0/4] add ZFS over NVMe/TCP
storage plugin":
- rebased onto current master; the three v2 patches are squashed
- behavior changes against v2:
- thin provisioning is unchecked by default, like the API default
(the storage v3 series no longer sets sparse on creation) and the
other ZFS editors
- Allowed Host NQNs: when editing, a validator refuses to drop a
host NQN that the storage already has, since the backend refuses
it too; an edit-only hint says that host NQNs can only be added
and that revoking a host needs a new storage with a new key
- DH-HMAC-CHAP key: on creation the field checks the DHHC-1 format
(DHHC-1:0[0-3]:<base64>:) and is required; when editing it shows
"Unchanged (rotation not supported)" and sends nothing (v2 showed
"Configured"); the label is "DH-HMAC-CHAP Key" instead of "DHCHAP
Key"
- Fast I/O Fail Timeout: a validator refuses a value above a finite
Controller Loss Timeout, revalidated when that timeout changes;
an empty field shows "Off"
- layout: I/O Queues moved to the first advanced column and shows
"Default" when empty; the fast I/O fail sentence moved from the
general hint to a new advanced hint, which also says that a
Controller Loss Timeout of -1 retries forever
- placeholders use example values only, and the portals placeholder
no longer shows the default port
- the options match the storage v3 series, whose schema did not change
for the editor (nvme-host-ifaces and nvme-host-nqns are required in
the schema now, which the editor already enforced); its online help
needs the docs patch
v2: https://lore.proxmox.com/pve-devel/cover.1785636980.git.joaquinvarela@neatech.ar/
www/manager6/Makefile | 1 +
www/manager6/Utils.js | 6 +
www/manager6/storage/ZFSNVMeEdit.js | 233 ++++++++++++++++++++++++++++
3 files changed, 240 insertions(+)
create mode 100644 www/manager6/storage/ZFSNVMeEdit.js
diff --git a/www/manager6/Makefile b/www/manager6/Makefile
index d2ea786b..fa634804 100644
--- a/www/manager6/Makefile
+++ b/www/manager6/Makefile
@@ -376,6 +376,7 @@ JSSRC= \
storage/Summary.js \
storage/TemplateView.js \
storage/ZFSEdit.js \
+ storage/ZFSNVMeEdit.js \
storage/ZFSPoolEdit.js \
storage/ESXIEdit.js \
Workspace.js \
diff --git a/www/manager6/Utils.js b/www/manager6/Utils.js
index 8b99371d..7022b8df 100644
--- a/www/manager6/Utils.js
+++ b/www/manager6/Utils.js
@@ -880,6 +880,12 @@ Ext.define('PVE.Utils', {
faIcon: 'building',
backups: false,
},
+ zfsnvme: {
+ name: 'ZFS over NVMe/TCP',
+ ipanel: 'ZFSNVMeInputPanel',
+ faIcon: 'building',
+ backups: false,
+ },
zfspool: {
name: 'ZFS',
ipanel: 'ZFSPoolInputPanel',
diff --git a/www/manager6/storage/ZFSNVMeEdit.js b/www/manager6/storage/ZFSNVMeEdit.js
new file mode 100644
index 00000000..c7649cfe
--- /dev/null
+++ b/www/manager6/storage/ZFSNVMeEdit.js
@@ -0,0 +1,233 @@
+Ext.define('PVE.storage.ZFSNVMeInputPanel', {
+ extend: 'PVE.panel.StorageBase',
+
+ onlineHelp: 'storage_zfsnvme',
+
+ onGetValues: function (values) {
+ if (this.isCreate) {
+ values.content = 'images';
+ }
+ return this.callParent([values]);
+ },
+
+ initComponent: function () {
+ let me = this;
+
+ let splitList = (value) =>
+ (value || '')
+ .split(',')
+ .map((item) => item.trim())
+ .filter((item) => item !== '');
+
+ me.column1 = [
+ {
+ xtype: me.isCreate ? 'textfield' : 'displayfield',
+ name: 'server',
+ fieldLabel: gettext('SSH Server'),
+ allowBlank: false,
+ },
+ {
+ xtype: me.isCreate ? 'textfield' : 'displayfield',
+ name: 'pool',
+ fieldLabel: gettext('ZFS Pool'),
+ emptyText: 'tank/pve-nvme',
+ allowBlank: false,
+ },
+ {
+ xtype: me.isCreate ? 'textfield' : 'displayfield',
+ name: 'subsysnqn',
+ fieldLabel: gettext('Subsystem NQN'),
+ emptyText: 'nqn.2026-01.com.example:pve-nvme',
+ allowBlank: false,
+ },
+ {
+ xtype: me.isCreate ? 'textfield' : 'displayfield',
+ name: 'blocksize',
+ value: '16k',
+ fieldLabel: gettext('Block Size'),
+ validator: PVE.Utils.validateZfsBlocksize,
+ allowBlank: false,
+ },
+ {
+ xtype: 'proxmoxcheckbox',
+ name: 'sparse',
+ checked: false,
+ uncheckedValue: 0,
+ fieldLabel: gettext('Thin provision'),
+ },
+ ];
+
+ me.column2 = [
+ {
+ xtype: me.isCreate ? 'textfield' : 'displayfield',
+ name: 'nvme-portals',
+ fieldLabel: gettext('NVMe/TCP Portals'),
+ emptyText: '192.0.2.10,198.51.100.10',
+ allowBlank: false,
+ },
+ {
+ xtype: 'textfield',
+ name: 'nvme-host-ifaces',
+ fieldLabel: gettext('Host Interfaces'),
+ emptyText: 'ens1f0,ens1f1',
+ allowBlank: false,
+ },
+ {
+ xtype: 'textfield',
+ name: 'nvme-host-nqns',
+ fieldLabel: gettext('Allowed Host NQNs'),
+ emptyText: 'nqn.2014-08.org.nvmexpress:uuid:...',
+ allowBlank: false,
+ validator: function (value) {
+ // the backend refuses to drop a host NQN from an existing storage
+ let current = splitList(value);
+ let removed = splitList(this.originalValue).filter(
+ (hostnqn) => !current.includes(hostnqn),
+ );
+ if (removed.length) {
+ return Ext.String.format(
+ gettext('Host NQN {0} cannot be removed from an existing storage'),
+ Ext.htmlEncode(removed[0]),
+ );
+ }
+ return true;
+ },
+ },
+ me.isCreate
+ ? {
+ xtype: 'textfield',
+ inputType: 'password',
+ name: 'dhchap-key',
+ fieldLabel: gettext('DH-HMAC-CHAP Key'),
+ emptyText: 'DHHC-1:xx:...:',
+ regex: /^DHHC-1:0[0-3]:[A-Za-z0-9+/]+={0,2}:$/,
+ regexText: gettext('Expected format: DHHC-1:xx:...:'),
+ allowBlank: false,
+ }
+ : {
+ xtype: 'displayfield',
+ fieldLabel: gettext('DH-HMAC-CHAP Key'),
+ value: gettext('Unchanged (rotation not supported)'),
+ },
+ {
+ xtype: 'proxmoxKVComboBox',
+ name: 'nvme-iopolicy',
+ value: 'round-robin',
+ fieldLabel: gettext('I/O Policy'),
+ comboItems: [
+ ['round-robin', 'round-robin'],
+ ['queue-depth', 'queue-depth'],
+ ['numa', 'numa'],
+ ],
+ allowBlank: false,
+ },
+ ];
+
+ me.advancedColumn1 = [
+ {
+ xtype: 'proxmoxintegerfield',
+ name: 'nvme-keep-alive-tmo',
+ value: 5,
+ minValue: 1,
+ maxValue: 120,
+ fieldLabel: gettext('Keep Alive Timeout'),
+ allowBlank: false,
+ },
+ {
+ xtype: 'proxmoxintegerfield',
+ name: 'nvme-reconnect-delay',
+ value: 2,
+ minValue: 1,
+ maxValue: 120,
+ fieldLabel: gettext('Reconnect Delay'),
+ allowBlank: false,
+ },
+ {
+ xtype: 'proxmoxintegerfield',
+ name: 'nvme-nr-io-queues',
+ minValue: 1,
+ maxValue: 1024,
+ fieldLabel: gettext('I/O Queues'),
+ emptyText: Proxmox.Utils.defaultText,
+ deleteEmpty: !me.isCreate,
+ allowBlank: true,
+ },
+ ];
+
+ me.advancedColumn2 = [
+ {
+ xtype: 'proxmoxintegerfield',
+ name: 'nvme-ctrl-loss-tmo',
+ value: 600,
+ minValue: -1,
+ maxValue: 86400,
+ fieldLabel: gettext('Controller Loss Timeout'),
+ allowBlank: false,
+ listeners: {
+ change: function (field) {
+ let panel = field.up('inputpanel');
+ if (panel) {
+ panel.down('field[name=nvme-fast-io-fail-tmo]').validate();
+ }
+ },
+ },
+ },
+ {
+ xtype: 'proxmoxintegerfield',
+ name: 'nvme-fast-io-fail-tmo',
+ minValue: 0,
+ maxValue: 86400,
+ fieldLabel: gettext('Fast I/O Fail Timeout'),
+ emptyText: gettext('Off'),
+ deleteEmpty: !me.isCreate,
+ allowBlank: true,
+ validator: function (value) {
+ let panel = this.up('inputpanel');
+ if (value === '' || !panel) {
+ return true;
+ }
+ let ctrlLossTmo = panel.down('field[name=nvme-ctrl-loss-tmo]').getValue();
+ if (ctrlLossTmo === null || ctrlLossTmo < 0) {
+ return true;
+ }
+ return (
+ Number(value) <= ctrlLossTmo ||
+ gettext('Must not exceed the Controller Loss Timeout')
+ );
+ },
+ },
+ ];
+
+ me.advancedColumnB = [
+ {
+ xtype: 'displayfield',
+ userCls: 'pmx-hint',
+ value: gettext(
+ 'A Controller Loss Timeout of -1 retries forever. Leave Fast I/O Fail Timeout empty to queue I/O until the controller is lost.',
+ ),
+ },
+ ];
+
+ me.columnB = [
+ {
+ xtype: 'displayfield',
+ userCls: 'pmx-hint',
+ value: gettext(
+ 'List /etc/nvme/hostnqn from every allowed cluster node. Host interface names are matched to portals by position and must exist on every selected node.',
+ ),
+ },
+ ];
+
+ if (!me.isCreate) {
+ me.columnB.push({
+ xtype: 'displayfield',
+ userCls: 'pmx-hint',
+ value: gettext(
+ 'Host NQNs can only be added. Revoking a host requires a new storage with a new key.',
+ ),
+ });
+ }
+
+ me.callParent();
+ },
+});
^ permalink raw reply related [flat|nested] 7+ messages in thread
end of thread, other threads:[~2026-10-06 8:54 UTC | newest]
Thread overview: 7+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-05 0:26 [PATCH storage v3 0/4] add ZFS over NVMe/TCP storage plugin Joaquin Varela
2026-10-05 0:26 ` [PATCH storage v3 1/4] zfsnvme: " Joaquin Varela
2026-10-05 0:26 ` [PATCH storage v3 2/4] test: add zfsnvme plugin tests Joaquin Varela
2026-10-05 0:26 ` [PATCH storage v3 3/4] zfsnvme: fence target commands of abandoned transactions Joaquin Varela
2026-10-05 0:26 ` [PATCH storage v3 4/4] zfsnvme: wait up to 30 seconds for the shared storage lock Joaquin Varela
2026-10-05 0:26 ` [PATCH docs v3] storage: document ZFS over NVMe/TCP Joaquin Varela
2026-10-05 0:26 ` [PATCH manager v3] ui: storage: add ZFS over NVMe/TCP editor Joaquin Varela
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox