From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from gate001.proxmox.com (gate001.proxmox.com [45.144.208.40]) by lore.proxmox.com (Postfix) with ESMTPS id 41CE61FF0AA for ; Tue, 06 Oct 2026 10:54:45 +0200 (CEST) Received: from gate001.proxmox.com (localhost.localdomain [127.0.0.1]) by gate001.proxmox.com (Proxmox) with ESMTP id C89D32168B; Tue, 06 Oct 2026 10:54:14 +0200 (CEST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=neatech-ar.20251104.gappssmtp.com; s=20251104; t=1791159977; x=1791764777; darn=lists.proxmox.com; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=rFWubJePtRxkqDmBgLuF8zzwR046KDYy9P+GLEw2VUc=; b=lrs1qwFJ3EOmMVci5t8DxyRO61r28k2gyz7ruDcfYjEAXWWgldPVunGnoKl34E1bdj cdzPEPhPtDkui5ILFe31Ehh+xiVg6/kIuwtgLq2FYkmNtRKNHcUP8f09NMNvahnKnAna xfYRMG0fUdFkTYoRC2V01I+nkENzubf2e5XZkIzEmKGQpWalHNZo3ysu1mzgf86NMs/D cW7AH8Gohshj3+NAqmwB69QMFGZj6q/n73pYYbYJp3ewPooJ3JXZe1sz5z7KVASbddF0 gVZZxnzA/GKiByIO94DNTJWAvduboLNGRdWX6S6MIj8GUr4LGr57FHSrrfM4RCndMghA v9yw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1791159977; x=1791764777; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=rFWubJePtRxkqDmBgLuF8zzwR046KDYy9P+GLEw2VUc=; b=F+AS+C23ykt798bpDVX00r7o98EScqQdC11MKCNWpHWTBk3uXnDINMS5OEJtp3HWcH dA1yGUN5MFzqrYZtnhmBglqqegw4ReQZmrNtYeh5LdwPaT0hITDXWKdfRdYOwM6MXfEu fa2NfbLbbweQ9luIcSrimqbsqbZ0ooarNGZvW4eKDRhS2SYUYdkZdhQ76Mgzzg523Icw JkdRegOLuAGH8nvOqPdtraevzxOBSrddtxtvBFs0ZnxdH4ilaDLSjt+OXL+YmUrk8e27 Zp8MSiXvrWvaULutP2MCnBN95hG3bfUzSe09B2Pug1pIKXgDf1vGvHtHQ0vA3fK0App7 XpAA== X-Gm-Message-State: AFq9FYLS3nbrzOw00mswmE3PRHC8tiSwubJKoUqZTv0i0zisZktNpVAu ilN16+smPdlgAXmeAjM3v8pRUsi7hSfAJ8ISsAkPS17qG5ACym9DdrSFXN6pwLQM2E5vRIxGMtx oscaT1Pk= X-Gm-Gg: AYBFou1uFWrByBdpYo7TJPlRbk2uPZPchFweUoDaZaWHdo55bEPpIw5riIiI2sVMENb XK1iz2L8WWLTCKAmTXd0X7HIH1ZG0KnkUQfiaT8SPkAn/E9+q37Ej5RDnY+9B1pwhKcEuRs4jh6 1YLFCNI04Xz57yVlHKX4+J5mzNSsJ4AhYaqRq/iz/e5Ah4BwhlfTj13KLhenv+DooOvjUcvfkQH TQ0WJy1p/zHJxnC8Do6F3QVOrd1UAg51k5SOn9rpTrBT/wmlfGT1URkIPLUkMcdk3QUN55XLCrv 7RkoFnyHkfDJqAyYm5CxPFFrxTnzIzySQ3WyslqCgzesQHljgTR1F4hMDligs4atYbLrVXRVBeC BjlCC6hW+hQPRWnDfwYvlwRcRVIO/2F2hj/28T2OOgfI3bnCGW2fVK4o5PbKnsPa+HD4HtTxrwh CiArVptbAdY0fBTpxG374IRcED5cfAwQppCKOyCN0aBmO0y6PWEJCDqmlOakoqhnwfLoUnn9qCX /os+ipAn9RO+ciPCJPOP1bmPekZH6+oBUB4x0R6X4hmAUJ3dDRpaETDtOWKZC/B22U3q7kC32aq X-Received: by 2002:a05:6122:1814:b0:5d4:41cb:8936 with SMTP id 71dfb90a1353d-5dab24f19eemr1777589e0c.4.1791159974462; Sun, 04 Oct 2026 17:26:14 -0700 (PDT) From: Joaquin Varela To: pve-devel@lists.proxmox.com Subject: [PATCH storage v3 1/4] zfsnvme: add ZFS over NVMe/TCP storage plugin Date: Sun, 4 Oct 2026 21:26:04 -0300 Message-ID: <20261005002609.571-2-joaquinvarela@neatech.ar> X-Mailer: git-send-email 2.54.0.windows.1 In-Reply-To: <20261005002609.571-1-joaquinvarela@neatech.ar> References: <20261005002609.571-1-joaquinvarela@neatech.ar> MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-SPAM-LEVEL: Spam detection results: 0 AWL -0.299 Adjusted score from AWL reputation of From: address DKIM_SIGNED 0.1 Message has a DKIM or DK signature, not necessarily valid DKIM_VALID -0.1 Message has at least one valid DKIM or DK signature DMARC_PASS -0.1 DMARC pass policy KAM_ASCII_DIVIDERS 0.8 Email that uses ascii formatting dividers and possible spam tricks POISEN_SPAM_PILL 0.1 Meta: its spam POISEN_SPAM_PILL_1 0.1 random spam to be learned in bayes POISEN_SPAM_PILL_3 0.1 random spam to be learned in bayes SPF_HELO_NONE 0.001 SPF: HELO does not publish an SPF Record SPF_PASS -0.001 SPF: sender matches SPF record X-MailFrom: joaquinvarela@neatech.ar X-Mailman-Rule-Hits: max-size X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; emergency; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; news-moderation; no-subject; digests; suspicious-header Message-ID-Hash: XVR6MOHIDUYY4PWQS4LRUM6DRQEMFM7N X-Message-ID-Hash: XVR6MOHIDUYY4PWQS4LRUM6DRQEMFM7N X-Mailman-Approved-At: Tue, 06 Oct 2026 10:54:04 +0200 X-Mailman-Version: 3.3.10 Precedence: list List-Id: Proxmox VE development discussion List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: Add the shared storage type 'zfsnvme': ZFS zvols on a remote Linux target, exported with the kernel NVMe target (nvmet, configured through configfs) over NVMe/TCP, and consumed by the cluster nodes with native NVMe multipath and DH-HMAC-CHAP authentication. The ZFS dataset is the source of truth for a namespace: the subsystem NQN, NSID and UUID of every volume are stored in the ZFS user properties 'proxmox:nvme-subsys', 'proxmox:nvme-nsid' and 'proxmox:nvme-uuid', and the pool keeps the highest NSID handed out in 'proxmox:nvme-last-nsid'. The nodes find a namespace by its UUID. Everything else on the target (nvmet_tcp module, configfs mount, subsystem, namespaces, ports, host ACLs) is derived state that an activation rebuilds after a target restart. All parsing, validation and planning happen locally. The target state is read with one 'zfs get' and one 'find' over configfs; other reads are 'zfs get' and 'zfs list', 'test -b' while a new zvol appears, and 'cat /proc/mounts' to decide whether configfs needs to be mounted. The target runs quoted simple commands over SSH, joined with '&&' where a step depends on the previous one: an export starts with 'test -b' on its zvol, and each publish link repeats two read-only checks right before it links the subsystem to a port. There are no scripts, loops or variables on the target. As with ZFS over iSCSI, the SSH key is /etc/pve/priv/zfs/_id_rsa. Target changes are serialized cluster-wide by a pmxcfs domain lock per target server, and an operation whose SSH connection broke is reported as unknown and repaired by the next activation rather than compensated from a read that could overtake it. Each storage needs its own NVMe/TCP listener (address family, address and port) on the target. nvmet answers a connect to a listener that does not publish the subsystem yet with DNR, and the host then deletes the controller whatever its loss timeout, so after a target restart the storage published second on a shared listener would lose its paths. Adding, updating and activating a storage therefore refuse a portal whose listener another zfsnvme storage uses, disabled ones included, and a target address that another storage reaches through a different 'server' value, since that value names the target lock. A data address thus identifies one target for all zfsnvme storages of the cluster. The DH-HMAC-CHAP key is a sensitive property, stored like the other storage secrets in /etc/pve/priv/storage/.nvme-dhchap. It reaches the target through stdin, and the plugin restricts the dhchap_key and dhchap_ctrl_key attributes there to 0600 whenever it writes the key. On the nodes, controllers are created with exactly one write to /dev/nvme-fabrics in a bounded child (run_fork_with_timeout), so the key never appears on a command line, and the secret attributes of a new controller are restricted before the child returns. Rescans and disconnects are sysfs writes; nvme-cli is not run. Keep-alive, reconnect delay, controller loss and optional fast I/O failure timeouts are configurable, with defaults that prefer waiting for the target over failing guest I/O (ctrl_loss_tmo 600, fast_io_fail_tmo unset). Outside the plugin: register the type in PVE::Storage and the Makefile, add it to @SHARED_STORAGE in PVE::Storage::Plugin, let 'pvesm add' and 'pvesm set' read --dhchap-key from a file, and depend on nvme-cli, which creates /etc/nvme/hostnqn and /etc/nvme/hostid and remains the administration tool, like the client packages of the other shared types. Signed-off-by: Joaquin Varela --- debian/control | 1 + src/PVE/CLI/pvesm.pm | 21 +- src/PVE/Storage.pm | 2 + src/PVE/Storage/Makefile | 1 + src/PVE/Storage/Plugin.pm | 2 +- src/PVE/Storage/ZFSNVMePlugin.pm | 3177 ++++++++++++++++++++++++++++++ 6 files changed, 3201 insertions(+), 3 deletions(-) create mode 100644 src/PVE/Storage/ZFSNVMePlugin.pm diff --git a/debian/control b/debian/control index 850cd57c..ab46e6b6 100644 --- a/debian/control +++ b/debian/control @@ -43,6 +43,7 @@ Depends: bzip2, lvm2, lzop, nfs-common, + nvme-cli, proxmox-backup-client (>= 2.1.10~), proxmox-backup-file-restore, pve-cluster (>= 5.0-32), diff --git a/src/PVE/CLI/pvesm.pm b/src/PVE/CLI/pvesm.pm index 1aaa9f3c..3c68ffb5 100755 --- a/src/PVE/CLI/pvesm.pm +++ b/src/PVE/CLI/pvesm.pm @@ -78,12 +78,29 @@ sub param_mapping { }, }; + my $dhchap_key_map = { + name => 'dhchap-key', + desc => 'a file containing the NVMe DH-HMAC-CHAP key', + func => sub { + my ($value) = @_; + # never repeat a key passed in place of the file name + die "dhchap-key expects the path of a file containing the key\n" + if $value =~ /^DHHC-1:/i || -d $value; + my ($key) = split(/\n/, PVE::Tools::file_get_contents($value), 2); + $key = PVE::Tools::trim($key // ''); + die "DH-HMAC-CHAP key file '$value' is empty\n" if $key eq ''; + return $key; + }, + }; + my $mapping = { 'cifsscan' => [$password_map], 'cifs' => [$password_map], 'pbs' => [$password_map], - 'create' => [$password_map, $enc_key_map, $master_key_map, $keyring_map], - 'update' => [$password_map, $enc_key_map, $master_key_map, $keyring_map], + 'create' => + [$password_map, $enc_key_map, $master_key_map, $keyring_map, $dhchap_key_map], + 'update' => + [$password_map, $enc_key_map, $master_key_map, $keyring_map, $dhchap_key_map], }; return $mapping->{$name}; } diff --git a/src/PVE/Storage.pm b/src/PVE/Storage.pm index fc3db812..2e5a8031 100755 --- a/src/PVE/Storage.pm +++ b/src/PVE/Storage.pm @@ -36,6 +36,7 @@ use PVE::Storage::CephFSPlugin; use PVE::Storage::ISCSIDirectPlugin; use PVE::Storage::ZFSPoolPlugin; use PVE::Storage::ZFSPlugin; +use PVE::Storage::ZFSNVMePlugin; use PVE::Storage::PBSPlugin; use PVE::Storage::BTRFSPlugin; use PVE::Storage::ESXiPlugin; @@ -61,6 +62,7 @@ PVE::Storage::CephFSPlugin->register(); PVE::Storage::ISCSIDirectPlugin->register(); PVE::Storage::ZFSPoolPlugin->register(); PVE::Storage::ZFSPlugin->register(); +PVE::Storage::ZFSNVMePlugin->register(); PVE::Storage::PBSPlugin->register(); PVE::Storage::BTRFSPlugin->register(); PVE::Storage::ESXiPlugin->register(); diff --git a/src/PVE/Storage/Makefile b/src/PVE/Storage/Makefile index a67dc25f..d1cbfe29 100644 --- a/src/PVE/Storage/Makefile +++ b/src/PVE/Storage/Makefile @@ -11,6 +11,7 @@ SOURCES= \ ISCSIDirectPlugin.pm \ ZFSPoolPlugin.pm \ ZFSPlugin.pm \ + ZFSNVMePlugin.pm \ PBSPlugin.pm \ BTRFSPlugin.pm \ LvmThinPlugin.pm \ diff --git a/src/PVE/Storage/Plugin.pm b/src/PVE/Storage/Plugin.pm index a9e17513..a5aedc2b 100644 --- a/src/PVE/Storage/Plugin.pm +++ b/src/PVE/Storage/Plugin.pm @@ -35,7 +35,7 @@ our @COMMON_TAR_FLAGS = qw( ); our @SHARED_STORAGE = ( - 'iscsi', 'nfs', 'cifs', 'rbd', 'cephfs', 'iscsidirect', 'zfs', 'drbd', 'pbs', + 'iscsi', 'nfs', 'cifs', 'rbd', 'cephfs', 'iscsidirect', 'zfs', 'zfsnvme', 'drbd', 'pbs', ); our $QCOW2_PREALLOCATION = { diff --git a/src/PVE/Storage/ZFSNVMePlugin.pm b/src/PVE/Storage/ZFSNVMePlugin.pm new file mode 100644 index 00000000..127680fb --- /dev/null +++ b/src/PVE/Storage/ZFSNVMePlugin.pm @@ -0,0 +1,3177 @@ +package PVE::Storage::ZFSNVMePlugin; + +use v5.36; + +use Compress::Zlib qw(crc32); +use Digest::SHA qw(sha256_hex); +use Errno qw(ENOENT); +use Fcntl qw(O_NOFOLLOW O_RDONLY O_RDWR S_ISCHR); +use File::Path qw(make_path); +use List::Util qw(max); +use MIME::Base64 qw(decode_base64); +use POSIX qw(SIG_BLOCK SIG_SETMASK SIGHUP SIGINT SIGQUIT SIGTERM sigprocmask); +use Socket qw(AF_INET6 inet_ntop inet_pton); +use Time::HiRes; + +use PVE::Cluster; +use PVE::File; +use PVE::JSONSchema; +use PVE::Network; +use PVE::RESTEnvironment qw(log_warn); +use PVE::RPCEnvironment; +use PVE::SysFSTools; +use PVE::Tools qw(file_read_firstline file_set_contents run_command trim); + +use base qw(PVE::Storage::Plugin); + +# ZFS zvols on a remote Linux target, exported through the kernel NVMe target +# (nvmet, configured through configfs) over NVMe/TCP and consumed with native +# NVMe multipath. +# +# The ZFS dataset is the source of truth for a namespace: its owner NQN, NSID +# and UUID are ZFS user properties. Everything else on the target is derived +# state that the plugin rebuilds after a target restart. +# +# Storage parsing, validation and planning happen locally. SSH carries quoted +# commands, sometimes joined with `&&`, to read or change ZFS/configfs state. +# A pmxcfs domain lock per target serializes the changes within the cluster, +# held like the storage lock of the other shared storage types. + +# --------------------------------------------------------------------------- +# Constants and regular expressions +# --------------------------------------------------------------------------- + +# LogLevel=ERROR keeps a login banner out of the first stderr line, which is +# the error text of a failed call. +my @ssh_cmd = ( + '/usr/bin/ssh', '-o', 'BatchMode=yes', '-o', 'ConnectTimeout=10', '-o', 'LogLevel=ERROR', +); +my $id_rsa_path = '/etc/pve/priv/zfs'; +my $secret_dir = '/etc/pve/priv/storage'; +my $max_paths = 16; +my $max_hosts = 64; +my $slow_path_backoff = 60; + +my $nvmet_root = '/sys/kernel/config/nvmet'; +my $nvmet_model = 'Proxmox ZFS NVMe'; +my $nvmet_marker = 'ZFSNVME-CONFIGFS'; +my $nvmet_max_nsid = 0xfffffffe; +my $nvmet_max_command = 65536; +# Seconds to wait for the target lock while another node changes the target; +# a destroy or rollback can hold it for several seconds. +my $nvmet_lock_wait = 30; +# On the pool dataset: the highest NSID handed out for a volume of the pool. +my $nvmet_last_nsid = 'proxmox:nvme-last-nsid'; + +my @nvmet_zfs_props = + ('type', 'proxmox:nvme-subsys', 'proxmox:nvme-nsid', 'proxmox:nvme-uuid', $nvmet_last_nsid); +my %nvmet_zfs_keys = ( + type => 'type', + 'proxmox:nvme-subsys' => 'nqn', + 'proxmox:nvme-nsid' => 'nsid', + 'proxmox:nvme-uuid' => 'uuid', + $nvmet_last_nsid => 'last_nsid', +); +my @nvmet_cfs_attrs = qw( + enable device_path device_uuid buffered_io + attr_model attr_serial attr_allow_any_host + addr_trtype addr_adrfam addr_traddr addr_trsvcid +); +# Every command the plugin runs on the target, besides the `dd | tee` pipe of +# a key write. The configfs read runs find through env to set its locale. +my %nvmet_commands = map { $_ => 1 } qw( + cat chmod env grep ln mkdir modprobe mount printf rm rmdir test zfs +); + +my $RE_NQN = qr{ + \A + nqn \. + [A-Za-z0-9] [A-Za-z0-9.-]* + : + [A-Za-z0-9] [A-Za-z0-9._:-]* + \z +}nxx; +my $RE_IPV4_PORTAL = qr{ + \A + (?
[^:]+) + (?: : (? [0-9]+))? + \z +}nxx; +my $RE_IPV6_PORTAL = qr{ + \A + \[ (?
[^\]]+) \] + (?: : (? [0-9]+))? + \z +}nxx; +# A canonical IPv4-mapped IPv6 address. +my $RE_IPV4_MAPPED = qr{\A ::ffff: (?
[0-9.]+) \z}nxx; +my $RE_HOST_IFACE = qr{\A [A-Za-z0-9_.-]+ \z}nxx; +my $RE_DHCHAP_KEY = qr{ + \A DHHC-1 : (? 0[0-3]) : (? [A-Za-z0-9+/]+ ={0,2}) : \z +}nxx; +# A stopped PVE task: the worker's signal handler dies with the first text, +# run_command adds the command context, and run_fork_with_timeout reports +# the signal in its own words. +my $RE_TASK_INTERRUPT = qr{ + \A (?: (?: command \x20 ' .* ' \x20 failed: \x20 )? received \x20 interrupt + | interrupted \x20 by \x20 unexpected \x20 signal ) \n \z +}nsxx; +my $RE_COMMAND_TIMEOUT = qr{ + \A (?: command \x20 ' .* ' \x20 failed: \x20 )? got \x20 timeout \n \z +}nsxx; +my $RE_COMMAND_EXIT = + qr{\A command \x20 ' .* ' \x20 failed: \x20 exit \x20 code \x20 (? [0-9]+) \n \z}nsxx; +my $RE_FABRICS_VALUE = qr{\A [^\s,\x00]+ \z}nxx; +my $RE_FABRICS_RESULT = qr{\A instance=(? [0-9]+) ,cntlid= [0-9]+ \n \z}nxx; +my $RE_ZERO_UUID = qr{\A 0{8} (?: -0{4}){3} -0{12} \z}nxx; +my $RE_CONFIG_INT = qr{\A -?[0-9]+ \z}nxx; +my $RE_NVME_CONTROLLER = qr{\A nvme [0-9]+ \z}nxx; +my $RE_NVME_SUBSYSTEM = qr{\A nvme-subsys [0-9]+ \z}nxx; +my $RE_NVME_NAMESPACE = qr{\A nvme [0-9]+ n [0-9]+ \z}nxx; +my $RE_NVME_PARTITION = qr{\A nvme [0-9]+ n [0-9]+ p [0-9]+ \z}nxx; +my $RE_TRADDR = qr{(?: \A | ,) traddr=(?[^,]+)}nxx; +my $RE_TRSVCID = qr{(?: \A | ,) trsvcid=(?[^,]+)}nxx; +my $RE_HOST_IFACE_ADDRESS = qr{(?: \A | ,) host_iface=(?[^,]+)}nxx; +my $RE_ZVOL_OWNER = qr{^ (?:vm|base|subvol|basevol)- (?\d+) - \S+ $}nxx; +my $RE_VOLNAME = qr{ + ^ + (?: (? (?:base|basevol)- (?\d+) - \S+) /)? + (? (?base|basevol|vm|subvol)- (?\d+) - \S+) + $ +}nxx; +my $RE_BASE_SNAPSHOT = qr{^ (?\S+) \@__base__ $}nxx; +my $RE_UNSIGNED_INTEGER = qr{^ (?\d+) $}nxx; + +my $RE_NVMET_UUID = qr{\A [0-9a-fA-F]{8} (?: - [0-9a-fA-F]{4}){3} - [0-9a-fA-F]{12} \z}nxx; +my $RE_NVMET_NSID = qr{\A [1-9] [0-9]* \z}nxx; +my $RE_NVMET_POOL = qr{\A [A-Za-z0-9] [A-Za-z0-9_.:/-]* \z}nxx; +my $RE_NVMET_DATASET_NAME = qr{\A [A-Za-z0-9] [A-Za-z0-9_.:-]* \z}nxx; +my $RE_NVMET_SNAPSHOT = qr{\A [A-Za-z0-9_.:-]+ \z}nxx; +my $RE_NVMET_TEMPLATE_ZVOL = qr{/ (? vm | base) - (? [^/]+) \z}nxx; +my $RE_NVMET_INHERITED = qr{\A inherited \x20 from \x20 (? .+) \z}nxx; +# Errors after which the target state is not known: an interrupted task, or a +# change whose command may still be running on the target. +my $RE_NVMET_ABANDONED = qr{ + (?: \A received \x20 interrupt | ; \x20 the \x20 target \x20 state \x20 is \x20 unknown) \n \z +}nxx; + +# Lines of the configfs read (nvmet_configfs_step below): `D `, +# `L `, `:` from grep, and ` ` from +# sha256sum. Attribute names anchor the split, so an NQN containing ':' is +# never ambiguous. +my $cfs_root = quotemeta($nvmet_root); +my $RE_CFS_TOP = qr{\A D \x20 $cfs_root (?: / (? hosts | ports | subsystems))? \z}nxx; +my $RE_CFS_SUBSYS_DIR = qr{\A D \x20 $cfs_root /subsystems/ (? [^/]+) \z}nxx; +my $RE_CFS_SUBSYS_GROUP = qr{ + \A D \x20 $cfs_root /subsystems/ [^/]+ / (?: namespaces | allowed_hosts) \z +}nxx; +my $RE_CFS_NS_DIR = qr{ + \A D \x20 $cfs_root /subsystems/ (? [^/]+) /namespaces/ (? [0-9]+) \z +}nxx; +my $RE_CFS_PORT_DIR = qr{\A D \x20 $cfs_root /ports/ (? [0-9]+) \z}nxx; +my $RE_CFS_PORT_GROUP = qr{\A D \x20 $cfs_root /ports/ [0-9]+ /subsystems \z}nxx; +my $RE_CFS_HOST_DIR = qr{\A D \x20 $cfs_root /hosts/ (? [^/]+) \z}nxx; +my $RE_CFS_PORT_LINK = qr{ + \A L \x20 $cfs_root /ports/ (? [0-9]+) /subsystems/ (? [^/]+) \z +}nxx; +my $RE_CFS_ACL_LINK = qr{ + \A L \x20 $cfs_root /subsystems/ (? [^/]+) /allowed_hosts/ (? [^/]+) \z +}nxx; +my $RE_CFS_OTHER_ENTRY = qr{\A [DL] \x20 $cfs_root /}nxx; +my $RE_CFS_NS_ATTR = qr{ + \A $cfs_root /subsystems/ (? [^/]+) /namespaces/ (? [0-9]+) + / (? enable | device_path | device_uuid | buffered_io) : (? .*) \z +}nxx; +my $RE_CFS_SUBSYS_ATTR = qr{ + \A $cfs_root /subsystems/ (? [^/]+) + / (? attr_model | attr_serial | attr_allow_any_host) : (? .*) \z +}nxx; +my $RE_CFS_PORT_ATTR = qr{ + \A $cfs_root /ports/ (? [0-9]+) + / (? addr_trtype | addr_adrfam | addr_traddr | addr_trsvcid) : (? .*) \z +}nxx; +my $RE_CFS_KEY_DIGEST = qr{ + \A (? [0-9a-f]{64}) \x20 [\x20*] $cfs_root /hosts/ (? [^/]+) /dhchap_key \z +}nxx; +my $RE_CFS_OTHER_ATTR = qr{\A $cfs_root /}nxx; + +# --------------------------------------------------------------------------- +# Target addressing and validated names +# --------------------------------------------------------------------------- + +my sub nvmet_server($scfg) { + return $scfg->{server}; +} + +my sub nvmet_ssh_key($scfg) { + return "$id_rsa_path/" . nvmet_server($scfg) . '_id_rsa'; +} + +my sub nvmet_lock_id($scfg) { + my $server = nvmet_server($scfg) // ''; + return 'zfsnvme-' . ($server =~ s/[^A-Za-z0-9.-]/_/gr); +} + +my sub secret_path($storeid) { + return "$secret_dir/$storeid.nvme-dhchap"; +} + +my sub nvmet_nqn($scfg) { + my $nqn = $scfg->{subsysnqn}; + die "invalid NVMe subsystem NQN\n" if !defined($nqn) || !verify_nvme_nqn($nqn, 1); + return $nqn; +} + +my sub nvmet_pool($scfg) { + my $pool = $scfg->{pool}; + die "invalid ZFS pool name\n" if !defined($pool) || $pool !~ $RE_NVMET_POOL; + return $pool; +} + +my sub nvmet_dataset($scfg, $name) { + die "invalid ZFS volume name\n" if !defined($name) || $name !~ $RE_NVMET_DATASET_NAME; + return nvmet_pool($scfg) . "/$name"; +} + +# zfs destroy reads '%' and ',' in a snapshot name as a range and a list. +my sub nvmet_snapshot($scfg, $name, $snap) { + die "invalid snapshot name\n" if !defined($snap) || $snap !~ $RE_NVMET_SNAPSHOT; + return nvmet_dataset($scfg, $name) . "\@$snap"; +} + +my sub nvmet_valid_nsid($value) { + return defined($value) && $value =~ $RE_NVMET_NSID && $value <= $nvmet_max_nsid; +} + +# IPv6 addresses have many spellings; compare and store only the canonical one. +my sub nvmet_canonical_address($address) { + return $address if !defined($address) || index($address, ':') < 0; + my $packed = inet_pton(AF_INET6, $address) // return $address; + return inet_ntop(AF_INET6, $packed); +} + +# The address family and address that a portal of parse_nvme_portals reaches: +# an IPv4-mapped IPv6 address reaches the IPv4 address. +my sub nvmet_listener_address($portal) { + return ('ipv4', $+{address}) + if $portal->{family} eq 'ipv6' && $portal->{address} =~ $RE_IPV4_MAPPED; + return $portal->@{qw(family address)}; +} + +# Target-provided names only appear in messages when they are well-formed. +my sub nvmet_label($name) { + return $name =~ $RE_NQN ? " '$name'" : ''; +} + +# The other name of a zvol that a template conversion renames: vm-* and base-*. +my sub nvmet_template_twin($name) { + return '' if $name !~ $RE_NVMET_TEMPLATE_ZVOL; + my $twin = ($+{type} eq 'vm' ? 'base' : 'vm') . "-$+{suffix}"; + return substr($name, 0, $-[0]) . "/$twin"; +} + +# The timeout of a call that changes ZFS: ZFS waits for transaction groups, +# which a busy pool can take minutes to sync. +my sub nvmet_long_timeout() { + return PVE::RPCEnvironment->is_worker() ? 60 * 60 : 60; +} + +# --------------------------------------------------------------------------- +# Remote steps +# +# A step is an argv array, a write `{ write => $path, value => $value }` or a +# key write `{ key => [$path, ...] }` fed from stdin. A unit is a list of steps +# that must stay in one call. +# --------------------------------------------------------------------------- + +my sub nvmet_subsys_path($nqn) { + return "$nvmet_root/subsystems/$nqn"; +} + +my sub nvmet_ns_path($nqn, $nsid) { + return "$nvmet_root/subsystems/$nqn/namespaces/$nsid"; +} + +my sub nvmet_port_path($id) { + return "$nvmet_root/ports/$id"; +} + +my sub nvmet_host_path($hostnqn) { + return "$nvmet_root/hosts/$hostnqn"; +} + +my sub nvmet_write($path, $value) { + return { write => $path, value => $value }; +} + +# Parts of a parsed configfs state; missing parts read as empty. +my sub nvmet_namespaces($cfs, $nqn) { + my $subsys = $cfs->{subsystems}->{$nqn}; + return ($subsys ? $subsys->{namespaces} : undef) // {}; +} + +my sub nvmet_port_linked($cfs, $id, $nqn) { + my $port = $cfs->{ports}->{$id}; + return $port && ($port->{links} // {})->{$nqn}; +} + +my sub nvmet_linked_hosts($cfs) { + my %linked; + for my $subsys (values $cfs->{subsystems}->%*) { + $linked{$_} = 1 for keys(($subsys->{acl} // {})->%*); + } + return \%linked; +} + +# A subsystem whose creation by this plugin stopped after its model was +# written, before its serial: nobody else writes our model, and it exports, +# allows and publishes nothing, so completing it takes nothing from anybody. +# A subsystem with the kernel's default model may belong to another tool. +my sub nvmet_unfinished_subsystem($cfs, $nqn) { + my $subsys = $cfs->{subsystems}->{$nqn}; + return 0 if !$subsys; + return + $subsys->{attr_model} eq $nvmet_model + && !($subsys->{namespaces} // {})->%* + && !($subsys->{acl} // {})->%* + && !grep { nvmet_port_linked($cfs, $_, $nqn) } keys $cfs->{ports}->%*; +} + +# Same write order as the kernel requires: the device before enabling it. +my sub nvmet_build_unit($nqn, $nsid, $uuid, $dev) { + my $ns = nvmet_ns_path($nqn, $nsid); + return [ + ['test', '-b', $dev], + ['mkdir', $ns], + nvmet_write("$ns/device_path", $dev), + nvmet_write("$ns/device_uuid", $uuid), + nvmet_write("$ns/buffered_io", 0), + nvmet_write("$ns/enable", 1), + ]; +} + +my sub nvmet_enable_unit($nqn, $nsid, $dev) { + return [['test', '-b', $dev], nvmet_write(nvmet_ns_path($nqn, $nsid) . '/enable', 1)]; +} + +my sub nvmet_unexport_unit($nqn, $nsid, $enabled) { + my $ns = nvmet_ns_path($nqn, $nsid); + return [($enabled ? (nvmet_write("$ns/enable", 0)) : ()), ['rmdir', $ns]]; +} + +my sub nvmet_identity($nqn, $nsid, $uuid) { + return ("proxmox:nvme-subsys=$nqn", "proxmox:nvme-nsid=$nsid", "proxmox:nvme-uuid=$uuid"); +} + +# The pool and its direct children with their identity properties. +my sub nvmet_zfs_inventory_step($pool) { + return [ + 'zfs', + 'get', + '-H', + '-p', + '-d', + '1', + '-t', + 'filesystem,volume', + '-o', + 'name,property,value,source', + join(',', @nvmet_zfs_props), + $pool, + ]; +} + +# Every configfs object the plugin uses, in one traversal. Keys are never +# read back; with host NQNs, only their sha256 digests are returned. Only the +# kernel's own groups of objects the plugin does not use are skipped, so an +# object of another tool with one of their names is still listed. +my sub nvmet_configfs_step($hostnqns) { + my @names = map { ('-o', '-name', $_) } @nvmet_cfs_attrs; + shift @names; + my @keys = map { ('-o', '-path', nvmet_host_path($_) . '/dhchap_key') } $hostnqns->@*; + shift @keys; + return [ + 'env', + 'LC_ALL=C', + 'find', + $nvmet_root, + '(', + '-path', + "$nvmet_root/subsystems/*/passthru", + '-o', + '-path', + "$nvmet_root/ports/*/ana_groups", + '-o', + '-path', + "$nvmet_root/ports/*/referrals", + ')', + '-type', + 'd', + '-prune', + '-o', + '-type', + 'd', + '-exec', + 'printf', + 'D %s\n', + '{}', + '+', + '-o', + '-type', + 'l', + '-exec', + 'printf', + 'L %s\n', + '{}', + '+', + '-o', + '-type', + 'f', + '(', + @names, + ')', + '-exec', + 'grep', + '', + '/dev/null', + '{}', + '+', + (@keys ? ('-o', '-type', 'f', '(', @keys, ')', '-exec', 'sha256sum', '{}', '+') : ()), + ]; +} + +# The whole target state: the ZFS inventory, a marker and configfs. Only a +# locked read loads nvmet first (modprobe). +my sub nvmet_state_steps($pool, $hostnqns, $modprobe) { + return [ + ($modprobe ? (['modprobe', 'nvmet_tcp']) : ()), + nvmet_zfs_inventory_step($pool), + ['printf', '%s\n', $nvmet_marker], + nvmet_configfs_step($hostnqns), + ]; +} + +# --------------------------------------------------------------------------- +# Renderer and runner +# --------------------------------------------------------------------------- + +my sub nvmet_valid_word($word) { + return defined($word) && !ref($word) && $word !~ /[\0\r\n]/; +} + +my sub nvmet_valid_path($path) { + return nvmet_valid_word($path) && $path =~ m{\A/}; +} + +my sub nvmet_quote($word) { + return PVE::Tools::shellquote($word); +} + +# Renders steps into one POSIX shell command line: simple commands joined with +# `&&`, plus `>` for configfs writes and one `dd | tee` pipe for a key read from +# stdin. There are no loops, variables, conditionals or substitutions. Every +# word is quoted; all operands are built from validated configuration. +sub _nvmet_render($steps) { + my $invalid = "internal error: invalid NVMe target step\n"; + die $invalid if ref($steps) ne 'ARRAY' || !$steps->@*; + + my $keys = 0; + my @commands; + for my $step ($steps->@*) { + if (ref($step) eq 'ARRAY') { + die $invalid if !$step->@* || grep { !nvmet_valid_word($_) } $step->@*; + die $invalid if !$nvmet_commands{ $step->[0] }; + push @commands, join(' ', map { nvmet_quote($_) } $step->@*); + } elsif (ref($step) eq 'HASH' && exists($step->{write})) { + die $invalid + if keys($step->%*) != 2 + || !nvmet_valid_path($step->{write}) + || !nvmet_valid_word($step->{value}); + push @commands, + q{printf '%s\n' } + . nvmet_quote($step->{value}) . ' > ' + . nvmet_quote($step->{write}); + } elsif (ref($step) eq 'HASH' && exists($step->{key})) { + my $paths = $step->{key}; + die $invalid + if keys($step->%*) != 1 + || ref($paths) ne 'ARRAY' + || !$paths->@* + || grep { !nvmet_valid_path($_) } $paths->@*; + $keys++; + die $invalid if $keys > 1; + # configfs needs the whole key in one write(2). dd collects its + # input into one output block: GNU dd without bs=, busybox dd only + # when ibs and obs differ. tee writes it to every host object. + push @commands, + 'dd ibs=4096 obs=8192 2>/dev/null | tee ' + . join(' ', map { nvmet_quote($_) } $paths->@*) + . ' >/dev/null'; + } else { + die $invalid; + } + } + return join(' && ', @commands); +} + +# Packs whole units into calls of at most $max bytes of rendered command, the +# single argument sshd runs, below its limit. A rendered command is ASCII +# (every operand is validated), so its length is its size in bytes. +sub _nvmet_chunk($units, $max = $nvmet_max_command) { + my (@chunks, @current); + for my $unit ($units->@*) { + my @next = (@current, $unit->@*); + if (length(_nvmet_render(\@next)) > $max) { + die "internal error: NVMe target command too long\n" + if !@current || length(_nvmet_render($unit)) > $max; + push @chunks, [@current]; + @next = $unit->@*; + } + @current = @next; + } + push @chunks, [@current] if @current; + return \@chunks; +} + +# Never propagate the context run_command adds to the task marker: it quotes +# the command line. +my sub rethrow_task_interrupt($error) { + die "received interrupt\n" if $error =~ $RE_TASK_INTERRUPT; +} + +# The only code that runs anything on the target. Returns +# { rc => exit code, or -1 when ssh did not run to its end, out => [lines], +# err => text } and never dies on a failing command; a stopped task dies with +# the task marker. %opts: op (label, required), timeout and input (stdin, +# used for keys). +sub _nvmet_run($scfg, $steps, %opts) { + die "internal error: NVMe target call without label\n" if !defined($opts{op}); + my $command = _nvmet_render($steps); + die "internal error: NVMe target command too long\n" if length($command) > $nvmet_max_command; + my $cmd = [@ssh_cmd, '-i', nvmet_ssh_key($scfg), 'root@' . nvmet_server($scfg), $command]; + my (@out, $err); + my $rc = 0; + # Not noerr: run_command would then also swallow the interrupt of a stopped + # task, which must end the call. The exit code is taken from the exception. + eval { + run_command( + $cmd, + timeout => $opts{timeout} // 15, + outfunc => sub($line) { push @out, $line }, + errfunc => sub($line) { $err = $line if !defined($err) && $line ne '' }, + (defined($opts{input}) ? (input => $opts{input}) : ()), + ); + }; + if (my $error = $@) { + rethrow_task_interrupt($error); + return { rc => -1, out => [], err => 'timeout' } if $error =~ $RE_COMMAND_TIMEOUT; + return { rc => -1, out => [], err => 'ssh failed' } if $error !~ $RE_COMMAND_EXIT; + $rc = $+{code}; + } + return { rc => $rc, out => \@out, err => $err // ($rc ? "exit code $rc" : '') }; +} + +# Monotonic clock for the activation backoff and local waits (a test seam). +sub _now() { + return Time::HiRes::clock_gettime(Time::HiRes::CLOCK_MONOTONIC()); +} + +# Every local wait of the plugin (a test seam). +sub _sleep($seconds) { + Time::HiRes::sleep($seconds); + return; +} + +# --------------------------------------------------------------------------- +# Parsers (pure) +# --------------------------------------------------------------------------- + +sub _nvmet_split_state($text) { + my @parts = split /^\Q$nvmet_marker\E\n/m, $text, -1; + die "malformed NVMe target state\n" if @parts != 2; + return @parts; +} + +sub _nvmet_property_value($value, $source) { + die "malformed ZFS property value\n" + if !defined($value) + || !defined($source) + || $value eq '' + || $source eq '' + || $value =~ /[\r\n\t]/ + || $source =~ /[\r\n\t]/; + return $value if $source eq 'local' || $source eq 'received'; + return '-' if $value eq '-' && $source eq '-'; + if ($source =~ $RE_NVMET_INHERITED) { + die "invalid ZFS pool name\n" if $+{source} !~ $RE_NVMET_POOL; + return '-'; + } + die "invalid ZFS property source\n"; +} + +# Parses the ZFS inventory into { types => { dataset => type }, volumes => +# { dataset => { nqn, nsid, uuid } }, last_nsid => highest NSID handed out }. +# The inventory must be complete: the pool rows and all rows of every +# dataset, or the read fails. Children whose names PVE never uses (for +# example with a space) are kept as datasets but never count as owned. +sub _nvmet_parse_zfs_inventory($text, $pool) { + die "invalid ZFS pool name\n" if $pool !~ $RE_NVMET_POOL; + + my (%rows, %foreign); + for my $line (split /\n/, $text) { + my @fields = split /\t/, $line, -1; + die "malformed ZFS inventory row\n" if @fields != 4 || $line =~ /\r/; + my ($dataset, $property, $value, $source) = @fields; + if ($dataset ne $pool) { + die "ZFS dataset is outside configured pool\n" if index($dataset, "$pool/") != 0; + my $child = substr($dataset, length($pool) + 1); + die "ZFS dataset is outside configured pool\n" + if $child eq '' || index($child, '/') >= 0; + $foreign{$dataset} = 1 if $child !~ $RE_NVMET_DATASET_NAME; + } + my $on = $foreign{$dataset} ? '' : " on '$dataset'"; + my $key = $nvmet_zfs_keys{$property} // die "unexpected ZFS property$on\n"; + die "duplicate ZFS property '$property'$on\n" if exists($rows{$dataset}->{$key}); + if ($key eq 'type') { + die "invalid ZFS dataset type$on\n" + if ($value ne 'filesystem' && $value ne 'volume') || $source ne '-'; + $rows{$dataset}->{type} = $value; + } else { + $rows{$dataset}->{$key} = _nvmet_property_value($value, $source); + } + } + die "ZFS dataset inventory is missing the configured pool\n" if !$rows{$pool}; + + my (%types, %volumes); + for my $dataset (keys %rows) { + my $row = $rows{$dataset}; + die "incomplete ZFS properties" . ($foreign{$dataset} ? '' : " on '$dataset'") . "\n" + if keys($row->%*) != scalar(@nvmet_zfs_props); + $types{$dataset} = $row->{type}; + next if $row->{type} ne 'volume'; + # nvmet shows device_uuid and udev names the by-id link in lower case, + # so identities are compared in lower case. + $volumes{$dataset} = + $foreign{$dataset} + ? { nqn => '-', nsid => '-', uuid => '-' } + : { nqn => $row->{nqn}, nsid => $row->{nsid}, uuid => lc($row->{uuid}) }; + } + my $last_nsid = $rows{$pool}->{last_nsid}; + return { + types => \%types, + volumes => \%volumes, + last_nsid => nvmet_valid_nsid($last_nsid) ? $last_nsid : 0, + }; +} + +# Parses the configfs read into +# { subsystems => { nqn => { attr_model, attr_serial, attr_allow_any_host, +# acl => { hostnqn => 1 }, namespaces => { nsid => { enable, device_path, +# device_uuid, buffered_io } } } }, +# ports => { id => { addr_trtype, addr_adrfam, addr_traddr, addr_trsvcid, +# links => { nqn => 1 } } }, +# hosts => { hostnqn => { key_sha256 => hex or undef } } } +# Every object must be complete; unknown objects of newer kernels are ignored. +sub _nvmet_parse_configfs($text) { + my $state = { subsystems => {}, ports => {}, hosts => {} }; + my (%top, %dir); + + my $subsystem = sub($nqn) { + return $state->{subsystems}->{$nqn} //= { acl => {}, namespaces => {} }; + }; + my $namespace = sub($nqn, $nsid) { + return $subsystem->($nqn)->{namespaces}->{$nsid} //= {}; + }; + my $port = sub($id) { return $state->{ports}->{$id} //= { links => {} } }; + my $host = sub($hostnqn) { return $state->{hosts}->{$hostnqn} //= { key_sha256 => undef } }; + my $set = sub($object, $attr, $value) { + die "duplicate NVMe target configfs attribute\n" if exists($object->{$attr}); + $object->{$attr} = $value; + }; + + for my $line (split /\n/, $text) { + if ($line =~ $RE_CFS_TOP) { + $top{ $+{top} // 'root' } = 1; + } elsif ($line =~ $RE_CFS_SUBSYS_DIR) { + my $nqn = $+{nqn}; + $subsystem->($nqn); + $dir{"subsystem $nqn"} = 1; + } elsif ($line =~ $RE_CFS_SUBSYS_GROUP || $line =~ $RE_CFS_PORT_GROUP) { + # default groups, always present with their parent + } elsif ($line =~ $RE_CFS_NS_DIR) { + my ($nqn, $nsid) = @+{qw(nqn nsid)}; + $namespace->($nqn, $nsid); + $dir{"namespace $nqn/$nsid"} = 1; + } elsif ($line =~ $RE_CFS_PORT_DIR) { + my $id = $+{port}; + $port->($id); + $dir{"port $id"} = 1; + } elsif ($line =~ $RE_CFS_HOST_DIR) { + my $hostnqn = $+{host}; + $host->($hostnqn); + $dir{"host $hostnqn"} = 1; + } elsif ($line =~ $RE_CFS_PORT_LINK) { + my ($id, $nqn) = @+{qw(port nqn)}; + $port->($id)->{links}->{$nqn} = 1; + } elsif ($line =~ $RE_CFS_ACL_LINK) { + my ($nqn, $hostnqn) = @+{qw(nqn host)}; + $subsystem->($nqn)->{acl}->{$hostnqn} = 1; + } elsif ($line =~ $RE_CFS_OTHER_ENTRY) { + # objects of a shape the plugin does not use + } elsif ($line =~ $RE_CFS_NS_ATTR) { + my ($nqn, $nsid, $attr, $value) = @+{qw(nqn nsid attr value)}; + $set->($namespace->($nqn, $nsid), $attr, $value); + } elsif ($line =~ $RE_CFS_SUBSYS_ATTR) { + my ($nqn, $attr, $value) = @+{qw(nqn attr value)}; + $set->($subsystem->($nqn), $attr, $value); + } elsif ($line =~ $RE_CFS_PORT_ATTR) { + my ($id, $attr, $value) = @+{qw(port attr value)}; + $set->($port->($id), $attr, $value); + } elsif ($line =~ $RE_CFS_KEY_DIGEST) { + my ($sha, $hostnqn) = @+{qw(sha host)}; + my $object = $host->($hostnqn); + die "duplicate DH-HMAC-CHAP key digest" . nvmet_label($hostnqn) . "\n" + if defined($object->{key_sha256}); + $object->{key_sha256} = $sha; + } elsif ($line =~ $RE_CFS_OTHER_ATTR) { + # attribute of an object shape the plugin does not use + } else { + die "unexpected NVMe target configfs line\n"; + } + } + + for my $name (qw(root hosts ports subsystems)) { + die "NVMe target configfs is unavailable\n" if !$top{$name}; + } + for my $nqn (keys $state->{subsystems}->%*) { + my $subsys = $state->{subsystems}->{$nqn}; + die "incomplete NVMe subsystem" . nvmet_label($nqn) . "\n" + if !$dir{"subsystem $nqn"} + || grep { !defined($subsys->{$_}) } qw(attr_model attr_serial attr_allow_any_host); + for my $nsid (keys $subsys->{namespaces}->%*) { + my $ns = $subsys->{namespaces}->{$nsid}; + die "incomplete NVMe namespace '$nsid' of subsystem" . nvmet_label($nqn) . "\n" + if !$dir{"namespace $nqn/$nsid"} + || grep { !defined($ns->{$_}) } qw(enable device_path device_uuid buffered_io); + } + } + for my $id (keys $state->{ports}->%*) { + my $object = $state->{ports}->{$id}; + my @attrs = qw(addr_trtype addr_adrfam addr_traddr addr_trsvcid); + die "incomplete NVMe port '$id'\n" + if !$dir{"port $id"} || grep { !defined($object->{$_}) } @attrs; + } + for my $hostnqn (keys $state->{hosts}->%*) { + die "DH-HMAC-CHAP key digest for unknown NVMe host" . nvmet_label($hostnqn) . "\n" + if !$dir{"host $hostnqn"}; + } + return $state; +} + +sub _nvmet_parse_mounts($text) { + for my $line (split /\n/, $text) { + my (undef, $mountpoint, $type) = split / /, $line; + return 1 if ($mountpoint // '') eq '/sys/kernel/config' && ($type // '') eq 'configfs'; + } + return 0; +} + +# --------------------------------------------------------------------------- +# Planners (pure) +# --------------------------------------------------------------------------- + +sub _nvmet_serial($nqn) { + return 'PVEZFS' . substr(sha256_hex($nqn), 0, 14); +} + +sub _nvmet_template_name($dataset) { + die "only VM zvols can become templates\n" + if $dataset !~ $RE_NVMET_TEMPLATE_ZVOL || $+{type} ne 'vm'; + return nvmet_template_twin($dataset); +} + +# The durable identity (NSID, UUID) of a volume owned by $nqn. Refuses a +# volume that another owned volume duplicates anywhere in the pool. +sub _nvmet_owned_identity($inv, $nqn, $dataset) { + my $type = $inv->{types}->{$dataset} // die "ZFS volume '$dataset' does not exist\n"; + die "'$dataset' is not a ZFS volume\n" if $type ne 'volume'; + my $row = $inv->{volumes}->{$dataset}; + die "ZFS volume '$dataset' is not owned by NVMe subsystem '$nqn'\n" if $row->{nqn} ne $nqn; + die "invalid NSID on '$dataset'\n" if !nvmet_valid_nsid($row->{nsid}); + die "invalid namespace UUID on '$dataset'\n" if $row->{uuid} !~ $RE_NVMET_UUID; + for my $other (keys $inv->{volumes}->%*) { + next if $other eq $dataset; + my $identity = $inv->{volumes}->{$other}; + next if $identity->{nqn} ne $nqn; + die "duplicate NVMe identity on '$dataset'\n" + if $identity->{nsid} eq $row->{nsid} || $identity->{uuid} eq $row->{uuid}; + } + return ($row->{nsid}, $row->{uuid}); +} + +# { nsid => [uuid, device] } for every volume owned by $nqn, after validating +# all of them, so a bad identity refuses the whole plan before any change. +sub _nvmet_desired_namespaces($inv, $nqn) { + my (%desired, %uuids); + for my $dataset (sort keys $inv->{volumes}->%*) { + my $row = $inv->{volumes}->{$dataset}; + next if $row->{nqn} ne $nqn; + die "invalid NSID on '$dataset'\n" if !nvmet_valid_nsid($row->{nsid}); + die "invalid namespace UUID on '$dataset'\n" if $row->{uuid} !~ $RE_NVMET_UUID; + die "duplicate NSID '$row->{nsid}'\n" if $desired{ $row->{nsid} }; + die "duplicate namespace UUID '$row->{uuid}'\n" if $uuids{ lc($row->{uuid}) }++; + $desired{ $row->{nsid} } = [$row->{uuid}, "/dev/zvol/$dataset"]; + } + return \%desired; +} + +# NSIDs grow monotonically, also past the last NSID handed out in the pool, +# so a host that missed a namespace removal never sees a new volume at an old +# NSID. Only after the last NSID is used does the allocation fall back to the +# lowest free one. +sub _nvmet_allocate_nsid($inv, $cfs, $nqn) { + my %used; + for my $row (values $inv->{volumes}->%*) { + $used{ $row->{nsid} } = 1 if $row->{nqn} eq $nqn && nvmet_valid_nsid($row->{nsid}); + } + $used{$_} = 1 for keys nvmet_namespaces($cfs, $nqn)->%*; + + my $highest = max(0, $inv->{last_nsid} // 0, keys %used); + return $highest + 1 if $highest < $nvmet_max_nsid; + my $nsid = 1; + $nsid++ while $used{$nsid}; + die "no free namespace ID\n" if $nsid > $nvmet_max_nsid; + return $nsid; +} + +# Export state of (NSID, UUID, device) in the subsystem: present, disabled, +# absent, incomplete (a disabled namespace with another identity, for example +# one being built), stale (enabled on the other template name of the zvol) or +# nosubsys. Dies on a conflict. +sub _nvmet_export_state($cfs, $nqn, $nsid, $uuid, $dev) { + return 'nosubsys' if !$cfs->{subsystems}->{$nqn}; + my $namespaces = nvmet_namespaces($cfs, $nqn); + for my $id (keys $namespaces->%*) { + next if $id eq $nsid; + die "namespace UUID '$uuid' is already in use\n" + if $namespaces->{$id}->{device_uuid} eq $uuid; + } + my $ns = $namespaces->{$nsid} // return 'absent'; + if ($ns->{device_uuid} eq $uuid && $ns->{device_path} eq $dev) { + return $ns->{enable} eq '1' ? 'present' : 'disabled'; + } + # A template conversion or its undo that ssh gave up on can rename the + # zvol on the target after an activation exported it under its old name. + my $twin = nvmet_template_twin($dev); + return 'stale' + if $ns->{enable} eq '1' + && $ns->{device_uuid} eq $uuid + && $twin ne '' + && $ns->{device_path} eq $twin; + die "NVMe namespace ID '$nsid' has a different identity; refusing to replace it\n" + if $ns->{enable} ne '0'; + return 'incomplete'; +} + +# Units that export (NSID, UUID, device) and the state they start from: +# absent, present, disabled, reclaim or stale. A disabled namespace with +# another identity, or a stale one, is rebuilt, which is only correct under +# the target lock. A stale namespace belongs to a template, or to a volume +# whose template conversion failed, so no guest uses it. +sub _nvmet_plan_export($cfs, $nqn, $nsid, $uuid, $dev) { + die "NVMe subsystem does not exist\n" if !$cfs->{subsystems}->{$nqn}; + my $state = _nvmet_export_state($cfs, $nqn, $nsid, $uuid, $dev); + return ([nvmet_build_unit($nqn, $nsid, $uuid, $dev)], 'absent') if $state eq 'absent'; + return ([], 'present') if $state eq 'present'; + return ([nvmet_enable_unit($nqn, $nsid, $dev)], 'disabled') if $state eq 'disabled'; + my $unexport = nvmet_unexport_unit($nqn, $nsid, $state eq 'stale'); + my $build = nvmet_build_unit($nqn, $nsid, $uuid, $dev); + return ([[$unexport->@*, $build->@*]], $state eq 'stale' ? 'stale' : 'reclaim'); +} + +# Units that remove the namespace exporting $uuid, and its previous state. +sub _nvmet_plan_unexport($cfs, $nqn, $uuid) { + my $namespaces = nvmet_namespaces($cfs, $nqn); + my @matches = grep { $namespaces->{$_}->{device_uuid} eq $uuid } + sort { $a <=> $b } keys $namespaces->%*; + die "duplicate namespace UUID '$uuid'\n" if @matches > 1; + return ([], undef) if !@matches; + my $enabled = $namespaces->{ $matches[0] }->{enable} eq '1'; + return ( + [nvmet_unexport_unit($nqn, $matches[0], $enabled)], + { nsid => $matches[0], enabled => $enabled }, + ); +} + +# The port serving a portal ({ family, address, port } of parse_nvme_portals). +# Among complete, matching ports, the one already linked to $nqn wins, then +# the lowest id. Incomplete ports never match. +sub _nvmet_find_port($cfs, $nqn, $portal) { + my $wanted = nvmet_canonical_address($portal->{address}); + my @matches = grep { + my $port = $cfs->{ports}->{$_}; + $port->{addr_trtype} eq 'tcp' + && $port->{addr_adrfam} eq $portal->{family} + && $port->{addr_traddr} ne '' + && $port->{addr_trsvcid} eq "$portal->{port}" + && nvmet_canonical_address($port->{addr_traddr}) eq $wanted + } sort { + $a <=> $b + } keys $cfs->{ports}->%*; + my @linked = grep { nvmet_port_linked($cfs, $_, $nqn) } @matches; + return $linked[0] // $matches[0]; +} + +# Units that bring the target to the configured state, except publishing: +# objects (subsystem, ports, hosts), keys, ACLs, namespaces in NSID order, +# then removal of undesired namespaces. +# $conf: { nqn, portals (of parse_nvme_portals), hostnqns, keysha } +sub _nvmet_plan_activation($inv, $cfs, $conf) { + my ($nqn, $keysha) = $conf->@{qw(nqn keysha)}; + my $desired = _nvmet_desired_namespaces($inv, $nqn); # refuse before any unit + my $subsys_path = nvmet_subsys_path($nqn); + my $subsys = $cfs->{subsystems}->{$nqn}; + my (@objects, @key_hosts, @acls, @namespaces, %port_ids); + + if (!$subsys) { + push @objects, + [ + ['mkdir', $subsys_path], + nvmet_write("$subsys_path/attr_model", $nvmet_model), + nvmet_write("$subsys_path/attr_serial", _nvmet_serial($nqn)), + nvmet_write("$subsys_path/attr_allow_any_host", 0), + ]; + } else { + my $serial = _nvmet_serial($nqn); + my @attrs; + if ($subsys->{attr_model} ne $nvmet_model || $subsys->{attr_serial} ne $serial) { + die "refusing to take over existing NVMe subsystem '$nqn': it was not created by" + . " Proxmox VE; remove it on the target or use another NQN\n" + if !nvmet_unfinished_subsystem($cfs, $nqn); + push @attrs, nvmet_write("$subsys_path/attr_serial", $serial); + } + push @attrs, nvmet_write("$subsys_path/attr_allow_any_host", 0) + if $subsys->{attr_allow_any_host} ne '0'; + push @objects, [@attrs] if @attrs; + } + + my %taken = map { $_ => 1 } keys $cfs->{ports}->%*; + for my $portal ($conf->{portals}->@*) { + my $id = _nvmet_find_port($cfs, $nqn, $portal); + if (!defined($id)) { + $id = 1; + $id++ while $taken{$id}; + $taken{$id} = 1; + my $port_path = nvmet_port_path($id); + push @objects, + [ + ['mkdir', $port_path], + nvmet_write("$port_path/addr_trtype", 'tcp'), + nvmet_write("$port_path/addr_adrfam", $portal->{family}), + nvmet_write("$port_path/addr_traddr", $portal->{address}), + nvmet_write("$port_path/addr_trsvcid", $portal->{port}), + ]; + } + $port_ids{$id} = 1; + } + + # nvmet keeps the key on the global host object. A key that differs from + # ours is only replaced while no subsystem links the host. + my $linked = nvmet_linked_hosts($cfs); + for my $hostnqn ($conf->{hostnqns}->@*) { + my $host_path = nvmet_host_path($hostnqn); + if (my $host = $cfs->{hosts}->{$hostnqn}) { + my $sha = $host->{key_sha256} + // die "NVMe target does not expose a DH-HMAC-CHAP key for host '$hostnqn'\n"; + if ($sha ne $keysha) { + die "refusing to replace an in-use DH-HMAC-CHAP key for '$hostnqn'\n" + if $linked->{$hostnqn}; + push @key_hosts, $host_path; + } + } else { + push @objects, [['mkdir', $host_path]]; + push @key_hosts, $host_path; + } + push @acls, [['ln', '-s', $host_path, "$subsys_path/allowed_hosts/$hostnqn"]] + if !$subsys || !($subsys->{acl} // {})->{$hostnqn}; + } + # nvmet creates the key attributes world-readable. The chmod before the + # write keeps every later open from reading the key; a descriptor opened + # while an attribute was still world-readable is not affected by it. + my @keys; + if (@key_hosts) { + my @attrs = map { ("$_/dhchap_key", "$_/dhchap_ctrl_key") } @key_hosts; + @keys = ([['chmod', '0600', @attrs], { key => [map { "$_/dhchap_key" } @key_hosts] }]); + } + + my $current = { subsystems => { $nqn => $subsys // { acl => {}, namespaces => {} } } }; + for my $nsid (sort { $a <=> $b } keys $desired->%*) { + my ($units) = _nvmet_plan_export($current, $nqn, $nsid, $desired->{$nsid}->@*); + push @namespaces, $units->@*; + } + my $existing = nvmet_namespaces($cfs, $nqn); + for my $nsid (sort { $a <=> $b } keys $existing->%*) { + next if $desired->{$nsid}; + push @namespaces, nvmet_unexport_unit($nqn, $nsid, $existing->{$nsid}->{enable} eq '1'); + } + + return { + prepublish => [@objects, @keys, @acls, @namespaces], + port_ids => [sort { $a <=> $b } keys %port_ids], + }; +} + +# One call per port: each link repeats the guards, so it only succeeds while +# the subsystem denies unknown hosts and every configured host has its ACL. +# Under the target lock, only a manager of the target outside this cluster +# can change that after the verify read; the guards stop the publish then. +# Links to ports that are no longer configured are removed. +# Returns [{ port => $id, action => 'link' | 'unlink', steps => [...] }, ...]. +sub _nvmet_plan_publish($cfs, $nqn, $port_ids, $hostnqns) { + my $subsys_path = nvmet_subsys_path($nqn); + my %wanted = map { $_ => 1 } $port_ids->@*; + my @guards = ( + ['grep', '-qx', '0', "$subsys_path/attr_allow_any_host"], + map { ['test', '-L', "$subsys_path/allowed_hosts/$_"] } $hostnqns->@*, + ); + + my @units; + for my $id ($port_ids->@*) { + next if nvmet_port_linked($cfs, $id, $nqn); + my $link = ['ln', '-s', $subsys_path, nvmet_port_path($id) . "/subsystems/$nqn"]; + push @units, { port => $id, action => 'link', steps => [@guards, $link] }; + } + for my $id (sort { $a <=> $b } keys $cfs->{ports}->%*) { + next if $wanted{$id} || !nvmet_port_linked($cfs, $id, $nqn); + my $unlink = ['rm', nvmet_port_path($id) . "/subsystems/$nqn"]; + push @units, { port => $id, action => 'unlink', steps => [$unlink] }; + } + return \@units; +} + +# Teardown units for the subsystem and the hosts that may become orphans. +sub _nvmet_plan_delete_target($inv, $cfs, $nqn, $hostnqns) { + for my $dataset (sort keys $inv->{volumes}->%*) { + die "refusing to delete NVMe subsystem '$nqn': owned ZFS volume '$dataset' exists\n" + if $inv->{volumes}->{$dataset}->{nqn} eq $nqn; + } + my $subsys = $cfs->{subsystems}->{$nqn} // return ([], [$hostnqns->@*]); + die "refusing to delete foreign NVMe subsystem '$nqn'\n" + if $subsys->{attr_model} ne $nvmet_model || $subsys->{attr_serial} ne _nvmet_serial($nqn); + die "refusing to delete NVMe subsystem '$nqn': namespaces remain\n" + if nvmet_namespaces($cfs, $nqn)->%*; + + my $subsys_path = nvmet_subsys_path($nqn); + my @acl = sort keys(($subsys->{acl} // {})->%*); + die "refusing to delete NVMe subsystem '$nqn': it allows a malformed host name\n" + if grep { $_ !~ $RE_NQN } @acl; + my @links = ( + ( + map { nvmet_port_path($_) . "/subsystems/$nqn" } + grep { nvmet_port_linked($cfs, $_, $nqn) } + sort { $a <=> $b } keys $cfs->{ports}->%* + ), + (map { "$subsys_path/allowed_hosts/$_" } @acl), + ); + # A retry after a teardown that failed at its rmdir finds no ACL left, so + # the configured hosts are candidates as well. + my %candidates = map { $_ => 1 } @acl, $hostnqns->@*; + return ( + [[(@links ? (['rm', @links]) : ()), ['rmdir', $subsys_path]]], [sort keys %candidates], + ); +} + +# Candidate hosts that exist and that no subsystem links any more. +sub _nvmet_plan_orphan_hosts($cfs, $candidates) { + my $linked = nvmet_linked_hosts($cfs); + my %seen; + my @orphans = map { nvmet_host_path($_) } + grep { !$seen{$_}++ && $_ =~ $RE_NQN && $cfs->{hosts}->{$_} && !$linked->{$_} } + sort $candidates->@*; + return @orphans ? [[['rmdir', @orphans]]] : []; +} + +# --------------------------------------------------------------------------- +# Target lock and remote execution +# --------------------------------------------------------------------------- + +my %nvmet_lock_owner; # lock id => pid of the process holding the domain lock + +# Runs $code under the pmxcfs domain lock of the target. Like the pmxcfs +# lock itself, it is not re-entrant, and a child forked inside is not the +# owner. +my sub nvmet_locked($scfg, $code) { + my $id = nvmet_lock_id($scfg); + die "cluster not quorate - refusing NVMe target changes\n" + if !PVE::Cluster::check_cfs_quorum(1); + my $res = PVE::Cluster::cfs_lock_domain( + $id, + $nvmet_lock_wait, + sub { + local $nvmet_lock_owner{$id} = $$; + return $code->(); + }, + ); + die $@ if $@; + return $res; +} + +# Runs steps on the target. %opts: op, timeout, input, change (the steps +# change the target, which needs the lock). +# +# ssh does not stop a command on the target when it gives up on it: a change +# whose connection broke (exit code 255), any call under the lock that ssh +# did not run to its end (-1: its timeout, or ssh was killed), and a call +# during which the task was stopped (it dies with the task marker) may still +# be running there. Its outcome is unknown, so it is never compensated from a +# read that could come before the rest of it. The next activation repairs +# from the target state. +my sub nvmet_exec($scfg, $steps, %opts) { + my $locked = ($nvmet_lock_owner{ nvmet_lock_id($scfg) } // 0) == $$; + die "internal error: NVMe target change without target lock\n" + if $opts{change} && !$locked; + my $res = _nvmet_run( + $scfg, $steps, + op => $opts{op}, + timeout => $opts{timeout} // 15, + (defined($opts{input}) ? (input => $opts{input}) : ()), + ); + die "NVMe target operation '$opts{op}' did not complete ($res->{err});" + . " the target state is unknown\n" + if ($locked && $res->{rc} == -1) + || ($opts{change} && $res->{rc} == 255); + return $res; +} + +# An abandoned operation is passed on instead of being compensated. +my sub nvmet_rethrow_abandoned($error) { + die $error if $error =~ $RE_NVMET_ABANDONED; +} + +my sub nvmet_unreachable($res) { + return $res->{rc} == 255 || $res->{rc} == -1; +} + +my sub nvmet_check($res, $op) { + die "NVMe target operation '$op' failed: $res->{err}\n" if $res->{rc}; +} + +my sub nvmet_read_failed($scfg, $res) { + die "NVMe target '" . nvmet_server($scfg) . "' is unreachable: $res->{err}\n" + if nvmet_unreachable($res); + die "cannot read NVMe target state: $res->{err}\n"; +} + +my sub nvmet_output($res) { + return join('', map { "$_\n" } $res->{out}->@*); +} + +# A read, retried twice after 200 ms unless the target is unreachable: a read +# without the lock can meet a configfs object that another node removes +# while find lists it, which makes find fail once. +my sub nvmet_read_call($scfg, $steps, %opts) { + my $res; + for my $attempt (0 .. 2) { + _sleep(0.2) if $attempt; + $res = nvmet_exec($scfg, $steps, %opts, timeout => 15); + last if !$res->{rc} || nvmet_unreachable($res); + } + return $res; +} + +# Reads the target state and returns (ZFS inventory, configfs state). +# keys => [hostnqns] adds the digests of their keys. modprobe => 1, for locked +# reads only, loads nvmet first and mounts configfs if that is what failed. +# With tolerant => 1, a failed or unparsable read returns an empty list, +# unless the target is unreachable. +my sub nvmet_read($scfg, %opts) { + my $pool = nvmet_pool($scfg); + my $hostnqns = $opts{keys} // []; + # loading a module that may still be loading is harmless + my @modprobe = $opts{modprobe} ? (change => 1) : (); + my $read = sub () { + return nvmet_read_call( + $scfg, + nvmet_state_steps($pool, $hostnqns, $opts{modprobe}), + op => 'read target state', + @modprobe, + ); + }; + my $res = $read->(); + if ($res->{rc} && $opts{modprobe} && !nvmet_unreachable($res)) { + my $mounts = nvmet_exec($scfg, [['cat', '/proc/mounts']], op => 'read mounts'); + nvmet_check($mounts, 'read mounts'); + if (!_nvmet_parse_mounts(nvmet_output($mounts))) { + my $mount = ['mount', '-t', 'configfs', 'none', '/sys/kernel/config']; + my $mounted = nvmet_exec($scfg, [$mount], op => 'mount configfs', change => 1); + nvmet_check($mounted, 'mount configfs'); + $res = $read->(); + } + } + if ($res->{rc}) { + return () if $opts{tolerant} && !nvmet_unreachable($res); + nvmet_read_failed($scfg, $res); + } + my @state = eval { + my ($zfs, $configfs) = _nvmet_split_state(nvmet_output($res)); + (_nvmet_parse_zfs_inventory($zfs, $pool), _nvmet_parse_configfs($configfs)); + }; + if (my $error = $@) { + return () if $opts{tolerant}; + die $error; + } + return @state; +} + +my sub nvmet_read_zfs($scfg) { + my $pool = nvmet_pool($scfg); + my $res = nvmet_read_call($scfg, [nvmet_zfs_inventory_step($pool)], op => 'read ZFS inventory'); + nvmet_read_failed($scfg, $res) if $res->{rc}; + return _nvmet_parse_zfs_inventory(nvmet_output($res), $pool); +} + +# Applies units in as few calls as possible and stops at the first failing +# call, whose result it returns. A chain only saves SSH round trips: every +# decision was taken locally before, and `&&` only stops at the first +# failing command. Only the call with the key step gets stdin. The default +# timeout suits configfs chains; ZFS changes pass their own. +my sub nvmet_apply($scfg, $units, %opts) { + my $input = delete $opts{input}; + $opts{timeout} //= 30; + for my $chunk (_nvmet_chunk($units)->@*) { + my $with_key = grep { ref($_) eq 'HASH' && exists($_->{key}) } $chunk->@*; + my @input = $with_key ? (input => $input) : (); + my $res = nvmet_exec($scfg, $chunk, %opts, @input, change => 1); + return $res if $res->{rc}; + } + return { rc => 0, out => [], err => '' }; +} + +my sub nvmet_change($scfg, $units, %opts) { + nvmet_check(nvmet_apply($scfg, $units, %opts), $opts{op}); +} + +my sub nvmet_wait_device($scfg, $dev) { + my $deadline = _now() + 10; + while (1) { + my $res = nvmet_exec($scfg, [['test', '-b', $dev]], op => 'wait for zvol'); + return if !$res->{rc}; + nvmet_read_failed($scfg, $res) if nvmet_unreachable($res); + last if _now() >= $deadline; + _sleep(0.25); + } + die "zvol '$dev' is not a block device after 10 seconds\n"; +} + +my sub nvmet_new_uuid($inv, $cfs) { + my %used = map { lc($_->{uuid}) => 1 } values $inv->{volumes}->%*; + for my $subsys (values $cfs->{subsystems}->%*) { + $used{ lc($_->{device_uuid}) } = 1 for values(($subsys->{namespaces} // {})->%*); + } + for (1 .. 10) { + my $uuid = lc(file_read_firstline('/proc/sys/kernel/random/uuid') // ''); + return $uuid if $uuid =~ $RE_NVMET_UUID && !$used{$uuid}; + } + die "cannot generate an unused namespace UUID\n"; +} + +my sub nvmet_has_identity($row, $nqn, $nsid, $uuid) { + return $row && $row->{nqn} eq $nqn && $row->{nsid} eq "$nsid" && $row->{uuid} eq $uuid; +} + +# The configfs state after a successful unexport of namespace $pre->{nsid}. +my sub nvmet_without_namespace($cfs, $nqn, $pre) { + return $cfs if !$pre; + my $subsys = $cfs->{subsystems}->{$nqn}; + my %namespaces = $subsys->{namespaces}->%*; + delete $namespaces{ $pre->{nsid} }; + my $changed = { $subsys->%*, namespaces => \%namespaces }; + return { $cfs->%*, subsystems => { $cfs->{subsystems}->%*, $nqn => $changed } }; +} + +# Links the subsystem to each desired port separately. One port that cannot +# be bound (for example, its address is missing on the target) only warns +# while another desired port serves the subsystem. +my sub nvmet_publish($scfg, $cfs, $nqn, $port_ids, $hostnqns) { + my @failed; + for my $unit (_nvmet_plan_publish($cfs, $nqn, $port_ids, $hostnqns)->@*) { + my $op = "$unit->{action} NVMe subsystem on port $unit->{port}"; + my $res = nvmet_apply($scfg, [$unit->{steps}], op => $op); + push @failed, { $unit->%*, err => $res->{err} } if $res->{rc}; + } + return if !@failed; + + my (undef, $fresh) = eval { nvmet_read($scfg) }; + nvmet_rethrow_abandoned($@) if $@; + my $published = $fresh && grep { nvmet_port_linked($fresh, $_, $nqn) } $port_ids->@*; + my ($first) = grep { $_->{action} eq 'link' } @failed; + $first //= $failed[0]; + die "cannot publish NVMe subsystem '$nqn': $first->{err}\n" if !$published; + for my $failure (@failed) { + my $what = $failure->{action} eq 'link' ? 'publish' : 'unpublish'; + log_warn("cannot $what NVMe subsystem '$nqn' on port $failure->{port}: $failure->{err}"); + } +} + +# --------------------------------------------------------------------------- +# Target flows +# --------------------------------------------------------------------------- + +# Brings the target to the configured state and publishes the subsystem. +# A converged target costs one read and no lock. Otherwise, under the lock: +# read, apply, verify with a fresh read, and publish only once the verify +# read plans nothing more. $portals and $hostnqns are the parsed and validated +# storage configuration. +sub _nvmet_activate_target($storeid, $scfg, $portals, $hostnqns, $key) { + my $nqn = nvmet_nqn($scfg); + my $conf = { + nqn => $nqn, + portals => $portals, + hostnqns => $hostnqns, + keysha => sha256_hex("$key\n"), + }; + my $converged = sub($inv, $cfs) { + my $plan = _nvmet_plan_activation($inv, $cfs, $conf); + return !$plan->{prepublish}->@* + && !_nvmet_plan_publish($cfs, $nqn, $plan->{port_ids}, $hostnqns)->@*; + }; + + # Without the lock, a refusal or failed read only means "take the lock". + my @state = nvmet_read($scfg, keys => $hostnqns, tolerant => 1); + return if @state && eval { $converged->(@state) }; + + nvmet_locked( + $scfg, + sub { + # Removing the storage deletes its key before it removes the target + # configuration, so an activation that waited for the lock meanwhile + # does not restore it. + die "storage '$storeid' is being removed or its key changed\n" + if (file_read_firstline(secret_path($storeid)) // '') ne $key; + my ($inv, $cfs) = nvmet_read($scfg, keys => $hostnqns, modprobe => 1); + my $plan = _nvmet_plan_activation($inv, $cfs, $conf); + my $error; + for my $round (1 .. 2) { + last if !$plan->{prepublish}->@*; + my $apply = nvmet_apply( + $scfg, $plan->{prepublish}, + op => 'configure NVMe target', + input => "$key\n", + ); + $error //= $apply->{err} if $apply->{rc}; + ($inv, $cfs) = nvmet_read($scfg, keys => $hostnqns); + $plan = _nvmet_plan_activation($inv, $cfs, $conf); + } + die "NVMe target did not converge: " . ($error // 'state mismatch') . "\n" + if $plan->{prepublish}->@*; + nvmet_publish($scfg, $cfs, $nqn, $plan->{port_ids}, $hostnqns); + }, + ); + return; +} + +# The namespace UUID of an owned volume, from one read-only ZFS query. +sub _nvmet_volume_uuid($scfg, $name) { + my $dataset = nvmet_dataset($scfg, $name); + my $inv = nvmet_read_zfs($scfg); + my (undef, $uuid) = _nvmet_owned_identity($inv, nvmet_nqn($scfg), $dataset); + return $uuid; +} + +# Creates (size => KiB) or clones (origin => pool/base@snap) a zvol with its +# identity set at creation, then exports it. A failure removes only what this +# call created, decided from a fresh read. +my sub nvmet_create_volume($scfg, $name, %opts) { + my $nqn = nvmet_nqn($scfg); + my $dataset = nvmet_dataset($scfg, $name); + my $dev = "/dev/zvol/$dataset"; + die "internal error: invalid zvol creation request\n" + if defined($opts{size}) == defined($opts{origin}) + || (defined($opts{size}) && $opts{size} !~ $RE_UNSIGNED_INTEGER); + + return nvmet_locked( + $scfg, + sub { + my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1); + die "zvol '$dataset' already exists\n" if exists($inv->{types}->{$dataset}); + die "NVMe subsystem does not exist; activate the storage first\n" + if !$cfs->{subsystems}->{$nqn}; + my $nsid = _nvmet_allocate_nsid($inv, $cfs, $nqn); + my $uuid = nvmet_new_uuid($inv, $cfs); + my (undef, $state) = _nvmet_plan_export($cfs, $nqn, $nsid, $uuid, $dev); + die "internal error: namespace ID '$nsid' is in use\n" if $state ne 'absent'; + + # Record the NSID on the pool before the zvol exists: a creation that + # completes on the target only after this call gave up on it must + # never share its NSID with a later volume. + my $timeout = nvmet_long_timeout(); + if ($nsid > $inv->{last_nsid}) { + my $reserve = ['zfs', 'set', "$nvmet_last_nsid=$nsid", nvmet_pool($scfg)]; + nvmet_change( + $scfg, [[$reserve]], + op => 'reserve namespace ID', + timeout => $timeout, + ); + } + + my @identity = map { ('-o', $_) } nvmet_identity($nqn, $nsid, $uuid); + my @create_opts = ( + ($scfg->{sparse} ? ('-s') : ()), + ($scfg->{blocksize} ? ('-b', $scfg->{blocksize}) : ()), + ); + my $create = + defined($opts{origin}) + ? ['zfs', 'clone', @identity, $opts{origin}, $dataset] + : ['zfs', 'create', @create_opts, @identity, '-V', "$opts{size}k", $dataset]; + my $res = nvmet_apply( + $scfg, [[$create]], + op => 'create zvol', + timeout => $timeout, + ); + if ($res->{rc}) { + # A definite command failure may still have created the zvol. + # Only a zvol with this call's identity counts. Ambiguous SSH + # outcomes have already aborted the transaction in nvmet_exec. + my $check = eval { nvmet_read_zfs($scfg) }; + nvmet_rethrow_abandoned($@) if $@; + my $row = $check ? $check->{volumes}->{$dataset} : undef; + die "NVMe target operation 'create zvol' failed: $res->{err}\n" + if !nvmet_has_identity($row, $nqn, $nsid, $uuid); + } + + my $verified; + eval { + nvmet_wait_device($scfg, $dev); + my $build_unit = nvmet_build_unit($nqn, $nsid, $uuid, $dev); + my $build = nvmet_apply($scfg, [$build_unit], op => 'export namespace'); + $verified = [nvmet_read($scfg)]; + my ($vinv, $vcfs) = $verified->@*; + my ($vnsid, $vuuid) = _nvmet_owned_identity($vinv, $nqn, $dataset); + die "NVMe identity of zvol '$dataset' changed\n" + if $vnsid ne $nsid || $vuuid ne $uuid; + if (_nvmet_export_state($vcfs, $nqn, $nsid, $uuid, $dev) ne 'present') { + nvmet_check($build, 'export namespace'); + die "NVMe namespace '$nsid' was not exported\n"; + } + }; + my $error = $@ or return $uuid; + nvmet_rethrow_abandoned($error); + + eval { + my ($cinv, $ccfs) = $verified ? $verified->@* : nvmet_read($scfg); + my $namespaces = nvmet_namespaces($ccfs, $nqn); + # The build writes the UUID before it enables the namespace, so an + # enabled namespace without our UUID was not built by this call. + my @unexport = + map { nvmet_unexport_unit($nqn, $_, $namespaces->{$_}->{enable} eq '1') } + grep { + my $ns = $namespaces->{$_}; + $ns->{device_uuid} eq $uuid + || ($ns->{enable} ne '1' && $ns->{device_path} eq $dev) + } sort { + $a <=> $b + } keys $namespaces->%*; + nvmet_change($scfg, \@unexport, op => 'remove namespace') if @unexport; + # -o at creation proves that this call created the zvol + if (nvmet_has_identity($cinv->{volumes}->{$dataset}, $nqn, $nsid, $uuid)) { + my $destroy = ['zfs', 'destroy', '-r', $dataset]; + nvmet_change( + $scfg, [[$destroy]], + op => 'destroy zvol', + timeout => $timeout, + ); + } + }; + log_warn("failed to clean up zvol '$name': $@") if $@; + die $error; + }, + ); +} + +# Unexports and destroys an owned zvol. A missing zvol is already destroyed. +my sub nvmet_destroy_volume($scfg, $name) { + my $nqn = nvmet_nqn($scfg); + my $dataset = nvmet_dataset($scfg, $name); + my $dev = "/dev/zvol/$dataset"; + my $destroy = ['zfs', 'destroy', '-r', $dataset]; + my $timeout = nvmet_long_timeout(); + + nvmet_locked( + $scfg, + sub { + my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1); + return if !exists($inv->{types}->{$dataset}); + my ($nsid, $uuid) = _nvmet_owned_identity($inv, $nqn, $dataset); + my ($unexport, $pre) = ([], undef); + if ($cfs->{subsystems}->{$nqn}) { + _nvmet_plan_export($cfs, $nqn, $nsid, $uuid, $dev); # refusals only + ($unexport, $pre) = _nvmet_plan_unexport($cfs, $nqn, $uuid); + } + my $units = [$unexport->@*, [$destroy]]; + my $res = nvmet_apply($scfg, $units, op => 'destroy zvol', timeout => $timeout); + return if !$res->{rc}; + my $error = $res->{err}; + + # Decide from a fresh read what the failed call changed. + my ($finv, $fcfs) = eval { nvmet_read($scfg) }; + if (my $read_error = $@) { + nvmet_rethrow_abandoned($read_error); + die "failed to destroy '$dataset': $error\n"; + } + return if !exists($finv->{types}->{$dataset}); + my $namespaces = nvmet_namespaces($fcfs, $nqn); + my ($current) = + grep { $namespaces->{$_}->{device_uuid} eq $uuid } keys $namespaces->%*; + if (defined($current)) { + die "cannot disable namespace '$current': $error\n" + if $namespaces->{$current}->{enable} eq '1'; + die "cannot remove namespace '$current': $error\n" if !$pre || !$pre->{enabled}; + my $enable = [nvmet_enable_unit($nqn, $current, $dev)]; + my $restore = nvmet_apply($scfg, $enable, op => 'enable namespace'); + die "cannot remove namespace '$current': $error" + . ($restore->{rc} ? "; cannot restore namespace: $restore->{err}" : '') . "\n"; + } + + # The namespace is gone and the zvol is still there, typically busy. + for my $attempt (1 .. 5) { + _sleep(1); + my $retry = + nvmet_apply($scfg, [[$destroy]], op => 'destroy zvol', timeout => $timeout); + return if !$retry->{rc}; + $error = $retry->{err}; + } + die "failed to destroy '$dataset': $error\n" if !$pre; + eval { + my (undef, $rcfs) = nvmet_read($scfg); + my ($units) = _nvmet_plan_export($rcfs, $nqn, $nsid, $uuid, $dev); + nvmet_change($scfg, $units, op => 'restore namespace'); + }; + die "failed to destroy '$dataset': $error; namespace restoration also failed: $@" + if $@; + die "failed to destroy '$dataset': $error; namespace restored\n"; + }, + ); + return; +} + +# Rolls an owned zvol back and always exports it again with the same identity. +my sub nvmet_rollback_volume($scfg, $name, $snap) { + my $snapshot = nvmet_snapshot($scfg, $name, $snap); + my $nqn = nvmet_nqn($scfg); + my $dataset = nvmet_dataset($scfg, $name); + my $dev = "/dev/zvol/$dataset"; + + my $timeout = nvmet_long_timeout(); + + nvmet_locked( + $scfg, + sub { + my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1); + my ($nsid, $uuid) = _nvmet_owned_identity($inv, $nqn, $dataset); + die "NVMe subsystem does not exist\n" if !$cfs->{subsystems}->{$nqn}; + _nvmet_plan_export($cfs, $nqn, $nsid, $uuid, $dev); # refusals only + my ($unexport, $pre) = _nvmet_plan_unexport($cfs, $nqn, $uuid); + my $set_identity = ['zfs', 'set', nvmet_identity($nqn, $nsid, $uuid), $dataset]; + my $rollback = ['zfs', 'rollback', $snapshot]; + my $res = nvmet_apply( + $scfg, [$unexport->@*, [$rollback], [$set_identity]], + op => 'rollback zvol', + timeout => $timeout, + ); + my $error; + $error = "NVMe target operation 'rollback zvol' failed: $res->{err}" if $res->{rc}; + + # Always export the volume again, from a fresh read if anything failed. + eval { + my ($rinv, $rcfs) = + $error ? nvmet_read($scfg) : ($inv, nvmet_without_namespace($cfs, $nqn, $pre)); + nvmet_wait_device($scfg, $dev); + my @units; + push @units, [$set_identity] + if !nvmet_has_identity($rinv->{volumes}->{$dataset}, $nqn, $nsid, $uuid); + my ($export) = _nvmet_plan_export($rcfs, $nqn, $nsid, $uuid, $dev); + push @units, $export->@*; + nvmet_change($scfg, \@units, op => 'restore namespace', timeout => $timeout); + }; + chomp(my $restore_error = $@); + + die "rollback failed: $error; namespace restoration failed: $restore_error\n" + if defined($error) && $restore_error; + die "namespace restoration failed after rollback: $restore_error\n" + if $restore_error; + die "$error\n" if defined($error); + }, + ); + return; +} + +# Renames vm-* to base-*, exports it under the same identity and takes the +# __base__ snapshot. A failed conversion is undone from a fresh read. +my sub nvmet_template_volume($scfg, $name) { + my $nqn = nvmet_nqn($scfg); + my $old = nvmet_dataset($scfg, $name); + my $new = _nvmet_template_name($old); + my ($old_dev, $new_dev) = ("/dev/zvol/$old", "/dev/zvol/$new"); + my $timeout = nvmet_long_timeout(); + + nvmet_locked( + $scfg, + sub { + my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1); + my ($nsid, $uuid) = _nvmet_owned_identity($inv, $nqn, $old); + die "template zvol '$new' already exists\n" if exists($inv->{types}->{$new}); + die "NVMe subsystem does not exist\n" if !$cfs->{subsystems}->{$nqn}; + _nvmet_plan_export($cfs, $nqn, $nsid, $uuid, $old_dev); # refusals only + my ($unexport) = _nvmet_plan_unexport($cfs, $nqn, $uuid); + + eval { + my $rename = [$unexport->@*, [['zfs', 'rename', $old, $new]]]; + nvmet_change($scfg, $rename, op => 'rename zvol', timeout => $timeout); + nvmet_wait_device($scfg, $new_dev); + my $build = nvmet_build_unit($nqn, $nsid, $uuid, $new_dev); + my $template = [$build, [['zfs', 'snapshot', "$new\@__base__"]]]; + nvmet_change($scfg, $template, op => 'create template', timeout => $timeout); + }; + my $error = $@ or return; + nvmet_rethrow_abandoned($error); + chomp($error); + + eval { + my ($rinv, $rcfs) = nvmet_read($scfg); + if (exists($rinv->{types}->{$new})) { + my ($units) = _nvmet_plan_unexport($rcfs, $nqn, $uuid); + my $back = [$units->@*, [['zfs', 'rename', $new, $old]]]; + nvmet_change($scfg, $back, op => 'rename zvol back', timeout => $timeout); + (undef, $rcfs) = nvmet_read($scfg); + } elsif (!exists($rinv->{types}->{$old})) { + die "zvol '$old' does not exist\n"; + } + nvmet_wait_device($scfg, $old_dev); + my ($units) = _nvmet_plan_export($rcfs, $nqn, $nsid, $uuid, $old_dev); + nvmet_change($scfg, $units, op => 'restore namespace'); + }; + die "template conversion failed: $error; restoration failed: $@" if $@; + die "template conversion failed: $error\n"; + }, + ); + return; +} + +# Grows an owned zvol and lets an exported namespace pick up the new size. +my sub nvmet_resize_volume($scfg, $name, $size) { + my $nqn = nvmet_nqn($scfg); + my $dataset = nvmet_dataset($scfg, $name); + die "internal error: invalid zvol size\n" if $size !~ $RE_UNSIGNED_INTEGER; + + # The volume is checked before it changes; under the lock, the namespace + # read here is still the one to revalidate after the resize. It is + # revalidated in the same call, so whenever the zvol grows, also after a + # lost reply, the hosts see the new size. A namespace that is not enabled + # reads the new size when it is enabled. + nvmet_locked( + $scfg, + sub { + my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1); + my (undef, $uuid) = _nvmet_owned_identity($inv, $nqn, $dataset); + my $namespaces = nvmet_namespaces($cfs, $nqn); + my @matches = + grep { $namespaces->{$_}->{device_uuid} eq $uuid } keys $namespaces->%*; + die "duplicate namespace UUID '$uuid'\n" if @matches > 1; + my @resize = (['zfs', 'set', "volsize=${size}k", $dataset]); + push @resize, nvmet_write(nvmet_ns_path($nqn, $matches[0]) . '/revalidate_size', 1) + if @matches && $namespaces->{ $matches[0] }->{enable} eq '1'; + nvmet_change( + $scfg, [\@resize], + op => 'resize zvol', + timeout => nvmet_long_timeout(), + ); + }, + ); + return; +} + +# Removes the subsystem once it has no volumes and namespaces, then the hosts +# that no subsystem links any more. Ports are never removed. +sub _nvmet_delete_target($scfg) { + my $nqn = nvmet_nqn($scfg); + my $hostnqns = parse_nvme_host_nqns($scfg->{'nvme-host-nqns'}, 1) // []; + + nvmet_locked( + $scfg, + sub { + my ($inv, $cfs) = nvmet_read($scfg, modprobe => 1); + my ($units, $candidates) = _nvmet_plan_delete_target($inv, $cfs, $nqn, $hostnqns); + my $fresh = $cfs; + if ($units->@*) { + my $res = nvmet_apply($scfg, $units, op => 'remove NVMe subsystem'); + (undef, $fresh) = eval { nvmet_read($scfg) }; + my $read_error = $@; + nvmet_rethrow_abandoned($read_error) if $read_error; + die "cannot remove NVMe subsystem '$nqn': $res->{err}\n" + if $res->{rc} && (!$fresh || $fresh->{subsystems}->{$nqn}); + die $read_error if !$fresh; + } + my $orphans = _nvmet_plan_orphan_hosts($fresh, $candidates); + return if !$orphans->@*; + my $res = nvmet_apply($scfg, $orphans, op => 'remove orphan NVMe hosts'); + log_warn("could not remove orphan NVMe host(s): $res->{err}") if $res->{rc}; + }, + ); + return; +} + +# --------------------------------------------------------------------------- +# Configuration +# --------------------------------------------------------------------------- + +sub verify_nvme_nqn($value, $noerr = undef) { + + if ( + length($value) > 223 + || $value !~ $RE_NQN + ) { + return undef if $noerr; + die "value is not a valid NVMe qualified name\n"; + } + + return $value; +} + +sub parse_nvme_portals($value, $noerr = undef) { + my $result = []; + my $seen = {}; + + for my $entry (split(/,/, $value // '')) { + $entry = trim($entry); + my ($address, $port, $family); + if ($entry =~ $RE_IPV6_PORTAL) { + ($address, $port, $family) = ($+{address}, $+{port} // 4420, 'ipv6'); + } elsif ($entry =~ $RE_IPV4_PORTAL) { + ($address, $port, $family) = ($+{address}, $+{port} // 4420, 'ipv4'); + } else { + return undef if $noerr; + die "invalid NVMe/TCP portal '$entry'\n"; + } + + if ( + !PVE::JSONSchema::pve_verify_ip($address, 1) + || ($family eq 'ipv4' && index($address, ':') >= 0) + || ($family eq 'ipv6' && index($address, ':') < 0) + || $port < 1 + || $port > 65535 + ) { + return undef if $noerr; + die "invalid NVMe/TCP portal '$entry'\n"; + } + $address = nvmet_canonical_address($address) if $family eq 'ipv6'; + # one listener: compare like _assert_unique_target + my $id = join( + "\0", + nvmet_listener_address({ family => $family, address => $address }), + int($port), + ); + if ($seen->{$id}++) { + return undef if $noerr; + die "duplicate NVMe/TCP portal '$entry'\n"; + } + push $result->@*, + { + address => $address, + port => int($port), + family => $family, + }; + if (scalar($result->@*) > $max_paths) { + return undef if $noerr; + die "at most $max_paths NVMe/TCP portals are supported\n"; + } + } + + if (!$result->@*) { + return undef if $noerr; + die "at least one NVMe/TCP portal is required\n"; + } + + return $result; +} + +my sub verify_nvme_portals($value, $noerr = undef) { + return undef if !parse_nvme_portals($value, $noerr); + return $value; +} + +sub parse_nvme_host_ifaces($value, $noerr = undef) { + my $result = []; + + for my $iface (split(/,/, $value // '')) { + $iface = trim($iface); + # the kernel never names an interface '.' or '..' + if ( + length($iface) < 1 + || length($iface) > 15 + || $iface !~ $RE_HOST_IFACE + || $iface eq '.' + || $iface eq '..' + ) { + return undef if $noerr; + die "invalid NVMe/TCP host interface '$iface'\n"; + } + push $result->@*, $iface; + } + + if (!$result->@*) { + return undef if $noerr; + die "at least one NVMe/TCP host interface is required\n"; + } + + return $result; +} + +sub parse_nvme_host_nqns($value, $noerr = undef) { + my $result = []; + my $seen = {}; + + for my $hostnqn (split(/,/, $value // '')) { + $hostnqn = trim($hostnqn); + if (!verify_nvme_nqn($hostnqn, 1) || $seen->{$hostnqn}++) { + return undef if $noerr; + die "invalid or duplicate NVMe host NQN '$hostnqn'\n"; + } + push $result->@*, $hostnqn; + if (scalar($result->@*) > $max_hosts) { + return undef if $noerr; + die "at most $max_hosts NVMe host NQNs are supported\n"; + } + } + + if (!$result->@*) { + return undef if $noerr; + die "at least one NVMe host NQN is required\n"; + } + + return $result; +} + +my sub verify_nvme_host_ifaces($value, $noerr = undef) { + return undef if !parse_nvme_host_ifaces($value, $noerr); + return $value; +} + +my sub verify_nvme_host_nqns($value, $noerr = undef) { + return undef if !parse_nvme_host_nqns($value, $noerr); + return $value; +} + +sub _configured_portals($scfg) { + my $portals = parse_nvme_portals($scfg->{'nvme-portals'}); + my $ifaces = parse_nvme_host_ifaces($scfg->{'nvme-host-ifaces'}); + die "nvme-host-ifaces must contain one interface for each nvme-portals entry\n" + if scalar($ifaces->@*) != scalar($portals->@*); + + for (my $i = 0; $i < scalar($portals->@*); $i++) { + $portals->[$i]->{host_iface} = $ifaces->[$i]; + } + return $portals; +} + +sub _local_iface_exists($iface) { + return -d "/sys/class/net/$iface"; +} + +sub _validate_local_ifaces($portals) { + for my $portal ($portals->@*) { + my $iface = $portal->{host_iface}; + die "NVMe/TCP host interface '$iface' does not exist on this node\n" + if !_local_iface_exists($iface); + } +} + +PVE::JSONSchema::register_format('pve-storage-nvme-nqn', \&verify_nvme_nqn); +PVE::JSONSchema::register_format('pve-storage-nvme-portals', \&verify_nvme_portals); +PVE::JSONSchema::register_format('pve-storage-nvme-host-ifaces', \&verify_nvme_host_ifaces); +PVE::JSONSchema::register_format('pve-storage-nvme-host-nqns', \&verify_nvme_host_nqns); + +sub type($class) { + return 'zfsnvme'; +} + +sub plugindata($class) { + return { + content => [{ images => 1 }, { images => 1 }], + 'sensitive-properties' => { 'dhchap-key' => 1 }, + }; +} + +sub properties($class) { + return { + subsysnqn => { + description => "NVMe subsystem qualified name.", + type => 'string', + format => 'pve-storage-nvme-nqn', + }, + 'nvme-portals' => { + description => + "Comma-separated NVMe/TCP target IP addresses, at most 16. Put IPv6 addresses in" + . " brackets. The default TCP port is 4420.", + type => 'string', + format => 'pve-storage-nvme-portals', + maxLength => 2048, + }, + 'nvme-host-ifaces' => { + description => + "Comma-separated local interfaces, in portal order. Interface names must be" + . " identical on every cluster node.", + type => 'string', + format => 'pve-storage-nvme-host-ifaces', + maxLength => 512, + }, + 'nvme-host-nqns' => { + description => + "Comma-separated /etc/nvme/hostnqn values for every cluster node allowed to use" + . " this storage, at most 64. Host NQNs can only be added.", + type => 'string', + format => 'pve-storage-nvme-host-nqns', + maxLength => 8192, + }, + 'dhchap-key' => { + description => + "NVMe DH-HMAC-CHAP key in secret representation format (DHHC-1). Required when the" + . " storage is created; it cannot be changed.", + type => 'string', + maxLength => 256, + }, + 'nvme-iopolicy' => { + description => "Native NVMe multipath I/O policy.", + type => 'string', + enum => ['numa', 'round-robin', 'queue-depth'], + default => 'round-robin', + }, + 'nvme-keep-alive-tmo' => { + description => "NVMe keep-alive timeout in seconds.", + type => 'integer', + minimum => 1, + maximum => 120, + default => 5, + }, + 'nvme-reconnect-delay' => { + description => "Delay between NVMe reconnect attempts in seconds.", + type => 'integer', + minimum => 1, + maximum => 120, + default => 2, + }, + 'nvme-ctrl-loss-tmo' => { + description => "Time to keep retrying a lost NVMe controller in seconds.", + type => 'integer', + minimum => -1, + maximum => 86400, + default => 600, + }, + 'nvme-fast-io-fail-tmo' => { + description => + "Optional time before failing I/O on a reconnecting NVMe controller. Unset queues" + . " I/O until controller loss timeout.", + type => 'integer', + minimum => 0, + maximum => 86400, + optional => 1, + }, + 'nvme-nr-io-queues' => { + description => "Number of NVMe/TCP I/O queues per controller.", + type => 'integer', + minimum => 1, + maximum => 1024, + optional => 1, + }, + }; +} + +sub options($class) { + return { + server => { fixed => 1 }, + subsysnqn => { fixed => 1 }, + 'nvme-portals' => { fixed => 1 }, + 'nvme-host-ifaces' => {}, + 'nvme-host-nqns' => {}, + pool => { fixed => 1 }, + blocksize => { fixed => 1 }, + sparse => { optional => 1 }, + 'dhchap-key' => { optional => 1 }, + 'nvme-iopolicy' => { optional => 1 }, + 'nvme-keep-alive-tmo' => { optional => 1 }, + 'nvme-reconnect-delay' => { optional => 1 }, + 'nvme-ctrl-loss-tmo' => { optional => 1 }, + 'nvme-fast-io-fail-tmo' => { optional => 1 }, + 'nvme-nr-io-queues' => { optional => 1 }, + nodes => { optional => 1 }, + disable => { optional => 1 }, + content => { optional => 1 }, + bwlimit => { optional => 1 }, + }; +} + +sub _validate_fail_fast_timeout($config, $default_ctrl_loss_tmo = undef) { + my $fast = $config->{'nvme-fast-io-fail-tmo'}; + return if !defined($fast); + + my $ctrl = $config->{'nvme-ctrl-loss-tmo'}; + $ctrl = $default_ctrl_loss_tmo if !defined($ctrl); + return if !defined($ctrl) || $ctrl < 0; + + die "nvme-fast-io-fail-tmo must not exceed nvme-ctrl-loss-tmo\n" + if $fast > $ctrl; +} + +# Defaults are not injected here: this also runs when storage.cfg is parsed. +# Every use site applies the schema default itself. +sub check_config($class, $section_id, $config, $create, $skip_schema_check = undef) { + nvmet_pool($config) if defined($config->{pool}); + verify_nvme_nqn($config->{subsysnqn}) if defined($config->{subsysnqn}); + parse_nvme_portals($config->{'nvme-portals'}) if defined($config->{'nvme-portals'}); + if (defined($config->{'nvme-host-ifaces'}) && defined($config->{'nvme-portals'})) { + _configured_portals($config); + } elsif (defined($config->{'nvme-host-ifaces'})) { + parse_nvme_host_ifaces($config->{'nvme-host-ifaces'}); + } + parse_nvme_host_nqns($config->{'nvme-host-nqns'}) + if defined($config->{'nvme-host-nqns'}); + _validate_fail_fast_timeout($config, $create ? 600 : undef); + return $class->SUPER::check_config($section_id, $config, $create, $skip_schema_check); +} + +# --------------------------------------------------------------------------- +# ZFS helpers (same formats as the other ZFS plugins) +# --------------------------------------------------------------------------- + +my sub zfs_request($scfg, $timeout, $method, @args) { + $timeout = PVE::RPCEnvironment->is_worker() ? 60 * 60 : 10 if !$timeout; + my $change = $method ne 'get' && $method ne 'list'; + my $res = nvmet_exec( + $scfg, [['zfs', $method, @args]], + op => "zfs $method", + timeout => $timeout, + change => $change, + ); + die "zfs error on '" . nvmet_server($scfg) . "': $res->{err}\n" if $res->{rc}; + return nvmet_output($res); +} + +# The direct children of the pool with PVE volume names. The plugin creates +# and owns zvols only; a filesystem child holds its name. +my sub zfs_parse_zvol_list($text, $pool) { + my $list = []; + for my $line (split /\n/, $text // '') { + my ($dataset, $size, $origin, $type) = split(/\s+/, $line); + next if !defined($type) || ($type ne 'volume' && $type ne 'filesystem'); + my @parts = split /\//, $dataset; + next if @parts < 2; + my $name = pop @parts; + next if join('/', @parts) ne $pool; + next if $name !~ $RE_ZVOL_OWNER; + push $list->@*, + { + name => $name, + owner => $+{owner}, + type => $type, + ($type eq 'volume' ? (size => $size + 0) : ()), + ($origin ne '-' ? (origin => $origin) : ()), + }; + } + return $list; +} + +# The pool's direct children with PVE volume names, so that a name is never +# handed out twice; with $owned_only, the zvols owned by this storage's NVMe +# subsystem. +my sub zfs_list_zvol($scfg, $owned_only) { + my $pool = nvmet_pool($scfg); + my $text = zfs_request( + $scfg, + 10, + 'list', + '-o', + 'name,volsize,origin,type', + '-t', + 'volume,filesystem', + '-d1', + '-Hp', + $pool, + ); + my $list = {}; + my $prefix = "$pool/"; + for my $zvol (zfs_parse_zvol_list($text, $pool)->@*) { + next if $owned_only && $zvol->{type} ne 'volume'; + my $parent = $zvol->{origin}; + # an origin in the pool is named relative to it, like the ZFS plugins do + $parent = substr($parent, length($prefix)) if $parent && index($parent, $prefix) == 0; + $list->{ $zvol->{name} } = { + name => $zvol->{name}, + size => $zvol->{size}, + parent => $parent, + format => 'raw', + vmid => $zvol->{owner}, + }; + } + return $list if !$owned_only; + + my $properties = zfs_request( + $scfg, + 10, + 'get', + '-H', + '-d', + '1', + '-o', + 'name,value,source', + 'proxmox:nvme-subsys', + $pool, + ); + my $owned = {}; + for my $line (split(/\n/, $properties)) { + my ($dataset, $nqn, $source) = split(/\t/, $line, 3); + next if !defined($source) || ($source ne 'local' && $source ne 'received'); + next if !defined($nqn) || $nqn ne $scfg->{subsysnqn}; + next if index($dataset, $prefix) != 0; + $owned->{ substr($dataset, length($prefix)) } = 1; + } + for my $name (keys $list->%*) { + delete $list->{$name} if !$owned->{$name}; + } + return $list; +} + +my sub zfs_get_properties($scfg, $properties, $dataset, $timeout = undef) { + my $text = zfs_request($scfg, $timeout, 'get', '-o', 'value', '-Hp', $properties, $dataset); + my @values = split /\n/, $text; + return wantarray ? @values : $values[0]; +} + +my sub zfs_get_sorted_snapshot_list($scfg, $name, $sort_params) { + return [ + map { s/^.*\@//r } split /\n/, + zfs_request( + $scfg, + undef, + 'list', + '-H', + '-r', + '-t', + 'snapshot', + '-o', + 'name', + $sort_params->@*, + nvmet_dataset($scfg, $name), + ), + ]; +} + +# --------------------------------------------------------------------------- +# Storage API: volumes +# --------------------------------------------------------------------------- + +sub parse_volname($class, $volname) { + if ($volname =~ $RE_VOLNAME) { + my $format = ($+{type} eq 'subvol' || $+{type} eq 'basevol') ? 'subvol' : 'raw'; + my $is_base = $+{type} eq 'base' || $+{type} eq 'basevol'; + return ('images', $+{name}, $+{vmid}, $+{base}, $+{base_vmid}, $is_base, $format); + } + die "unable to parse zfs volume name '$volname'\n"; +} + +sub list_images($class, $storeid, $scfg, $vmid = undef, $vollist = undef, $cache = undef) { + my $res = []; + for my $info (values zfs_list_zvol($scfg, 1)->%*) { + my $volname = + $info->{parent} && $info->{parent} =~ $RE_BASE_SNAPSHOT + ? "$storeid:$+{base}/$info->{name}" + : "$storeid:$info->{name}"; + next + if $vollist + ? !grep { $_ eq $volname } $vollist->@* + : defined($vmid) && $info->{vmid} ne $vmid; + $info->{volid} = $volname; + push $res->@*, $info; + } + return $res; +} + +# Names are allocated against every direct child of the pool, owned or not, so +# a leftover zvol that list_images hides is never handed out again. +sub find_free_diskname($class, $storeid, $scfg, $vmid, $fmt = undef, $add_fmt_suffix = undef) { + my $disk_list = [map { "$storeid:$_" } sort keys zfs_list_zvol($scfg, 0)->%*]; + return PVE::Storage::Plugin::get_next_vm_diskname( + $disk_list, $storeid, $vmid, $fmt, $scfg, $add_fmt_suffix, + ); +} + +sub status($class, $storeid, $scfg, $cache = undef) { + my ($available, $used) = + eval { zfs_get_properties($scfg, 'available,used', nvmet_pool($scfg)) }; + if (my $err = $@) { + warn "storage '$storeid': $err"; + return (0, 0, 0, 0); + } + if ( + !defined($available) + || !defined($used) + || $available !~ $RE_UNSIGNED_INTEGER + || $used !~ $RE_UNSIGNED_INTEGER + ) { + warn "unexpected ZFS pool usage for storage '$storeid'\n"; + return (0, 0, 0, 0); + } + return ($available + $used, $available, $used, 1); +} + +sub volume_size_info($class, $scfg, $storeid, $volname, $timeout = undef) { + my (undef, $name, undef, $parent, undef, undef, $format) = $class->parse_volname($volname); + die "volume_size_info requires a ZFS volume\n" if $format ne 'raw'; + my ($size, $used) = zfs_get_properties( + $scfg, 'volsize,usedbydataset', nvmet_dataset($scfg, $name), $timeout, + ); + die "Could not get zfs volume size\n" if !defined($size) || $size !~ $RE_UNSIGNED_INTEGER; + $used = defined($used) && $used =~ $RE_UNSIGNED_INTEGER ? $used + 0 : 0; + return wantarray ? ($size + 0, 'raw', $used, $parent) : $size + 0; +} + +sub volume_snapshot($class, $scfg, $storeid, $volname, $snap) { + my (undef, $name, undef, undef, undef, undef, $format) = $class->parse_volname($volname); + die "volume_snapshot requires a ZFS volume\n" if $format ne 'raw'; + my $snapshot = nvmet_snapshot($scfg, $name, $snap); + nvmet_locked($scfg, sub { zfs_request($scfg, undef, 'snapshot', $snapshot) }); + return; +} + +sub volume_snapshot_delete($class, $scfg, $storeid, $volname, $snap, $running = undef) { + my $name = ($class->parse_volname($volname))[1]; + my $snapshot = nvmet_snapshot($scfg, $name, $snap); + nvmet_locked($scfg, sub { zfs_request($scfg, undef, 'destroy', $snapshot) }); + return; +} + +sub volume_rollback_is_possible($class, $scfg, $storeid, $volname, $snap, $blockers = undef) { + my $name = ($class->parse_volname($volname))[1]; + nvmet_snapshot($scfg, $name, $snap); # validation only + my $found; + $blockers //= []; + for my $snapshot (zfs_get_sorted_snapshot_list($scfg, $name, ['-s', 'creation'])->@*) { + $found = 1 if $snapshot eq $snap; + push $blockers->@*, $snapshot if $found && $snapshot ne $snap; + } + die "can't rollback, snapshot '$snap' does not exist on '${storeid}:${volname}'\n" if !$found; + die "can't rollback, '$snap' is not most recent snapshot on '${storeid}:${volname}'\n" + if $blockers->@*; + return 1; +} + +sub volume_snapshot_rollback($class, $scfg, $storeid, $volname, $snap) { + my (undef, $name, undef, undef, undef, undef, $format) = $class->parse_volname($volname); + die "snapshot rollback requires a ZFS volume\n" if $format ne 'raw'; + nvmet_rollback_volume($scfg, $name, $snap); +} + +sub volume_snapshot_info($class, $scfg, $storeid, $volname) { + my $name = ($class->parse_volname($volname))[1]; + my $info = {}; + my $text = zfs_request( + $scfg, + undef, + 'list', + '-Hp', + '-r', + '-t', + 'snapshot', + '-o', + 'name,guid,creation', + nvmet_dataset($scfg, $name), + ); + for my $line (split /\n/, $text) { + my ($snapshot, $guid, $creation) = split /\s+/, $line; + $snapshot =~ s/^.*\@//; + $info->{$snapshot} = { id => $guid, timestamp => $creation }; + } + return $info; +} + +sub free_image($class, $storeid, $scfg, $volname, $is_base = undef, $format = undef) { + my (undef, $name, undef, undef, undef, undef, $volformat) = $class->parse_volname($volname); + die "free_image requires a ZFS volume\n" if $volformat ne 'raw'; + nvmet_destroy_volume($scfg, $name); + return undef; +} + +sub create_base($class, $storeid, $scfg, $volname) { + my (undef, $name, undef, $basename, undef, $is_base, $format) = $class->parse_volname($volname); + die "create_base not possible with base image\n" if $is_base; + die "create_base requires a ZFS volume\n" if $format ne 'raw'; + my $newname = $name =~ s/^vm-/base-/r; + nvmet_template_volume($scfg, $name); + return $basename ? "$basename/$newname" : $newname; +} + +sub clone_image($class, $scfg, $storeid, $volname, $vmid, $snap = undef) { + $snap ||= '__base__'; + my (undef, $basename, undef, undef, undef, $is_base, $format) = $class->parse_volname($volname); + die "clone_image only works on base images\n" if !$is_base; + die "clone_image requires a ZFS volume\n" if $format ne 'raw'; + my $origin = nvmet_snapshot($scfg, $basename, $snap); + my $name = $class->find_free_diskname($storeid, $scfg, $vmid, $format); + nvmet_create_volume($scfg, $name, origin => $origin); + return "$basename/$name"; +} + +sub alloc_image($class, $storeid, $scfg, $vmid, $fmt, $name, $size) { + die "unsupported format '$fmt'" if $fmt ne 'raw'; + die "illegal name '$name' - should be 'vm-$vmid-*'\n" + if $name && index($name, "vm-$vmid-") != 0; + my $volname = $name || $class->find_free_diskname($storeid, $scfg, $vmid, $fmt); + $size += (1024 - $size % 1024) % 1024; + nvmet_create_volume($scfg, $volname, size => $size); + return $volname; +} + +sub volume_resize($class, $scfg, $storeid, $volname, $size, $running = undef, $snapname = undef) { + # QEMU can resize a host_device once the device itself has grown, but the + # plugin cannot yet wait for the local NVMe namespace to report the new + # size. Until it can, refuse online resize before changing the zvol, so a + # running VM never sees a size different from its configuration. + die "online resize is not supported for NVMe/TCP block devices; stop the VM first\n" + if $running; + die "resizing a snapshot is not supported for $class\n" if $snapname; + my $new_size = int($size / 1024); + $new_size += (1024 - $new_size % 1024) % 1024; + my $name = ($class->parse_volname($volname))[1]; + nvmet_resize_volume($scfg, $name, $new_size); + return $new_size; +} + +sub volume_has_feature( + $class, $scfg, $feature, $storeid, $volname, + $snapname = undef, + $running = undef, + $opts = undef, +) { + my $features = { + snapshot => { current => 1, snap => 1 }, + clone => { base => 1 }, + template => { current => 1 }, + copy => { base => 1, current => 1 }, + }; + my $is_base = ($class->parse_volname($volname))[5]; + my $key = $snapname ? 'snap' : $is_base ? 'base' : 'current'; + return $features->{$feature}->{$key}; +} + +# No stream transfers or renames; the storage has no path, so the base class +# offers no export or import format either. +sub volume_export($class, @args) { + die "ZFS stream export is not supported for NVMe/TCP storage\n"; +} + +sub volume_import($class, @args) { + die "ZFS stream import is not supported for NVMe/TCP storage\n"; +} + +sub rename_volume($class, @args) { + die "renaming volumes is not supported for NVMe/TCP storage\n"; +} + +sub rename_snapshot($class, @args) { + die "rename_snapshot is not supported for $class"; +} + +# --------------------------------------------------------------------------- +# Secrets and configuration hooks +# --------------------------------------------------------------------------- + +# Each storage needs its own NVMe/TCP listener (address family, address and +# port), and a disabled storage keeps its listener. nvmet answers a connect to +# a listener that does not publish the subsystem yet with DNR, and the host +# then deletes the controller whatever its loss timeout: after a target +# restart, the storage published second would lose its paths. Storages that +# reach the same data address must spell the server alike, as it names the +# target lock. +sub _assert_unique_target($storeid, $scfg, $cfg = undef) { + $cfg //= PVE::Storage::config(); + my (%listeners, %addresses); + for my $portal (parse_nvme_portals($scfg->{'nvme-portals'})->@*) { + my $address = join("\0", nvmet_listener_address($portal)); + $listeners{"$address\0$portal->{port}"} = 1; + $addresses{$address} = 1; + } + my $ids = $cfg->{ids} // {}; + for my $other_id (sort keys $ids->%*) { + next if $other_id eq $storeid; + my $other = $ids->{$other_id}; + next if ($other->{type} // '') ne 'zfsnvme'; + + die "NVMe subsystem NQN is already used by storage '$other_id'\n" + if ($other->{subsysnqn} // '') eq $scfg->{subsysnqn}; + my $other_server = $other->{server} // ''; + # a storage whose portals do not parse cannot be activated + for my $portal ((parse_nvme_portals($other->{'nvme-portals'}, 1) // [])->@*) { + my $address = join("\0", nvmet_listener_address($portal)); + die "NVMe/TCP portal '$portal->{address}' port $portal->{port} is already used by" + . " storage '$other_id'; give each storage its own address or port\n" + if $listeners{"$address\0$portal->{port}"}; + die "storage '$other_id' reaches target address '$portal->{address}' through server" + . " '$other_server'; use the same server value for both storages\n" + if $addresses{$address} && $other_server ne $scfg->{server}; + } + next if $other_server ne $scfg->{server}; + my ($pool, $other_pool) = ($scfg->{pool}, $other->{pool} // ''); + die "ZFS pool '$pool' on '$scfg->{server}' is already used by storage '$other_id'\n" + if $other_pool eq $pool; + die "ZFS pool '$pool' on '$scfg->{server}' overlaps the pool of storage '$other_id'\n" + if index("$other_pool/", "$pool/") == 0 || index("$pool/", "$other_pool/") == 0; + } +} + +# nvmet keeps one DH-HMAC-CHAP key per host NQN, so storages on the same target +# whose host NQNs overlap must use the same key. +sub _assert_shared_host_key($storeid, $scfg, $key, $cfg = undef) { + $cfg //= PVE::Storage::config(); + my %hosts = map { $_ => 1 } (parse_nvme_host_nqns($scfg->{'nvme-host-nqns'}, 1) // [])->@*; + my $ids = $cfg->{ids} // {}; + for my $other_id (sort keys $ids->%*) { + next if $other_id eq $storeid; + my $other = $ids->{$other_id}; + next if ($other->{type} // '') ne 'zfsnvme'; + next if ($other->{server} // '') ne ($scfg->{server} // ''); + my $other_hosts = parse_nvme_host_nqns($other->{'nvme-host-nqns'}, 1) // []; + next if !grep { $hosts{$_} } $other_hosts->@*; + my $other_key = file_read_firstline(secret_path($other_id)); + die "storage '$other_id' uses a different DH-HMAC-CHAP key for the same NVMe host NQNs" + . " on '$scfg->{server}'\n" + if defined($other_key) && $other_key ne $key; + } +} + +# DHHC-1::: as generated by +# `nvme gen-dhchap-key`. Messages never contain the key. +sub _validate_secret($key) { + die "missing NVMe DH-HMAC-CHAP key\n" if !defined($key) || $key eq ''; + my $invalid = "invalid NVMe DH-HMAC-CHAP key representation\n"; + die $invalid if $key !~ $RE_DHCHAP_KEY; + my ($hash, $encoded) = @+{qw(hash secret)}; + die $invalid if length($encoded) % 4; + my $decoded = decode_base64($encoded); + my $length = length($decoded) - 4; + my %lengths = ('00' => [32, 48, 64], '01' => [32], '02' => [48], '03' => [64]); + die $invalid if !grep { $_ == $length } $lengths{$hash}->@*; + die $invalid if pack('V', crc32(substr($decoded, 0, $length))) ne substr($decoded, $length); + return $key; +} + +my sub set_secret($storeid, $key) { + _validate_secret($key); + make_path($secret_dir, { mode => 0700 }); + file_set_contents(secret_path($storeid), "$key\n", 0600); +} + +my sub get_secret($storeid) { + my $key = file_read_firstline(secret_path($storeid)); + return _validate_secret($key); +} + +# Removes the key file of a storage (a test seam). +sub _unlink_file($path) { + return unlink($path); +} + +my sub delete_secret($storeid) { + _unlink_file(secret_path($storeid)); +} + +sub on_add_hook($class, $storeid, $scfg, %sensitive) { + _configured_portals($scfg); + parse_nvme_host_nqns($scfg->{'nvme-host-nqns'}); + _assert_unique_target($storeid, $scfg); + my $key = _validate_secret($sensitive{'dhchap-key'}); + _assert_shared_host_key($storeid, $scfg, $key); + set_secret($storeid, $key); + return; +} + +sub on_update_hook_full($class, $storeid, $scfg, $update, $delete = undef, $sensitive = undef) { + $sensitive //= {}; + my %prospective = ($scfg->%*, $update->%*); + delete @prospective{ $delete->@* } if $delete; + verify_nvme_nqn($prospective{subsysnqn}); + _configured_portals(\%prospective); + my $new_hostnqns = parse_nvme_host_nqns($prospective{'nvme-host-nqns'}); + my %new_hosts = map { $_ => 1 } $new_hostnqns->@*; + for my $old_hostnqn (parse_nvme_host_nqns($scfg->{'nvme-host-nqns'})->@*) { + die "removing NVMe host NQN '$old_hostnqn' is not supported; move every volume off" + . " the storage and recreate it with a new subsystem NQN and a new DH-HMAC-CHAP" + . " key\n" + if !$new_hosts{$old_hostnqn}; + } + _validate_fail_fast_timeout(\%prospective, 600); + _assert_unique_target($storeid, \%prospective); + + my $old_key = file_read_firstline(secret_path($storeid)); + my $key = exists($sensitive->{'dhchap-key'}) ? $sensitive->{'dhchap-key'} : $old_key; + _validate_secret($key); + + # nvmet stores authentication on the global Host NQN object, which can be + # shared by multiple subsystems. Replacing an active key in place can make + # unrelated storages unrecoverable on their next reconnect. Until PVE can + # coordinate a rolling rotation on every node, fail before mutating the + # cluster-wide secret. + die "NVMe DH-HMAC-CHAP key rotation is not supported; create a new storage with a new" + . " subsystem NQN\n" + if defined($old_key) && $key ne $old_key; + _assert_shared_host_key($storeid, \%prospective, $key); + + set_secret($storeid, $key) if exists($sensitive->{'dhchap-key'}); + return; +} + +# --------------------------------------------------------------------------- +# Local NVMe host side +# --------------------------------------------------------------------------- + +my %slow_path_last; # storeid => time of the last slow-path activation + +# Probes of the local host (test seams). +sub _block_device($path) { + return -b $path; +} + +sub _link_target($path) { + return readlink($path); +} + +sub _fabrics_device() { + return '/dev/nvme-fabrics'; +} + +# Runs $code in a child and gives up on it after $timeout seconds. Returns the +# child's result; the child's errors pass through as exceptions, a timeout +# becomes one. The bound is not hard: as with PVE::Tools::run_command, a +# child stuck in the kernel (a write or delete in D state) is only reaped +# once the kernel returns, and a timeout does not mean that the operation did +# not happen. +# +# A stopped task (TERM, also HUP, INT and QUIT) stays pending in the parent +# while the child runs, and the child inherits the mask: a connect that ends +# before PVE escalates the stop to KILL, 5 seconds after the TERM, creates +# the controller and restricts its secret attributes without interruption, +# and run_fork_with_timeout, whose own handlers would consume the signal when +# the child delivers a result, never sees it. Once the child is reaped and the +# mask is restored, the handler of the worker dies with the task marker, also +# over a result or an error of the child. ALRM, the timeout, is not blocked. +# The KILL ends the parent and the child together; a controller that the +# child had created but not restricted yet keeps its secret attributes +# world-readable until the next activation restricts them. +# +# run_fork_with_timeout is called in list context on purpose: only there does +# run_with_timeout report a timeout as its second return value; in scalar +# context the timeout would be a warning and an undef result. +my sub run_bounded($timeout, $what, $code) { + my $previous = POSIX::SigSet->new(); + sigprocmask(SIG_BLOCK, POSIX::SigSet->new(SIGHUP, SIGINT, SIGQUIT, SIGTERM), $previous) + or die "cannot block signals: $!\n"; + my $outer_warn = $SIG{__WARN__}; + my ($res, $timed_out) = eval { + # Only a child that leaves no result makes run_fork_with_timeout warn + # in the parent; that case is an exception below. The child keeps the + # handler of the caller, so its own warnings reach the task log. + local $SIG{__WARN__} = sub($message) { }; + PVE::Tools::run_fork_with_timeout( + $timeout, + sub { + local $SIG{__WARN__} = $outer_warn; + return $code->(); + }, + ); + }; + my $error = $@; + sigprocmask(SIG_SETMASK, $previous) or die "cannot restore signals: $!\n"; + if ($error) { + rethrow_task_interrupt($error); + die $error; + } + die "$what did not complete within $timeout seconds\n" if $timed_out; + die "$what left no result\n" if !defined($res); + return $res; +} + +# The connect options of one controller, as the fabrics device takes them, +# and their names. The values are validated configuration and host identities +# checked by the caller; the kernel reads each one up to the next ','. +sub _fabrics_options($scfg, $portal, $hostnqn, $hostid, $key) { + my @options = ( + [transport => 'tcp'], + [traddr => $portal->{address}], + [trsvcid => $portal->{port}], + [host_iface => $portal->{host_iface}], + [nqn => $scfg->{subsysnqn}], + [hostnqn => $hostnqn], + [hostid => $hostid], + [dhchap_secret => $key], + [keep_alive_tmo => $scfg->{'nvme-keep-alive-tmo'} // 5], + [reconnect_delay => $scfg->{'nvme-reconnect-delay'} // 2], + [ctrl_loss_tmo => $scfg->{'nvme-ctrl-loss-tmo'} // 600], + ); + # unlike unset, which leaves fast I/O failure off, 0 is a real timeout + push @options, [fast_io_fail_tmo => $scfg->{'nvme-fast-io-fail-tmo'}] + if defined($scfg->{'nvme-fast-io-fail-tmo'}); + push @options, [nr_io_queues => $scfg->{'nvme-nr-io-queues'}] + if defined($scfg->{'nvme-nr-io-queues'}); + for my $option (@options) { + die "internal error: invalid NVMe connect option '$option->[0]'\n" + if !defined($option->[1]) || $option->[1] !~ $RE_FABRICS_VALUE; + } + return (join(',', map { "$_->[0]=$_->[1]" } @options), [map { $_->[0] } @options]); +} + +# The child of a connect (a test seam): creates exactly one controller with +# one write of $options to the fabrics device and returns its instance. The +# write is never repeated, because the kernel may have created the controller +# even when it reports an error. Errors carry the errno text of the failed +# operation and never the options, which hold the key. +sub _fabrics_connect($options, $names) { + my $device = _fabrics_device(); + my $probe; + if (!sysopen($probe, $device, O_RDONLY | O_NOFOLLOW)) { + die "'$device' does not exist; load the nvme-tcp kernel module\n" if $! == ENOENT; + die "cannot open '$device': $!\n"; + } + my @stat = stat($probe); + die "'$device' is not a character device\n" if !@stat || !S_ISCHR($stat[2]); + # without a controller, a read lists the supported options as patterns + my $supported = ''; + sysread($probe, $supported, 4096) // die "cannot read '$device': $!\n"; + close($probe); + my %supported = map { s/=.*//r => 1 } split(/,/, $supported =~ s/\n\z//r); + for my $name ($names->@*) { + die "the running kernel does not support the NVMe connect option '$name'\n" + if !$supported{$name}; + } + + # Not PVE::SysFSTools::file_write: the kernel returns the instance of the + # new controller on a read of the file descriptor that wrote the options. + sysopen(my $fh, $device, O_RDWR | O_NOFOLLOW) or die "cannot open '$device': $!\n"; + my $written = _fabrics_write($fh, $options); + die "connect failed: $!\n" if !defined($written); + die "connect failed: short write\n" if $written != length($options); + my $result = ''; + sysread($fh, $result, 128) // die "cannot read the connect result: $!\n"; + die "unexpected connect result\n" if $result !~ $RE_FABRICS_RESULT; + my $instance = int($+{instance}); + # The kernel creates the secret attributes of a controller world-readable; + # the caller keeps a stopped task from interrupting before this, short of + # a KILL. + _restrict_attr("/sys/class/nvme/nvme$instance/$_") for qw(dhchap_secret dhchap_ctrl_secret); + close($fh); + return $instance; +} + +# The write that creates a controller (a test seam). +sub _fabrics_write($fh, $options) { + return syswrite($fh, $options); +} + +# Writes a sysfs attribute and returns the errno text of a failure, or undef. +# PVE::SysFSTools::file_write returns undef without a warning when the open +# fails, leaving its errno in $!, and 0 after warning "error writing ...: +# " when the write fails; that warning is not passed on. With +# $missing_ok, an attribute that does not exist, because its controller went +# away, is not a failure. +my sub write_attribute($path, $value, $missing_ok = 0) { + my $warning; + my $written = do { + local $SIG{__WARN__} = sub($message) { $warning = $message }; + PVE::SysFSTools::file_write($path, $value); + }; + return if $written; + return if !defined($warning) && $missing_ok && $! == ENOENT; + return "$!" if !defined($warning); + return $warning =~ s/\A.*: //sr =~ s/\n\z//r; +} + +# Deletes controllers (the body of a child, a test seam): the deletion can +# block in the kernel. The error names the controller and the errno text; a +# controller that is already gone counts as deleted. +sub _delete_controllers($controllers) { + for my $controller ($controllers->@*) { + my $path = "/sys/class/nvme/$controller->{name}/delete_controller"; + my $error = write_attribute($path, '1', 1) // next; + die "cannot delete NVMe controller '$controller->{name}': $error\n"; + } + return scalar($controllers->@*); +} + +# Every local controller of the subsystem. +my sub nvme_controllers($nqn) { + my $controllers = []; + PVE::File::dir_glob_foreach( + '/sys/class/nvme', + $RE_NVME_CONTROLLER, + sub($entry) { + my $base = "/sys/class/nvme/$entry"; + my $subsys = file_read_firstline("$base/subsysnqn"); + return if !defined($subsys) || $subsys ne $nqn; + my $address = file_read_firstline("$base/address") // ''; + my $traddr = $address =~ $RE_TRADDR ? $+{value} : undef; + my $trsvcid = $address =~ $RE_TRSVCID ? $+{value} : undef; + my $host_iface = $address =~ $RE_HOST_IFACE_ADDRESS ? $+{value} : undef; + push $controllers->@*, + { + name => $entry, + state => file_read_firstline("$base/state") // 'unknown', + traddr => nvmet_canonical_address($traddr), + trsvcid => $trsvcid, + host_iface => $host_iface, + }; + }, + ); + return $controllers; +} + +# Controllers by "address:port", with canonical IPv6 addresses. A portal has +# more than one controller while its path moves to another interface. +my sub controller_states($nqn) { + my $states = {}; + for my $controller (nvme_controllers($nqn)->@*) { + next if !defined($controller->{traddr}) || !defined($controller->{trsvcid}); + push $states->{"$controller->{traddr}:$controller->{trsvcid}"}->@*, $controller; + } + return $states; +} + +# The multipath namespace devices of the subsystem and their partitions. One +# dir_glob_foreach per directory level, because PVE::File::dir_glob_regex +# returns only the first match. +my sub namespace_devices($nqn) { + my $devices = {}; + + PVE::File::dir_glob_foreach( + '/sys/class/nvme-subsystem', + $RE_NVME_SUBSYSTEM, + sub($entry) { + my $base = "/sys/class/nvme-subsystem/$entry"; + my $subsys = file_read_firstline("$base/subsysnqn"); + return if !defined($subsys) || $subsys ne $nqn; + + PVE::File::dir_glob_foreach( + $base, + $RE_NVME_NAMESPACE, + sub($device) { + $devices->{"/dev/$device"} = $device if _block_device("/dev/$device"); + PVE::File::dir_glob_foreach( + "/sys/class/block/$device", + $RE_NVME_PARTITION, + sub($partition) { + $devices->{"/dev/$partition"} = $partition + if _block_device("/dev/$partition"); + }, + ); + }, + ); + }, + ); + return $devices; +} + +sub _namespace_openers($nqn) { + my $devices = namespace_devices($nqn); + return [] if !$devices->%*; + + my %openers; + PVE::File::dir_glob_foreach( + '/proc', + '[0-9]+', + sub($pid) { + PVE::File::dir_glob_foreach( + "/proc/$pid/fd", + '[0-9]+', + sub($fd) { + my $target = _link_target("/proc/$pid/fd/$fd"); + return if !defined($target) || !exists($devices->{$target}); + my $comm = eval { file_read_firstline("/proc/$pid/comm") } // 'unknown'; + $openers{"$pid:$target"} = "$comm (PID $pid, $target)"; + }, + ); + }, + ); + + # Kernel consumers such as device-mapper do not necessarily keep a userspace + # file descriptor open, but expose their dependency in the holders directory. + for my $path (keys $devices->%*) { + my $device = $devices->{$path}; + my $holders = "/sys/class/block/$device/holders"; + PVE::File::dir_glob_foreach( + $holders, + '[^\\.].*', + sub($holder) { + $openers{"holder:$device:$holder"} = "$path held by $holder"; + }, + ); + } + + return [sort values %openers]; +} + +my sub portal_reachable($portal) { + return PVE::Network::tcp_ping($portal->{address}, $portal->{port}, 2) // 0; +} + +# The number of portals with a live controller on their configured interface, +# or, with $any_iface, on any interface: such a path carries I/O, even while +# it waits to be moved to its configured interface. +sub _live_portal_count($states, $portals, $any_iface = 0) { + my $live = 0; + for my $portal ($portals->@*) { + my $id = "$portal->{address}:$portal->{port}"; + $live++ if grep { + $_->{state} eq 'live' + && ($any_iface || ($_->{host_iface} // '') eq $portal->{host_iface}) + } ($states->{$id} // [])->@*; + } + return $live; +} + +# Creates the controller of a portal in a child, giving up after 10 seconds. +sub _connect_portal($scfg, $portal, $hostnqn, $hostid, $key) { + my ($options, $names) = _fabrics_options($scfg, $portal, $hostnqn, $hostid, $key); + return run_bounded(10, 'NVMe/TCP connect', sub { return _fabrics_connect($options, $names) }); +} + +my sub set_iopolicy($nqn, $policy) { + PVE::File::dir_glob_foreach( + '/sys/class/nvme-subsystem', + $RE_NVME_SUBSYSTEM, + sub($entry) { + my $base = "/sys/class/nvme-subsystem/$entry"; + my $subsys = file_read_firstline("$base/subsysnqn"); + return if !defined($subsys) || $subsys ne $nqn; + my $error = write_attribute("$base/iopolicy", "$policy\n"); + die "unable to set NVMe multipath policy: $error\n" if defined($error); + }, + ); +} + +# Applies the reconnect timeouts to connected controllers, which otherwise +# keep the values of their connect. sysfs shows "off" for -1, and shows the +# loss timeout rounded up to a multiple of the reconnect delay. +my sub set_tunables($scfg) { + my $delay = $scfg->{'nvme-reconnect-delay'} // 2; + my $loss = $scfg->{'nvme-ctrl-loss-tmo'} // 600; + my $fast = $scfg->{'nvme-fast-io-fail-tmo'}; + my $shown_loss = $loss < 0 ? 'off' : int(($loss + $delay - 1) / $delay) * $delay; + my $shown_fast = defined($fast) && $fast >= 0 ? $fast : 'off'; + + for my $controller (nvme_controllers($scfg->{subsysnqn})->@*) { + my $base = "/sys/class/nvme/$controller->{name}"; + my $current_delay = file_read_firstline("$base/reconnect_delay") // next; + my $write = sub($attr, $value) { + my $error = write_attribute("$base/$attr", "$value\n") // return; + log_warn("cannot set $attr of NVMe controller '$controller->{name}': $error"); + }; + my $delay_changed = $current_delay ne "$delay"; + $write->('reconnect_delay', $delay) if $delay_changed; + # the loss timeout is stored as a number of reconnects of the delay + $write->('ctrl_loss_tmo', $loss < 0 ? -1 : $loss) + if $delay_changed + || (file_read_firstline("$base/ctrl_loss_tmo") // '') ne "$shown_loss"; + $write->('fast_io_fail_tmo', $shown_fast eq 'off' ? -1 : $fast) + if (file_read_firstline("$base/fast_io_fail_tmo") // '') ne "$shown_fast"; + } +} + +# Attempts to make a controller attribute readable by root only (a test seam). +sub _restrict_attr($path) { + my @stat = stat($path); + if (!@stat) { + log_warn("cannot stat NVMe controller attribute '$path': $!") if $! != ENOENT; + return; + } + return if !($stat[2] & 077); + chmod(0600, $path) or log_warn("cannot restrict permissions of '$path': $!"); +} + +# The kernel creates the DH-HMAC-CHAP secret attributes world-readable. +my sub restrict_secret_attrs($nqn) { + for my $controller (nvme_controllers($nqn)->@*) { + _restrict_attr("/sys/class/nvme/$controller->{name}/$_") + for qw(dhchap_secret dhchap_ctrl_secret); + } +} + +# Best effort: a controller that went away since it was listed needs no scan, +# and one that cannot be rescanned is a warning. +my sub rescan_namespaces($nqn) { + for my $controller (nvme_controllers($nqn)->@*) { + next if $controller->{state} ne 'live'; + my $path = "/sys/class/nvme/$controller->{name}/rescan_controller"; + my $error = write_attribute($path, "1\n", 1) // next; + log_warn("cannot rescan NVMe controller '$controller->{name}': $error"); + } +} + +# Deletes a controller that is dead or that moved to another interface. The +# activation reconnects or keeps the path either way, so this only warns. +my sub disconnect_controller($controller) { + my $name = $controller->{name}; + eval { + run_bounded( + 5, + "deleting NVMe controller '$name'", + sub { return _delete_controllers([$controller]) }, + ); + }; + if (my $error = $@) { + rethrow_task_interrupt($error); + log_warn($error); + } +} + +# Waits up to 10 seconds for a live controller of the portal on its +# configured interface. +my sub wait_for_live_path($nqn, $portal) { + for (my $attempt = 0; $attempt < 40; $attempt++) { + return 1 if _live_portal_count(controller_states($nqn), [$portal]); + _sleep(0.25); + } + return 0; +} + +# Connects every missing or dead path on its configured interface. A path +# that is connected on another interface moves make-before-break: its old +# controller is only disconnected once the new one is live, so changing an +# interface mapping never takes away a path that carries I/O. +my sub connect_portals($scfg, $portals, $hostnqn, $hostid, $key) { + my $nqn = $scfg->{subsysnqn}; + for my $portal ($portals->@*) { + my $id = "$portal->{address}:$portal->{port}"; + my $iface = $portal->{host_iface}; + my @controllers = (controller_states($nqn)->{$id} // [])->@*; + my ($configured) = grep { ($_->{host_iface} // '') eq $iface } @controllers; + my @moved = grep { ($_->{host_iface} // '') ne $iface } @controllers; + + if (!$configured || $configured->{state} eq 'dead') { + disconnect_controller($configured) if $configured; + if (!portal_reachable($portal)) { + log_warn("NVMe/TCP portal '$id' is unreachable"); + next; + } + eval { _connect_portal($scfg, $portal, $hostnqn, $hostid, $key) }; + my $connect_error = $@; + # A failed or timed-out connect can still have created a controller. + eval { restrict_secret_attrs($nqn); 1 }; + if (my $error = $@) { + rethrow_task_interrupt($error); + log_warn("cannot restrict NVMe controller attributes for '$nqn'"); + } + rethrow_task_interrupt($connect_error) if $connect_error; + if ($connect_error) { + log_warn("$id: $connect_error"); + next; + } + } + next if !@moved; + if (!wait_for_live_path($nqn, $portal)) { + log_warn("NVMe/TCP portal '$id' is not live on '$iface' yet; keeping its path on" + . " another interface"); + next; + } + disconnect_controller($_) for @moved; + } +} + +sub activate_storage($class, $storeid, $scfg, $cache = undef) { + $cache //= {}; + my $nqn = $scfg->{subsysnqn}; + # First, whatever the rest of the activation does, also for controllers + # connected by hand. Also try after each connect attempt below, including + # failures that may have created a controller. + restrict_secret_attrs($nqn); + + my $multipath = file_read_firstline('/sys/module/nvme_core/parameters/multipath') + // die "the NVMe kernel modules are not loaded; load the nvme-tcp kernel module\n"; + die "native NVMe multipath is disabled in the running kernel\n" if $multipath ne 'Y'; + + my $portals = _configured_portals($scfg); + verify_nvme_nqn($nqn); + _validate_local_ifaces($portals); + _assert_unique_target($storeid, $scfg); + my $hostnqn = file_read_firstline('/etc/nvme/hostnqn') + // die "missing /etc/nvme/hostnqn; the nvme-cli package generates it\n"; + verify_nvme_nqn($hostnqn); + my $hostnqns = parse_nvme_host_nqns($scfg->{'nvme-host-nqns'}); + die "local NVMe host NQN '$hostnqn' is missing from nvme-host-nqns\n" + if !grep { $_ eq $hostnqn } $hostnqns->@*; + my $policy = $scfg->{'nvme-iopolicy'} // 'round-robin'; + my $force_reconcile = delete($cache->{'zfsnvme-force-reconcile'}->{$storeid}) // 0; + my $states = controller_states($nqn); + my $healthy = _live_portal_count($states, $portals); + my $usable = _live_portal_count($states, $portals, 1); + my $last = $slow_path_last{$storeid}; + my $recent = defined($last) && _now() - $last < $slow_path_backoff; + + # This method is called by the periodic storage status loop. Once every + # configured path is live, lifecycle operations keep the target converged, + # so no remote call is needed. A degraded storage retries the target at + # most once a minute: in between, one with a live path (also one waiting + # to move to its configured interface) is usable, and one without fails + # at once instead of waiting for a path again. Missing namespaces + # explicitly force the slow path from activate_volume(). + if (!$force_reconcile && ($healthy == scalar($portals->@*) || $recent)) { + die "no live NVMe/TCP path for storage '$storeid'\n" if !$usable; + die "missing NVMe DH-HMAC-CHAP key\n" + if (file_read_firstline(secret_path($storeid)) // '') eq ''; + set_iopolicy($nqn, $policy); + set_tunables($scfg); + return 1; + } + + $slow_path_last{$storeid} = _now(); + my $hostid = file_read_firstline('/etc/nvme/hostid') + // die "missing /etc/nvme/hostid; the nvme-cli package generates it\n"; + die "invalid NVMe host ID\n" + if $hostid !~ $RE_NVMET_UUID || $hostid =~ $RE_ZERO_UUID; + my $key = get_secret($storeid); + for my $name (qw(keep-alive-tmo reconnect-delay ctrl-loss-tmo fast-io-fail-tmo nr-io-queues)) { + my $value = $scfg->{"nvme-$name"}; + die "invalid NVMe connection parameter '$name'\n" + if defined($value) && $value !~ $RE_CONFIG_INT; + } + _nvmet_activate_target($storeid, $scfg, $portals, $hostnqns, $key); + connect_portals($scfg, $portals, $hostnqn, $hostid, $key); + + # Existing controllers can be in the middle of their kernel reconnect delay + # after a target restart. Do not create duplicates, but give that recovery + # cycle enough time to complete before declaring the storage unavailable. + for (my $attempt = 0; $attempt < 60; $attempt++) { + $states = controller_states($nqn); + last if _live_portal_count($states, $portals, 1); + _sleep(0.25); + } + die "no live NVMe/TCP path for storage '$storeid'\n" + if !_live_portal_count($states, $portals, 1); + $healthy = _live_portal_count($states, $portals); + log_warn("storage '$storeid' is degraded: $healthy/" . scalar($portals->@*) . " paths live") + if $healthy < scalar($portals->@*); + + set_iopolicy($nqn, $policy); + set_tunables($scfg); + return 1; +} + +sub deactivate_storage($class, $storeid, $scfg, $cache = undef) { + my $nqn = $scfg->{subsysnqn}; + my $openers = _namespace_openers($nqn); + die "refusing to disconnect NVMe storage '$storeid': namespace in use by " + . join(', ', $openers->@*) . "\n" + if $openers->@*; + + # One child deletes every controller, bounded by 15 seconds in all. When it + # gives up, the controllers that are left stay connected and the + # deactivation fails. + my $controllers = nvme_controllers($nqn); + run_bounded( + 15, + "disconnecting NVMe subsystem '$nqn'", + sub { return _delete_controllers($controllers) }, + ) if $controllers->@*; + return 1; +} + +# Removing a storage definition never destroys data and never depends on the +# target, like for every other storage type. Target cleanup is best effort +# and only happens once the storage owns no volume. Other nodes keep their +# connections until they are disconnected there. +sub on_delete_hook($class, $storeid, $scfg) { + # First: an activation elsewhere that waits for the target lock then + # refuses instead of restoring what is removed below. + delete_secret($storeid); + eval { $class->deactivate_storage($storeid, $scfg) }; + if (my $err = $@) { + rethrow_task_interrupt($err); + chomp($err); + log_warn("not disconnecting NVMe storage '$storeid': $err"); + } + eval { _nvmet_delete_target($scfg) }; + if (my $err = $@) { + rethrow_task_interrupt($err); + chomp($err); + log_warn("keeping the NVMe target configuration of storage '$storeid': $err"); + } + return; +} + +sub path($class, $scfg, $volname, $storeid, $snapname = undef) { + die "direct access to snapshots not implemented\n" if defined($snapname); + my ($vtype, $name, $vmid) = $class->parse_volname($volname); + my $uuid = _nvmet_volume_uuid($scfg, $name); + my $path = "/dev/disk/by-id/nvme-uuid.$uuid"; + return ($path, $vmid, $vtype); +} + +sub qemu_blockdev_options( + $class, $scfg, $storeid, $volname, + $machine_version = undef, + $options = undef, +) { + die "direct access to snapshots not implemented\n" if $options->{'snapshot-name'}; + my ($path) = $class->path($scfg, $volname, $storeid); + return { driver => 'host_device', filename => $path }; +} + +# Whether the local device behind a namespace link is the namespace (NSID, +# UUID) of this subsystem, so a UUID claimed by another subsystem is never +# handed to a guest. +sub _nvmet_local_namespace_ok($path, $nqn, $nsid, $uuid) { + my $link = _link_target($path) // return 0; + my $device = (split m{/}, $link)[-1]; + return 0 if !defined($device) || $device !~ $RE_NVME_NAMESPACE; + my $sys = "/sys/block/$device"; + return 0 if lc(file_read_firstline("$sys/uuid") // '') ne lc($uuid); + return 0 if (file_read_firstline("$sys/nsid") // '') ne "$nsid"; + return 0 if (file_read_firstline("$sys/device/subsysnqn") // '') ne $nqn; + return 1; +} + +sub activate_volume( + $class, $storeid, $scfg, $volname, + $snapname = undef, + $cache = undef, + $hints = undef, +) { + die "unable to activate snapshot from remote zfs storage\n" if $snapname; + my (undef, $name) = $class->parse_volname($volname); + my $nqn = nvmet_nqn($scfg); + my $dataset = nvmet_dataset($scfg, $name); + + my ($inv, $cfs) = nvmet_read($scfg); + my ($nsid, $uuid) = _nvmet_owned_identity($inv, $nqn, $dataset); + my $state = _nvmet_export_state($cfs, $nqn, $nsid, $uuid, "/dev/zvol/$dataset"); + my $path = "/dev/disk/by-id/nvme-uuid.$uuid"; + my $ok = sub { + return _block_device($path) && _nvmet_local_namespace_ok($path, $nqn, $nsid, $uuid); + }; + return 1 if $state eq 'present' && $ok->(); + + $cache //= {}; + $cache->{'zfsnvme-force-reconcile'}->{$storeid} = 1; + $class->activate_storage($storeid, $scfg, $cache); + for (my $attempt = 0; $attempt < 40 && !_block_device($path); $attempt++) { + # a host can miss the change notice; rescanning is harmless + rescan_namespaces($nqn) if $attempt % 8 == 0; + _sleep(0.25); + } + die "NVMe namespace for '$volname' did not appear\n" if !_block_device($path); + die "NVMe namespace for '$volname' has an unexpected identity\n" if !$ok->(); + return 1; +} + +sub deactivate_volume( + $class, $storeid, $scfg, $volname, $snapname = undef, $cache = undef, +) { + die "unable to deactivate snapshot from remote zfs storage\n" if $snapname; + return 1; +} + +1;