From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from gate001.proxmox.com (gate001.proxmox.com [IPv6:2a0f:8001:1:32::40]) by lore.proxmox.com (Postfix) with ESMTPS id CE21F1FF09C for ; Mon, 05 Oct 2026 02:26:49 +0200 (CEST) Received: from gate001.proxmox.com (localhost.localdomain [127.0.0.1]) by gate001.proxmox.com (Proxmox) with ESMTP id B97D9216C8; Mon, 05 Oct 2026 02:26:33 +0200 (CEST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=neatech-ar.20251104.gappssmtp.com; s=20251104; t=1791159983; x=1791764783; darn=lists.proxmox.com; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=PfBm+Ta+/vPJ3A2caCtjg/ruWAo2zvZuhw3IRXKIsNA=; b=cAZT+0bA5UXt1xHrlLVlSv/4A5UvK6+llg9mj4pry8dgAtRl9TvdpbIUATL1Km7qaF DOPEjdddVxeZy/DECGgCrmYekX6K6M6qHwPxhnElii9TPTXg4hYcjVxwiaIAeRaGrc8C UnxoIzlPin1AQ5OlWhdiYRSjodFQphWRevrniIvr71MPUbQtLFW3glaa1btY1FWp4Bcu 5mv/YbIP8nFFr0Ly7fpypq0BYU3QuVxMwv91p0LTy8wsAst3A5wbXcJs172IVF7UqEfZ 8MDhI+gyWy+5EhS+uHs1V6xZzC/FHrqH+dVP5IN+TpAAX771z+PG4Tg4Mqh8XgS3hFF2 6AGA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1791159983; x=1791764783; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=PfBm+Ta+/vPJ3A2caCtjg/ruWAo2zvZuhw3IRXKIsNA=; b=G3+PQcgR0jzXsR19o4n4Txw4HGjL9Iuhc76ln7T+LkM32JjAl8YZ9Y4ec3t2iv06dr j62/cRmzcIIDbFz1BQfw8b4jqEk96MKunO0Dkujd83/I1ZC5pyEpdQ1FKNC6NAtsGnBy mXpuR4eMzFH+nC62ISyvaksmW97z7hRZ/SUwpmf8fE3Ro652bnxgR8V63PYZRpEXJC22 lkwLP53TBf5nNzgLPxBw2HYkjtzANEP7xIcSkDk/WWdX3KBxVUiiHySbgZYJfONOLz6W uxoZPOSTwxS0GUgIS9FQBZ+fNKCm0CP289dn5LJ4cUy8V3hVUqSXoBBldxaH4Yd22LI9 ZosQ== X-Gm-Message-State: AFq9FYKHORbuAmJ52y38aJAz5aG6lt9CwDlDLweyfCqw0ks0gL+f1a4T s5hduI43GjEwlIakNNvjCQ/CuQXI1YDImEt3Icao8icMg4Se+rjg3lELXZXVG1mRsnh/btuL1Gu f4xv4QLU= X-Gm-Gg: AYBFou0mIvEwszTjiRNEw2xHcN6CxMIPL7h6SOiIJHuQl5YhN8zdXa4QyPbqH1DFxOu wYqE2mUtx2tvP8USpG2PgobD+91HArR02YcOrGyqz20ihPhM/EBI7SfI67XCjflALYK6WxZ15Ak miTN/iwjGjUEZB+TBmXPg1yHsaEe2hvP4mFJzEdD0Xi/9Uebi38e5biY4Yl4QotaRABmUpjsr1I xLfcpj77ZNFKod+27+HciJqwm4Ozzo1WBYSEh8vTibuQ0JPdXlkoy9R9EeyQxdWvx/xM6P7YTs2 sFXavBEcbj7QPi/AjnZkSAy0bcOacJ+PifmgjXA/Dx2ds+WJVdRy1HSGdZ0LbLqY+SFXrXCzR/D /eog8cc6Mt0Mmv2kGSqxPLIWE2kFEn/OywPDL83ldMIGoU3gwVcoxEXu97gsT+SEd3TqEn5thiw glyEUEYBU8DyFaZDdC6M+ieHqFIfnwwarQd22RscGZlz76cE0/VKI10rDXoTZ+KSxPXx1UNaYK3 X2u+eCBINexLV53yAEJoisBzYbSEAVgIRWMRCEYwoxJsJGFQPBC17zbCMAhN9znCtLj0Zp34yqF X-Received: by 2002:ac5:c891:0:b0:5ce:27d2:aa47 with SMTP id 71dfb90a1353d-5dc177cd2e1mr1130212e0c.1.1791159982314; Sun, 04 Oct 2026 17:26:22 -0700 (PDT) From: Joaquin Varela To: pve-devel@lists.proxmox.com Subject: [PATCH docs v3] storage: document ZFS over NVMe/TCP Date: Sun, 4 Oct 2026 21:26:08 -0300 Message-ID: <20261005002609.571-6-joaquinvarela@neatech.ar> X-Mailer: git-send-email 2.54.0.windows.1 In-Reply-To: <20261005002609.571-1-joaquinvarela@neatech.ar> References: <20261005002609.571-1-joaquinvarela@neatech.ar> MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-SPAM-LEVEL: Spam detection results: 0 AWL -0.198 Adjusted score from AWL reputation of From: address DKIM_SIGNED 0.1 Message has a DKIM or DK signature, not necessarily valid DKIM_VALID -0.1 Message has at least one valid DKIM or DK signature DMARC_PASS -0.1 DMARC pass policy KAM_ASCII_DIVIDERS 0.8 Email that uses ascii formatting dividers and possible spam tricks SPF_HELO_NONE 0.001 SPF: HELO does not publish an SPF Record SPF_PASS -0.001 SPF: sender matches SPF record Message-ID-Hash: IJM5OU5N3MT64D6L55VFEPWRTMGGEBYA X-Message-ID-Hash: IJM5OU5N3MT64D6L55VFEPWRTMGGEBYA X-MailFrom: joaquinvarela@neatech.ar X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; emergency; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header X-Mailman-Version: 3.3.10 Precedence: list List-Id: Proxmox VE development discussion List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: Document the zfsnvme storage backend: target and node requirements, its properties with a configuration example, native multipath and DH-HMAC-CHAP authentication, security and availability limits, and the lifecycle of the target configuration. Include the chapter in pvesm and link its wiki page. Every storage needs its own NVMe/TCP listeners on the target, and they must not publish any other nvmet subsystem: nvmet rejects a connection to a subsystem that is not published on the listener yet with DNR, and the node deletes the controller regardless of the controller loss timeout. The chapter says which portals and server values the backend refuses for that reason, and that adding, updating or activating a storage fails on them. Nodes connect each path with a single write to /dev/nvme-fabrics and delete controllers through sysfs, so nvme-cli is only needed for the host identity files and as the administration tool. The chapter states which connect options the running kernel must support, that the key is never put on a command line or in a temporary file, and which secret attributes are restricted to mode 0600 on the nodes and on the target. Revoking a node also requires removing it from the cluster and replacing the SSH key of every target, since every cluster member can read the storage keys and the SSH keys give root access to the targets. It also explains what happens when every path stays down longer than the controller loss timeout: the kernel removes the controllers and the multipath block devices, and running guests have to be stopped and started once the paths are back. A controller loss timeout of -1 needs a finite fast I/O fail timeout to keep guest I/O from blocking for the whole outage. Signed-off-by: Joaquin Varela --- v3, accompanying "[PATCH storage v3 0/4] add ZFS over NVMe/TCP storage plugin": - the five v2 patches are squashed into one; the structure of the chapter is unchanged - changes that follow the storage v3 implementation: - target requirements: the plugin only runs standard utilities over SSH (listed), and the highest namespace ID handed out is kept in the proxmox:nvme-last-nsid property of the pool - nodes: paths are created through /dev/nvme-fabrics and deleted through sysfs, so nvme-cli is only needed for the host identity files (hostid must be a nonzero UUID) and as the administration tool; the running kernel must support every connect option used, and a failed connect is logged per portal - every storage needs its own listeners (address and port), which must not publish any other nvmet subsystem, because nvmet rejects a subsystem that is not published on a listener yet with DNR and the node then deletes the controller regardless of the controller loss timeout; the plugin refuses a listener used by another zfsnvme storage and a target address reached through a different server value, and the server value is the key of the cluster-wide lock per target - target changes, including the restore after a target restart, take that lock and therefore need quorum; an uncertain command result aborts the operation and the next one reads the target again; queries outside a transaction, such as capacity, do not take it - one cluster per target, pools of one target must not overlap, storages whose host NQNs overlap share the key - nvme-host-ifaces and nvme-host-nqns are required in the schema, an interface change is make-before-break, host NQNs can only be added, so revoking a node has its own procedure - sparse is no longer set by default, like in the web interface - pvesm takes --dhchap-key as the path of a key file - how the key reaches the kernel and the target, and the mode 0600 restriction of the secret attributes on the nodes and the target - the connection tunables are applied to connected controllers on the next activation - no pvesm export/import - removing the storage never deletes volumes; what deactivation and the best-effort target cleanup do - the optional transaction guard of storage 3/4: one sentence in the target requirements and one paragraph in "Security and Availability"; both go away if that patch is dropped - corrections of v2 statements that were wrong or misleading: - DH-HMAC-CHAP only authenticates the hosts to the target, with one key shared by all nodes (v2: "authenticates the endpoints") - the target kernel needs NVMe in-band authentication, which v2 already required but did not list - the plugin itself loads nvmet_tcp and mounts configfs on the target; loading the transport at boot is recommended, not required - all-path loss: after the controller loss timeout the kernel deletes the controllers and the multipath devices, and running guests need a stop and start; a rejection deletes the controller at once; -1 needs a finite fast I/O fail timeout. v2 only said that I/O stays queued, and "600 seconds in the example" was wrong because the example sets a fast I/O fail timeout of 30 seconds - the configuration example uses documentation addresses and an example NQN, and drops "shared 1", which is not a property of this type - the pvesm wiki link uses the page name of the chapter title - nothing else was added, apart from wording ("controller loss timeout" without hyphens) and the anchor for the cross-references to "Security and Availability"; happy to trim further v2: https://lore.proxmox.com/pve-devel/cover.1785636981.git.joaquinvarela@neatech.ar/ pve-storage-zfsnvme.adoc | 373 +++++++++++++++++++++++++++++++++++++++ pvesm.adoc | 4 + 2 files changed, 377 insertions(+) create mode 100644 pve-storage-zfsnvme.adoc diff --git a/pve-storage-zfsnvme.adoc b/pve-storage-zfsnvme.adoc new file mode 100644 index 0000000..c590636 --- /dev/null +++ b/pve-storage-zfsnvme.adoc @@ -0,0 +1,373 @@ +[[storage_zfsnvme]] +ZFS over NVMe/TCP Backend +------------------------- +ifdef::wiki[] +:pve-toplevel: +:title: Storage: ZFS over NVMe/TCP +endif::wiki[] + +Storage pool type: `zfsnvme` + +This backend accesses a remote Linux machine with a ZFS pool and the kernel +NVMe target through `ssh`. For each guest disk it creates a ZVOL, exports it as +an NVMe namespace, and connects the {pve} nodes through native Linux NVMe/TCP +multipath. + +The backend supports thin provisioning, snapshots, rollback, templates, linked +clones, offline resize, and shared-storage live migration. Namespace UUIDs are +stored as ZFS user properties, and the stable `nvme-uuid` device link is used +for guest disks. + +Configuration +~~~~~~~~~~~~~ + +The target needs OpenZFS, configfs, and the `nvmet` and `nvmet-tcp` kernel +modules with NVMe in-band authentication support. The backend runs target +utilities over SSH and does all parsing and planning on the {pve} node, so +neither Perl nor Bash is required on the target. Root's login shell must be a +POSIX-compatible shell, and the target needs these standard utilities: `find` +(with `-prune`, `-path`, and `-exec ... {} +`), `grep`, `env`, `printf`, `test`, +`cat`, `mkdir`, `rmdir`, `rm`, `ln`, `chmod`, `dd`, `tee`, `sha256sum`, +`modprobe`, and `mount`. The transaction guard described in +xref:storage_zfsnvme_security[Security and Availability] also needs `/bin/sh` to +be a POSIX-compatible shell, util-linux `flock` with `-w` and `-E` support, and +`stat` with `-c '%u:%a'` support. Configure root SSH access like for the ZFS +over iSCSI backend. The key for the server address is stored at +`/etc/pve/priv/zfs/_id_rsa`. + +The NVMe target modules must be installed on the target. The backend loads +`nvmet_tcp` and, if necessary, mounts configfs before it restores the target +configuration after a reboot, but loading the transport at boot is recommended. +On a Linux target using systemd, load the transport at boot and verify the +configfs mount: + +---- +# echo nvmet_tcp >/etc/modules-load.d/pve-nvmet.conf +# modprobe nvmet_tcp +# mountpoint /sys/kernel/config +---- + +The target configuration below configfs is derived state. It does not need a +separate persistence service: after the module and ZFS pool are available, the +backend reconstructs namespaces and UUIDs from ZFS user properties, and +reconstructs the subsystem, host ACLs, and ports from the shared storage +configuration during activation. The backend also records the highest +namespace ID it handed out in the `proxmox:nvme-last-nsid` property of the +pool dataset, so that a namespace ID is not reused for a new volume. + +Do not leave cloud-init device discovery enabled on a dedicated Linux target. +A guest cloud-init ZVOL contains a `cidata` filesystem and can otherwise be +mistaken for the target host's own NoCloud datasource during boot. After the +target host is provisioned, remove cloud-init or disable it according to the +distribution's documentation. For distributions supporting the standard +disable marker, use: + +---- +# touch /etc/cloud/cloud-init.disabled +---- + +Each {pve} node needs the `nvme-tcp` module, native NVMe multipath, a unique +`/etc/nvme/hostnqn`, and a nonzero UUID in `/etc/nvme/hostid`. The `nvme-cli` +package creates both files and provides the NVMe administration tools. +Interface names listed in the storage configuration must exist on every node +where the storage is enabled. + +The backend does not use `nvme-cli` to connect or disconnect. It creates each +missing path with a single write to the kernel's `/dev/nvme-fabrics` interface +from a short-lived child process, supplying both host identities explicitly, +and deletes controllers through sysfs. The running kernel must support every +connect option the backend uses, including DH-HMAC-CHAP and `host_iface`; +otherwise the connection fails with an error naming the option. A failed +connection is logged as a warning for its portal, with the error text reported +by the kernel. + +The following properties are specific to ZFS over NVMe/TCP: + +server:: + +IP address or DNS name used for the SSH control connection. Connected guest I/O +uses the independent NVMe/TCP data paths, but capacity reporting and lifecycle +operations require this endpoint. Use a redundant management DNS name or VIP +where the storage appliance provides one. Use the same value for a target in +all storage definitions. A cluster-wide lock serializes transactions using this +value. The backend can only detect another spelling of the same target, such as +its IP address and its DNS name, when the storage definitions share a target +address in `nvme-portals`. + +Use each target with only one {pve} cluster. Storage definitions on the same +target need their own pool, NQN, and listener, see `nvme-portals`. They share +the cluster-wide lock, which does not coordinate with other clusters. + +pool:: + +ZFS pool or child dataset used exclusively by this storage definition. It must +not contain, or be contained in, the pool of another storage definition on the +same target. Do not share an NQN or pool with another cluster. + +subsysnqn:: + +NVMe qualified name of the target subsystem. It must be unique across the +cluster's storage definitions. + +nvme-portals:: + +Comma-separated NVMe/TCP target addresses. The default service is `4420`. +Specify IPv6 addresses in brackets, for example `[2001:db8::10]:4420`. ++ +Each portal is a listener on the target, that is, an address and port. Every +storage definition needs its own listeners, and they must not publish any other +nvmet subsystem, including target configuration not made by {pve}. nvmet +rejects a connection to a subsystem that is not published on that listener yet +with the Do Not Retry (DNR) status, for example while the target configuration +is restored after a target restart, and the node then deletes the controller +regardless of the controller loss timeout. The backend compares the portals +with those of every other ZFS over NVMe/TCP storage definition, including +disabled ones. Adding a storage, any update of its configuration, and its +activation on a node fail if one of its portals uses an address and port that +another storage definition already uses, whatever their `server` values, or an +address, on any port, that another storage definition reaches through a +different `server` value. IPv6 addresses are compared in canonical form, and an +IPv4-mapped IPv6 address counts as the IPv4 address. Use different addresses or +ports, and the same `server` value, for storage definitions on the same target. + +nvme-host-ifaces:: + +Required. Comma-separated local interfaces, matched to `nvme-portals` by +position. An explicit interface prevents a failed data path from reconnecting +over the management network. Activation fails before changing target state if +any configured interface is missing on the local node. When an interface +changes, the backend connects the path on the new interface before it +disconnects the old one. + +nvme-host-nqns:: + +Required. Comma-separated contents of `/etc/nvme/hostnqn` from every cluster +node that may activate the storage. The complete allow-list lets any one node +restore all host ACLs before publishing the subsystem after a target reboot. +Activation fails if the local node's Host NQN is absent. Update this property +before enabling the storage on a newly added cluster node. Host NQNs can only +be added. To revoke a host, follow the procedure in +xref:storage_zfsnvme_security[Security and Availability], which requires +removing the node from the cluster, a new SSH key for the target, and a new key. + +dhchap-key:: + +NVMe DH-HMAC-CHAP secret in `DHHC-1` representation. The value is handled as a +sensitive property and stored below `/etc/pve/priv/storage/` with mode `0600`. +The target stores the key per Host NQN, so storage definitions on the same +target whose `nvme-host-nqns` overlap must use the same key. Key rotation is +not supported, see +xref:storage_zfsnvme_security[Security and Availability]. + +nvme-iopolicy:: + +Native multipath policy: `round-robin`, `queue-depth`, or `numa`. + +nvme-keep-alive-tmo:: + +Keep-alive timeout in seconds. + +nvme-reconnect-delay:: + +Delay between controller reconnect attempts in seconds. + +nvme-ctrl-loss-tmo:: + +Time in seconds during which the kernel keeps reconnecting a lost controller +before it deletes the controller. `-1` retries indefinitely. See +xref:storage_zfsnvme_security[Security and Availability] for the consequences. + +nvme-fast-io-fail-tmo:: + +Optional time in seconds before outstanding I/O fails while a controller is +reconnecting. If this property is absent, the kernel queues I/O until +`nvme-ctrl-loss-tmo` expires. A configured value must not exceed a finite +controller loss timeout. A value of `0` selects immediate fail-fast behavior. + +nvme-nr-io-queues:: + +Optional number of I/O queues per controller. + +blocksize:: + +ZFS volume block size. + +sparse:: + +Use ZFS thin provisioning instead of reserving the full virtual size. If this +property is not set, volumes are thick-provisioned, which is also the default +in the web interface. + +.Configuration Example (`/etc/pve/storage.cfg`) +---- +zfsnvme: nvme-shared + server 203.0.113.10 + pool tank/pve-nvme + subsysnqn nqn.2026-01.com.example:pve-nvme + nvme-portals 192.0.2.10:4420,198.51.100.10:4420 + nvme-host-ifaces ens1f0,ens1f1 + nvme-host-nqns nqn.2014-08.org.nvmexpress:uuid:11111111-1111-1111-1111-111111111111,nqn.2014-08.org.nvmexpress:uuid:22222222-2222-2222-2222-222222222222 + blocksize 16k + sparse 1 + nvme-iopolicy round-robin + nvme-keep-alive-tmo 5 + nvme-reconnect-delay 2 + nvme-ctrl-loss-tmo 600 + nvme-fast-io-fail-tmo 30 + content images +---- + +The DH-HMAC-CHAP key is intentionally not shown in `storage.cfg`. Set it with +the storage creation API or web interface. With `pvesm`, pass the path of a file +that contains the key, for example `--dhchap-key /root/nvme-dhchap.key`. Create +the file with mode `0600`, for example after `umask 077`, and remove it once the +storage is created. + +[[storage_zfsnvme_security]] +Security and Availability +~~~~~~~~~~~~~~~~~~~~~~~~~ + +The target uses an ACL for the unique Host NQN of every node and never enables +`allow_any_host`. All nodes share the DH-HMAC-CHAP key of the storage, and it +only authenticates the hosts to the target. The target is not authenticated to +the hosts, no Diffie-Hellman group is configured, and data is not encrypted. +Host NQNs are not secret, so any holder of the key can connect as any allowed +host. This backend does not configure NVMe/TCP TLS, so use isolated storage +networks or equivalent protection. + +Apart from its key file, the backend keeps the key in memory. It passes the key +to the local kernel in the connect write to `/dev/nvme-fabrics`, and to the +target on the standard input of an SSH command. The key never appears on a +command line, and the backend writes no temporary files. + +The kernel exposes the key with mode `0644` in the `dhchap_secret` sysfs +attributes of the NVMe controllers on the nodes and in the host entries below +`/sys/kernel/config/nvmet/hosts/` on the target. The backend attempts to +restrict the controller attributes on the {pve} nodes to mode `0600` right after +it creates a controller, during storage activation, and after each connection +attempt. This is best effort: permission changes can fail, and controller +attributes can appear or reappear after the permission check. Until then, any +process on the node that can read sysfs, including processes in containers, can +read the key. There is no guaranteed upper bound on the exposure interval. Check +the actual permissions and activation warnings; do not rely on this mitigation +to protect keys from untrusted local users. On the target, the backend restricts +the `dhchap_key` and `dhchap_ctrl_key` attributes of a host entry to mode `0600` +before it writes the key. A process on the target that opened one of these +attributes before the change can still read the key through that descriptor. +Host entries that already hold the key, for example ones created by hand, keep +their mode. Limit shell access on both the target and the {pve} nodes to trusted +administrators. + +To rotate the key, or to revoke a node: + +. To revoke a node, first remove it from the cluster. Every cluster member can + read the keys below `/etc/pve/priv/`: the DH-HMAC-CHAP key of every ZFS over + NVMe/TCP storage and the SSH key of every target, which gives root access to + the target, including its keys and volumes. Then replace the SSH key of each + target: create a new key pair at `/etc/pve/priv/zfs/_id_rsa` (for + every `server` value used for the target), replace the old public key with + the new one in `/root/.ssh/authorized_keys` on the target, and check that + file for other unknown keys. +. Move or remove every volume of the storage, including unused disks, + templates, VM state volumes, and disks with snapshots, which cannot be moved + while their snapshots exist. `pvesm list ` must show no volume. +. Remove this storage definition. The backend removes the subsystem from the + target once it owns no volumes. If the removal task warns that it kept the + NVMe target configuration, remove the remaining volumes, or the subsystem on + the target, before you continue. +. On every other node that had the storage active, disconnect it with + `nvme disconnect --nqn ` or reboot the node. +. Create a new storage with a new key and a new `subsysnqn`, without the Host + NQN of a revoked node. + +To revoke a node, repeat this for every ZFS over NVMe/TCP storage of the +cluster, since the node could read all of their keys. Other storage definitions +on the same target that list any of the same Host NQNs share the key, because +the target stores it per Host NQN. Rotate their keys together: remove all of +them before you create the new ones. + +Shared access relies on {pve} cluster locking and fencing. Test quorum and +fencing before placing production guests on the storage. Use independent +failure domains for the configured paths and make sure the ZFS target itself +is not a single point of failure. + +Changes to the target configuration, including restoring it after a target +restart, are serialized by a cluster-wide lock per target and therefore require +quorum. An uncertain command result aborts the operation; a subsequent +operation reads the target again. Queries outside a transaction, such as +capacity and namespace identity queries, do not take this lock. They can see +intermediate state, and the underlying ZFS queries can still be delayed by +storage activity. + +In addition, the transaction guard of the backend serializes the commands of +each transaction with a lock on the target, in the runtime directory +`/run/pve-storage-nvmet`. Before claiming a transaction, the backend observes +the current owner without taking the target lock. The claim compares that +observation under the lock before publishing a new token, so an abandoned claim +cannot replace an owner that has since changed. Subsequent command groups, +including reads, acquire the same lock and validate the token. This prevents a +command delayed by an SSH failure or an expired cluster lock from modifying a +volume after a newer operation has replaced it. The lock and token are shared by +all pools and storage definitions on the target host, so they also reject +commands from transactions that used another spelling of the server address. +Queries outside a transaction do not wait for the target lock. The runtime +directory must be owned by root with mode `0700`. Do not delete it or replace +its `lock` or `owner` files while commands can still be running. A stuck command +can block further changes; the backend does not break this lock to restore +availability. + +Choose the all-path outage policy for the workload. The default configuration +leaves `nvme-fast-io-fail-tmo` unset: transient outages can recover without +guest block errors, but an I/O request can stall until the controller is lost. +The default controller loss timeout is 600 seconds; the example instead +explicitly sets a fast I/O fail timeout of 30 seconds. Set a shorter fast I/O +fail timeout when the application prefers a prompt block error over a long +stall. This is a service-level decision, not a universally safer default. + +When every path of a node stays down past the controller loss timeout, the +kernel deletes the controllers and the multipath block devices of the storage on +that node, and queued and new I/O fails. A running guest keeps the removed +device open, so once the paths are back, stop and start the guest, for example +with `qm stop ` and `qm start `; a reboot inside the guest is not +enough. A target that rejects the connection, for example because of an unknown +Host NQN, a wrong key, or a listener that does not publish the subsystem yet +(see `nvme-portals`), makes the kernel delete the controller regardless of the +timeout. A controller loss timeout of `-1` keeps the devices, but combine it +with a finite fast I/O fail timeout. Otherwise guest I/O, and stopping or +migrating the guest, can block for the whole outage. + +Changes to `nvme-iopolicy`, `nvme-reconnect-delay`, `nvme-ctrl-loss-tmo`, and +`nvme-fast-io-fail-tmo` are applied to connected controllers on the next +storage activation. They do not affect an outage of all paths that is already +in progress. `nvme-keep-alive-tmo` and `nvme-nr-io-queues` only apply to new +controller connections, for example after a node reboot. + +Online resize of a running guest disk is not supported. Stop the VM before +resizing; the new size is detected when the block device is reopened. + +The backend does not support volume transfer streams through `pvesm export` +or `pvesm import`, including `raw+size`. Offline or remote migration workflows +that require these streams cannot transfer disks to or from this backend. +Shared-storage live migration remains supported: it does not transfer the disk +contents. + +Removing the storage definition never deletes volumes. The node that removes it +deactivates the storage, which deletes its controllers. Deactivation fails if +one of the namespaces is still in use locally, or if the controllers cannot be +deleted within 15 seconds (a deletion that is stuck in the kernel keeps the +task waiting until the kernel returns); the removal then logs a warning and the +node keeps its connections. Other nodes keep their connections. On each of +them, run `nvme disconnect --nqn ` or reboot the node. If no volumes +of the storage remain, the backend also tries to remove the subsystem and +unused host entries from the target. This cleanup is best effort: if it fails, +or if volumes remain, the target configuration is kept and a warning is logged. + +Storage Features +~~~~~~~~~~~~~~~~ + +.Storage features for backend `zfsnvme` +[width="100%",cols="m,m,3*d",options="header"] +|============================================================================== +|Content types |Image formats |Shared |Snapshots |Clones +|images |raw |yes |yes |yes +|============================================================================== diff --git a/pvesm.adoc b/pvesm.adoc index 5bd24b2..4ada3f0 100644 --- a/pvesm.adoc +++ b/pvesm.adoc @@ -439,6 +439,8 @@ See Also * link:/wiki/Storage:_ZFS_over_ISCSI[Storage: ZFS over ISCSI] +* link:/wiki/Storage:_ZFS_over_NVMe/TCP[Storage: ZFS over NVMe/TCP] + endif::wiki[] ifndef::wiki[] @@ -471,6 +473,8 @@ include::pve-storage-btrfs.adoc[] include::pve-storage-zfs.adoc[] +include::pve-storage-zfsnvme.adoc[] + ifdef::manvolnum[] include::pve-copyright.adoc[]