From: Joaquin Varela <joaquinvarela@neatech.ar>
To: pve-devel@lists.proxmox.com
Subject: [PATCH storage v3 0/4] add ZFS over NVMe/TCP storage plugin
Date: Sun, 4 Oct 2026 21:26:03 -0300 [thread overview]
Message-ID: <20261005002609.571-1-joaquinvarela@neatech.ar> (raw)
Hi,
this is v3 of the zfsnvme backend: ZFS volumes on a remote Linux
target, exported with the kernel's nvmet over NVMe/TCP and used through
native NVMe multipath and DH-HMAC-CHAP on the nodes.
The scope is one remote Linux ZFS/nvmet target and its integration with
the PVE shared-storage API. This series does not implement target high
availability, pool ownership between target machines, or target fencing.
PVE remains responsible for VM placement and node fencing. The namespace
identity and reconnect behavior are intended to allow an independently
managed target to recover, but a multi-head target has not been qualified.
v2: https://lore.proxmox.com/pve-devel/cover.1785636979.git.joaquinvarela@neatech.ar/
Changes since v2, following Max's review:
1. The plugin inherits from PVE::Storage::Plugin directly. The ZFS
helpers it needs are private copies; volume names and their parsing
are the same as in the other ZFS plugins.
2. PVE::Storage::LunCmd::NVMET is gone; the nvmet code lives in the
plugin with an nvmet_ prefix.
3. There is no remote helper anymore. ZFS/configfs parsing,
storage-configuration validation and operation planning happen locally,
in functions that are tested on their own. Optional patch 3/4 adds a
fixed target-side transport guard, described below.
The target state is read in one call: 'zfs get' on the pool and its
direct children, and one 'find' over the nvmet configfs tree that
lists its objects, prints the attributes the plugin uses with 'grep'
and reads host keys only as 'sha256sum' digests. The first such read
of an operation under the target lock starts with
'modprobe nvmet_tcp'. Other reads are 'zfs get' and 'zfs list' like
in the other ZFS plugins, 'test -b' while a new zvol appears, and
'cat /proc/mounts' when that first locked read fails, to decide
whether configfs needs to be mounted. Changes are quoted simple
commands, joined with '&&' so that a step only runs after the
previous one succeeded, plus redirections into configfs attributes
and a 'dd | tee' pipe that writes the key from standard input,
without bash-isms. An export starts with 'test -b' on its zvol, and
each publish link repeats two read-only checks right before it links
the subsystem to a port: that the subsystem denies unknown hosts and
that every configured host has its ACL. A pmxcfs domain lock per
target serializes changes within the cluster. Patch 3/4 adds the only
shell control flow, see below.
4. opendir/readdir and glob were replaced by
PVE::File::dir_glob_foreach, nested where a sysfs walk spans several
directory levels, because dir_glob_regex returns only the first
match; IO::Socket::IP was replaced by PVE::Network::tcp_ping, and
IO::File by PVE::SysFSTools::file_write.
Other changes:
- Each storage needs its own NVMe/TCP listeners (address and port) on
the target. nvmet answers a connect to a listener that does not
publish the subsystem yet with DNR, and the node then deletes the
controller regardless of the controller loss timeout. After a target
restart the plugin restores one storage at a time, so storages that
shared a listener could lose their controllers. Adding, updating and
activating a storage now refuse a portal whose listener another
zfsnvme storage uses, also a disabled one, and a target address that
another storage reaches through a different server value, since the
server value names the lock per target. IPv6 addresses are compared
in canonical form.
- Nodes create each missing path with a single write to
/dev/nvme-fabrics in a short-lived child process
(PVE::Tools::run_fork_with_timeout) instead of running
'nvme connect --config'. With the Debian trixie versions tested here,
nvme-cli 2.13-2 and libnvme 1.13-2, the invocation with an explicit
host NQN and host ID re-issues connects to existing fabrics controllers
of that host, and its sysfs scan replaces the configured DH-HMAC-CHAP
host secret with the secret of an existing controller. On a node with
another NVMe/TCP controller
of the same host NQN, such as a second zfsnvme storage on a different
target, a new path can then be attempted with the other target's key.
The same happens with nvme-cli 2.16 and libnvme 1.16.2; nvme-cli 3.1
reworked 'nvme connect --config' and did not reproduce this behavior
in our tests. The direct write also keeps the key off command lines
and avoids a temporary nvme-cli configuration file. The key still
lives in the PVE private key file and the kernel attributes described
below. It uses sysopen rather than PVE::SysFSTools::file_write,
because the kernel returns the instance of the new controller on a
read of the descriptor that wrote the options. Controllers are
deleted and rescanned through sysfs. nvme-cli stays a dependency
because it creates /etc/nvme/hostnqn and /etc/nvme/hostid.
- The plugin restricts the DH-HMAC-CHAP secret attributes of the node
controllers, and of the target host entries it writes the key to, to
mode 0600. The kernel creates them world-readable, so on the nodes
this is best effort; the docs describe the remaining exposure.
- A correction to the v2 cover letter: the "verified 45.7 GB mixed-I/O
soak" ran for only 741 seconds and was a short checksum exercise; it
is not evidence of cache-independent readback or of crash durability.
The patches:
1/4 adds the backend.
2/4 adds the tests: the plugin against mocked host interfaces, and an
emulated target that checks every command the plugin plans and
replays single and double faults during each operation.
3/4 is optional, and a question for you. ssh does not stop a remote
command when the client gives up, so a delayed command of an
abandoned operation can still change the target after the next
operation started. This patch wraps every call made inside a locked
transaction, reads included, in a small, fixed POSIX sh guard: a
flock on a file below /run on the target, a per-transaction token
that every later command of the transaction checks, and a
compare-and-set of that token when a transaction starts. Read-only
calls outside a transaction do not use it. It excludes stale
control-plane commands; it does not fence hosts or guest I/O. It
adds two SSH round trips (owner observation and claim) to every
locked operation. Without it, the backend relies on the cluster lock
like ZFS over iSCSI does. Would you prefer to keep it, a smaller
variant with only the flock and the token check, or to drop it? 4/4
applies and passes its tests with or without 3/4.
4/4 is also a question: it overrides cluster_lock_storage so that the
shared storage lock waits up to 30 seconds instead of 10 when the
caller passes no timeout, since a target change over SSH can hold
it for several seconds. Explicit timeouts are unchanged. Would you
rather have this in the plugin, as a parameter in the core, or
should I reduce the SSH round trips per operation instead?
The series is much larger than v2: the plugin has about 3,300 lines,
because the local parsing and planning replace the remote helper, and
the tests about 7,800 (two test files and two fixtures). I can split
1/4 into the schema and pvesm, the target planning, the node connect
and the volume operations, and split or trim the test patch, whatever
is easier to review.
Testing: every patch passes 'make test' on its own, unprivileged and
with stock libpve-common-perl, and the Perl files are formatted with
proxmox-perltidy. On a nested two-node PVE 9 cluster (kernel 7.0.14)
with an Ubuntu 24.04 target (kernel 6.8), I ran the checks below with
a package built from the code of this series. The only exception is
the reproduction of the shared listener problem, which deliberately
used the code before the listener check:
- activation and DH-HMAC-CHAP on both nodes and both paths through the
new connect path: one read-write open of /dev/nvme-fabrics per path,
no nvme-cli process;
- allocation, snapshot, rollback, resize, template, linked clone and
free on both nodes;
- a node that already has another NVMe/TCP controller with the same host
NQN and a different key: the plugin connects with its own key, while a
single 'nvme connect --config' in the same situation issued five
connect writes and failed authentication;
- live migration in both directions under guest writes, with readback;
- the loss and restore of one path, and a 15-second loss of both paths,
under guest writes, without guest I/O errors;
- a target reboot under guest writes: the plugin rebuilt the target
configuration and the committed data verified;
- stopping an activation task while a connect was in flight, with stock
libpve-common-perl: no process was left behind and none of the six
cancellations left a controller; a controller created out of band, as
a cancelled connect can leave behind, is adopted by the next
activation;
- the controller loss behavior the docs describe: with a 20-second
controller loss timeout, the controllers and the multipath device were
removed and an open descriptor stayed dead, while -1 with a 5-second
fast I/O fail timeout failed I/O after about 15 seconds and kept the
device, which worked again once the paths were back;
- two storages that share a listener: with the code before the
listener check, after their configuration was removed from the
target, as a target restart does, nvmet rejected the reconnects of
the storage restored second with DNR, both nodes deleted its
controllers, and an open descriptor on its multipath device stayed
dead; with this series, adding the second storage on the same
listener is refused, and with a listener per storage both recovered
from the same removal, in either order, without deleting a
controller, while an open O_DIRECT descriptor kept reading the
written data, with one read blocked for about 5 seconds in each of
the two runs.
This is lab qualification, not hardware qualification.
I also saw Dietmar's RFC "add guided remote storage setup and SAN
visibility", which adds NVMe-oF transport and path status helpers to
Diskmanage. Once it is applied, the plugin can use them instead of its
own sysfs parsing:
https://lore.proxmox.com/pve-devel/20260731102156.3947857-1-dietmar@proxmox.com/
The documentation and the editor follow as replies to this cover
letter, as separate patches for pve-docs and pve-manager.
Joaquin Varela (4):
zfsnvme: add ZFS over NVMe/TCP storage plugin
test: add zfsnvme plugin tests
zfsnvme: fence target commands of abandoned transactions
zfsnvme: wait up to 30 seconds for the shared storage lock
debian/control | 1 +
src/PVE/CLI/pvesm.pm | 21 +-
src/PVE/Storage.pm | 2 +
src/PVE/Storage/Makefile | 1 +
src/PVE/Storage/Plugin.pm | 2 +-
src/PVE/Storage/ZFSNVMePlugin.pm | 3319 +++++++++++
src/test/run_plugin_tests.pl | 2 +
.../zfsnvme_fixtures/configfs_snapshot.txt | 38 +
src/test/zfsnvme_fixtures/zfs_inventory.txt | 15 +
src/test/zfsnvme_target_test.pm | 5299 +++++++++++++++++
src/test/zfsnvme_test.pm | 2429 ++++++++
11 files changed, 11126 insertions(+), 3 deletions(-)
create mode 100644 src/PVE/Storage/ZFSNVMePlugin.pm
create mode 100644 src/test/zfsnvme_fixtures/configfs_snapshot.txt
create mode 100644 src/test/zfsnvme_fixtures/zfs_inventory.txt
create mode 100644 src/test/zfsnvme_target_test.pm
create mode 100644 src/test/zfsnvme_test.pm
next reply other threads:[~2026-10-05 0:26 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-05 0:26 Joaquin Varela [this message]
2026-10-05 0:26 ` [PATCH storage v3 1/4] zfsnvme: add ZFS over NVMe/TCP storage plugin Joaquin Varela
2026-10-05 0:26 ` [PATCH storage v3 2/4] test: add zfsnvme plugin tests Joaquin Varela
2026-10-05 0:26 ` [PATCH storage v3 3/4] zfsnvme: fence target commands of abandoned transactions Joaquin Varela
2026-10-05 0:26 ` [PATCH storage v3 4/4] zfsnvme: wait up to 30 seconds for the shared storage lock Joaquin Varela
2026-10-05 0:26 ` [PATCH docs v3] storage: document ZFS over NVMe/TCP Joaquin Varela
2026-10-05 0:26 ` [PATCH manager v3] ui: storage: add ZFS over NVMe/TCP editor Joaquin Varela
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261005002609.571-1-joaquinvarela@neatech.ar \
--to=joaquinvarela@neatech.ar \
--cc=pve-devel@lists.proxmox.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox