public inbox for pve-devel@lists.proxmox.com
 help / color / mirror / Atom feed
From: Joaquin Varela <joaquinvarela@neatech.ar>
To: pve-devel@lists.proxmox.com
Subject: [PATCH storage v3 0/4] add ZFS over NVMe/TCP storage plugin
Date: Sun,  4 Oct 2026 21:26:03 -0300	[thread overview]
Message-ID: <20261005002609.571-1-joaquinvarela@neatech.ar> (raw)

Hi,

this is v3 of the zfsnvme backend: ZFS volumes on a remote Linux
target, exported with the kernel's nvmet over NVMe/TCP and used through
native NVMe multipath and DH-HMAC-CHAP on the nodes.

The scope is one remote Linux ZFS/nvmet target and its integration with
the PVE shared-storage API. This series does not implement target high
availability, pool ownership between target machines, or target fencing.
PVE remains responsible for VM placement and node fencing. The namespace
identity and reconnect behavior are intended to allow an independently
managed target to recover, but a multi-head target has not been qualified.

v2: https://lore.proxmox.com/pve-devel/cover.1785636979.git.joaquinvarela@neatech.ar/

Changes since v2, following Max's review:

1. The plugin inherits from PVE::Storage::Plugin directly. The ZFS
   helpers it needs are private copies; volume names and their parsing
   are the same as in the other ZFS plugins.
2. PVE::Storage::LunCmd::NVMET is gone; the nvmet code lives in the
   plugin with an nvmet_ prefix.
3. There is no remote helper anymore. ZFS/configfs parsing,
   storage-configuration validation and operation planning happen locally,
   in functions that are tested on their own. Optional patch 3/4 adds a
   fixed target-side transport guard, described below.
   The target state is read in one call: 'zfs get' on the pool and its
   direct children, and one 'find' over the nvmet configfs tree that
   lists its objects, prints the attributes the plugin uses with 'grep'
   and reads host keys only as 'sha256sum' digests. The first such read
   of an operation under the target lock starts with
   'modprobe nvmet_tcp'. Other reads are 'zfs get' and 'zfs list' like
   in the other ZFS plugins, 'test -b' while a new zvol appears, and
   'cat /proc/mounts' when that first locked read fails, to decide
   whether configfs needs to be mounted. Changes are quoted simple
   commands, joined with '&&' so that a step only runs after the
   previous one succeeded, plus redirections into configfs attributes
   and a 'dd | tee' pipe that writes the key from standard input,
   without bash-isms. An export starts with 'test -b' on its zvol, and
   each publish link repeats two read-only checks right before it links
   the subsystem to a port: that the subsystem denies unknown hosts and
   that every configured host has its ACL. A pmxcfs domain lock per
   target serializes changes within the cluster. Patch 3/4 adds the only
   shell control flow, see below.
4. opendir/readdir and glob were replaced by
   PVE::File::dir_glob_foreach, nested where a sysfs walk spans several
   directory levels, because dir_glob_regex returns only the first
   match; IO::Socket::IP was replaced by PVE::Network::tcp_ping, and
   IO::File by PVE::SysFSTools::file_write.

Other changes:

- Each storage needs its own NVMe/TCP listeners (address and port) on
  the target. nvmet answers a connect to a listener that does not
  publish the subsystem yet with DNR, and the node then deletes the
  controller regardless of the controller loss timeout. After a target
  restart the plugin restores one storage at a time, so storages that
  shared a listener could lose their controllers. Adding, updating and
  activating a storage now refuse a portal whose listener another
  zfsnvme storage uses, also a disabled one, and a target address that
  another storage reaches through a different server value, since the
  server value names the lock per target. IPv6 addresses are compared
  in canonical form.
- Nodes create each missing path with a single write to
  /dev/nvme-fabrics in a short-lived child process
  (PVE::Tools::run_fork_with_timeout) instead of running
  'nvme connect --config'. With the Debian trixie versions tested here,
  nvme-cli 2.13-2 and libnvme 1.13-2, the invocation with an explicit
  host NQN and host ID re-issues connects to existing fabrics controllers
  of that host, and its sysfs scan replaces the configured DH-HMAC-CHAP
  host secret with the secret of an existing controller. On a node with
  another NVMe/TCP controller
  of the same host NQN, such as a second zfsnvme storage on a different
  target, a new path can then be attempted with the other target's key.
  The same happens with nvme-cli 2.16 and libnvme 1.16.2; nvme-cli 3.1
  reworked 'nvme connect --config' and did not reproduce this behavior
  in our tests. The direct write also keeps the key off command lines
  and avoids a temporary nvme-cli configuration file. The key still
  lives in the PVE private key file and the kernel attributes described
  below. It uses sysopen rather than PVE::SysFSTools::file_write,
  because the kernel returns the instance of the new controller on a
  read of the descriptor that wrote the options. Controllers are
  deleted and rescanned through sysfs. nvme-cli stays a dependency
  because it creates /etc/nvme/hostnqn and /etc/nvme/hostid.
- The plugin restricts the DH-HMAC-CHAP secret attributes of the node
  controllers, and of the target host entries it writes the key to, to
  mode 0600. The kernel creates them world-readable, so on the nodes
  this is best effort; the docs describe the remaining exposure.
- A correction to the v2 cover letter: the "verified 45.7 GB mixed-I/O
  soak" ran for only 741 seconds and was a short checksum exercise; it
  is not evidence of cache-independent readback or of crash durability.

The patches:

1/4 adds the backend.

2/4 adds the tests: the plugin against mocked host interfaces, and an
    emulated target that checks every command the plugin plans and
    replays single and double faults during each operation.

3/4 is optional, and a question for you. ssh does not stop a remote
    command when the client gives up, so a delayed command of an
    abandoned operation can still change the target after the next
    operation started. This patch wraps every call made inside a locked
    transaction, reads included, in a small, fixed POSIX sh guard: a
    flock on a file below /run on the target, a per-transaction token
    that every later command of the transaction checks, and a
    compare-and-set of that token when a transaction starts. Read-only
    calls outside a transaction do not use it. It excludes stale
    control-plane commands; it does not fence hosts or guest I/O. It
    adds two SSH round trips (owner observation and claim) to every
    locked operation. Without it, the backend relies on the cluster lock
    like ZFS over iSCSI does. Would you prefer to keep it, a smaller
    variant with only the flock and the token check, or to drop it? 4/4
    applies and passes its tests with or without 3/4.

4/4 is also a question: it overrides cluster_lock_storage so that the
    shared storage lock waits up to 30 seconds instead of 10 when the
    caller passes no timeout, since a target change over SSH can hold
    it for several seconds. Explicit timeouts are unchanged. Would you
    rather have this in the plugin, as a parameter in the core, or
    should I reduce the SSH round trips per operation instead?

The series is much larger than v2: the plugin has about 3,300 lines,
because the local parsing and planning replace the remote helper, and
the tests about 7,800 (two test files and two fixtures). I can split
1/4 into the schema and pvesm, the target planning, the node connect
and the volume operations, and split or trim the test patch, whatever
is easier to review.

Testing: every patch passes 'make test' on its own, unprivileged and
with stock libpve-common-perl, and the Perl files are formatted with
proxmox-perltidy. On a nested two-node PVE 9 cluster (kernel 7.0.14)
with an Ubuntu 24.04 target (kernel 6.8), I ran the checks below with
a package built from the code of this series. The only exception is
the reproduction of the shared listener problem, which deliberately
used the code before the listener check:
- activation and DH-HMAC-CHAP on both nodes and both paths through the
  new connect path: one read-write open of /dev/nvme-fabrics per path,
  no nvme-cli process;
- allocation, snapshot, rollback, resize, template, linked clone and
  free on both nodes;
- a node that already has another NVMe/TCP controller with the same host
  NQN and a different key: the plugin connects with its own key, while a
  single 'nvme connect --config' in the same situation issued five
  connect writes and failed authentication;
- live migration in both directions under guest writes, with readback;
- the loss and restore of one path, and a 15-second loss of both paths,
  under guest writes, without guest I/O errors;
- a target reboot under guest writes: the plugin rebuilt the target
  configuration and the committed data verified;
- stopping an activation task while a connect was in flight, with stock
  libpve-common-perl: no process was left behind and none of the six
  cancellations left a controller; a controller created out of band, as
  a cancelled connect can leave behind, is adopted by the next
  activation;
- the controller loss behavior the docs describe: with a 20-second
  controller loss timeout, the controllers and the multipath device were
  removed and an open descriptor stayed dead, while -1 with a 5-second
  fast I/O fail timeout failed I/O after about 15 seconds and kept the
  device, which worked again once the paths were back;
- two storages that share a listener: with the code before the
  listener check, after their configuration was removed from the
  target, as a target restart does, nvmet rejected the reconnects of
  the storage restored second with DNR, both nodes deleted its
  controllers, and an open descriptor on its multipath device stayed
  dead; with this series, adding the second storage on the same
  listener is refused, and with a listener per storage both recovered
  from the same removal, in either order, without deleting a
  controller, while an open O_DIRECT descriptor kept reading the
  written data, with one read blocked for about 5 seconds in each of
  the two runs.
This is lab qualification, not hardware qualification.

I also saw Dietmar's RFC "add guided remote storage setup and SAN
visibility", which adds NVMe-oF transport and path status helpers to
Diskmanage. Once it is applied, the plugin can use them instead of its
own sysfs parsing:
https://lore.proxmox.com/pve-devel/20260731102156.3947857-1-dietmar@proxmox.com/

The documentation and the editor follow as replies to this cover
letter, as separate patches for pve-docs and pve-manager.

Joaquin Varela (4):
  zfsnvme: add ZFS over NVMe/TCP storage plugin
  test: add zfsnvme plugin tests
  zfsnvme: fence target commands of abandoned transactions
  zfsnvme: wait up to 30 seconds for the shared storage lock

 debian/control                                |    1 +
 src/PVE/CLI/pvesm.pm                          |   21 +-
 src/PVE/Storage.pm                            |    2 +
 src/PVE/Storage/Makefile                      |    1 +
 src/PVE/Storage/Plugin.pm                     |    2 +-
 src/PVE/Storage/ZFSNVMePlugin.pm              | 3319 +++++++++++
 src/test/run_plugin_tests.pl                  |    2 +
 .../zfsnvme_fixtures/configfs_snapshot.txt    |   38 +
 src/test/zfsnvme_fixtures/zfs_inventory.txt   |   15 +
 src/test/zfsnvme_target_test.pm               | 5299 +++++++++++++++++
 src/test/zfsnvme_test.pm                      | 2429 ++++++++
 11 files changed, 11126 insertions(+), 3 deletions(-)
 create mode 100644 src/PVE/Storage/ZFSNVMePlugin.pm
 create mode 100644 src/test/zfsnvme_fixtures/configfs_snapshot.txt
 create mode 100644 src/test/zfsnvme_fixtures/zfs_inventory.txt
 create mode 100644 src/test/zfsnvme_target_test.pm
 create mode 100644 src/test/zfsnvme_test.pm




             reply	other threads:[~2026-10-05  0:26 UTC|newest]

Thread overview: 7+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-05  0:26 Joaquin Varela [this message]
2026-10-05  0:26 ` [PATCH storage v3 1/4] zfsnvme: add ZFS over NVMe/TCP storage plugin Joaquin Varela
2026-10-05  0:26 ` [PATCH storage v3 2/4] test: add zfsnvme plugin tests Joaquin Varela
2026-10-05  0:26 ` [PATCH storage v3 3/4] zfsnvme: fence target commands of abandoned transactions Joaquin Varela
2026-10-05  0:26 ` [PATCH storage v3 4/4] zfsnvme: wait up to 30 seconds for the shared storage lock Joaquin Varela
2026-10-05  0:26 ` [PATCH docs v3] storage: document ZFS over NVMe/TCP Joaquin Varela
2026-10-05  0:26 ` [PATCH manager v3] ui: storage: add ZFS over NVMe/TCP editor Joaquin Varela

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261005002609.571-1-joaquinvarela@neatech.ar \
    --to=joaquinvarela@neatech.ar \
    --cc=pve-devel@lists.proxmox.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Service provided by Proxmox Server Solutions GmbH | Privacy | Legal