From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from gate001.proxmox.com (gate001.proxmox.com [45.144.208.40]) by lore.proxmox.com (Postfix) with ESMTPS id 7B6D71FF0C1 for ; Wed, 26 Aug 2026 11:30:18 +0200 (CEST) Received: from gate001.proxmox.com (localhost.localdomain [127.0.0.1]) by gate001.proxmox.com (Proxmox) with ESMTP id 525572161B; Wed, 26 Aug 2026 11:29:23 +0200 (CEST) ARC-Seal: i=1; a=rsa-sha256; t=1787665928; cv=none; d=google.com; s=arc-20260327; b=LG56KOrAOmWpFl6wkltwLXxfvG/WeqlAHdH02euSYEJ9fzz5FTGDmuMMcGvP96l6eg oMRqxui9874Gfenv5nchr7IJZPeWTpAFuNFuDytGobbdCTpM60vM3kSjFf8JYleUpfiA LeotpQ8/fwBxiFnB7xhVMFm1OXn+rZe7o1jhA/+pF7nMkkHnxuVmJn31kpXlDbcQJV6l EAGxlRkumQ6HwuOKBZv3cz6uZt+5tl5sZxuwiT5TLYF4v5ygpzzbE2Wl8+DUiC3tCJmu hxFeVNRoy9EJYs+TzvyxozR+YsAmdkqFg9gZTMnAdMs6fBE8FDiMFUs43oXgCe2q+VXS puqQ== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20260327; h=content-transfer-encoding:cc:to:subject:message-id:date:from :in-reply-to:references:mime-version:dkim-signature; bh=IzLJuDvzYeGEOL831XE8EDb4Fs8Tgb9PrGGDPQrwfEU=; fh=U76Nv9QJU7QU+cM588J3XqmDcQafw1LmdlINO4jFtVI=; b=JS3n0HbYiHVrnfYybqTux5CIYHAL2+sa6UMNIQ92A/PNoLqrWZMXew2OrY0ANsQmy0 yT6GLLkvLODLJwEoyEu4cejutv9SABYpyIgMQ7rAS1ZJhhwIVPNE9yJ/KnszXL7AAnF0 K4g95O2RZZSubR6lxapbsLplkWhoy+UuTzKH22piizqtKBuiNFPLnecwRtF0aNTmVpCe UM1j0oKOtYpHQPjMK37oJVZxzOl8AnQtJN6hCUskICUVmnjSk2idf5295zmf3HkHmJ1H S3d0SuHi8s1r7MYSHN/zA3L/e9GxOLkCRLu4FbfDFk/EnL7HKGA+WJncOs+qp/Bpxz3s gh0w==; darn=lists.proxmox.com ARC-Authentication-Results: i=1; mx.google.com; arc=none DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787665928; x=1788270728; darn=lists.proxmox.com; h=content-transfer-encoding:content-type:cc:to:subject:message-id :date:from:in-reply-to:references:mime-version:from:to:cc:subject :date:message-id:reply-to:content-type; bh=IzLJuDvzYeGEOL831XE8EDb4Fs8Tgb9PrGGDPQrwfEU=; b=dKfth0icViZl/v36K7KpI7H4Qbvw4rezvJlWGOk/8kMwk67TmEPTLSrt89vci/UC1T IvwMKLmbFMh61tUMpwTiU+IbJlVEmRyYaTyKRlt70T7QRPdS+mu619RDy45ggqTHtosw ZpmtqN/hrjrLX4Ka+0zjf5jwbc7jSZzRSz72DHrK4DYu+VSSQqbBBClSAh41dE8/IFeJ ndE5QDHOkx0YMEIAPETkeuEtHH2dLoJtSh8ec8EpkRuB7grgLcNBxMVupZYq93x7aOH1 i7g/rrDuGZcsOreNcW5YRPhMvpSU8++oSajVLs9exi1Rw5NFVN1L51hezWT7OqFNzbZ0 IMUQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787665928; x=1788270728; h=content-transfer-encoding:content-type:cc:to:subject:message-id :date:from:in-reply-to:references:mime-version:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=IzLJuDvzYeGEOL831XE8EDb4Fs8Tgb9PrGGDPQrwfEU=; b=KwIMU37crsvIriPHED3pDWh4WMUsZiJ0pqV/n75toPrDojEhk2PcUuNVDFbjISU4ob aoI28u39Br6wGB+YqBamgaRAZnDhjBUWfAMM/08J4L3QNXz2nouuxmMCytT5Oij/gqeL l7tvJLfQ5eDKxa9QVRZa0Ix5D/g2jP2optoFl+pcOu44kRFb+fyOHPVmKEyC8/u7u6+d ffrYAgSRv4LYjvnUvQi78vkS0aCrpkni4YdLr5N1kg3jEFri0ObJGO0d8eVXxu+mkBze CNtmcJtZBPFJStXcfiEFXFRiyXWlUS3K1sicKIecTs/oaoFmgY+62YgrkNKleZlJIRxj qs5Q== X-Gm-Message-State: AFuF++mtblLF1cLKAoH9ECPDZbioJsWW+N2rHScR18oSSxL2V0TepvSU TYu4jNehSeZtx5OUuAmwb9TiCAMCqB936we5Fhx3iMIDrJdkjS0fmNXjf9IN3WoIbFg+l6TYA4i eQxJzRfRCxf7/0CvwnSg9INnlRCGNI20= X-Gm-Gg: AR+sD12GLjJipQtgYdal2FvDKZpOC6vQQsS/uM9Xm0DJ+AOE/q/Qj2dbHsAKNzbURfF vTR+eEC8zS/RsqYfTXNM23dGFC1NGKEYg4p+Xkx8uUl3uGK9bSuoATU9iqcf0WWek6hYmUVKDdL RG/aNoC0mdBa70fOiAdDfRGBmIGxl+G+L9AdrZmMUaHsMTudyVVqgX4tYY7KBzRhidGEN8mWgiz 5ODvHmPds8Ujm9Hv8HU9se8ZYSZkiW2uvyznrDlgpZPKeU4XAntKNHbnsQTzVTj7+/7mtOIs+Pt hFDyLPyE5w330ILJymVxAuiIMIADefGiJR560GNF2lZvGct66yRfvZUFZeYWjOnkgBNC9cuq48r 3fMnMtj5qCJ43XU7avOZkq93eJfM= X-Received: by 2002:a05:690e:78e:b0:668:9ea4:775d with SMTP id 956f58d0204a3-66ce4b06530mr9531888d50.37.1787665927492; Tue, 25 Aug 2026 06:52:07 -0700 (PDT) MIME-Version: 1.0 References: <20260824170828.42821-1-Xy2462381442@gmail.com> In-Reply-To: From: Yao Xu Date: Tue, 25 Aug 2026 15:51:56 +0200 X-Gm-Features: AcwNN1VBTTYjF42-PQFRkRZb1zOtMafiNz8YPqJhqpidrMzYHimeWx6eq9lsS18 Message-ID: Subject: Re: [PATCH pve-docs 0/2] qm: pci passthrough: the AtomicOps caveat of the all-functions form To: Dominik Csapak Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-SPAM-LEVEL: Spam detection results: 1 DKIM_SIGNED 0.1 Message has a DKIM or DK signature, not necessarily valid DKIM_VALID -0.1 Message has at least one valid DKIM or DK signature DKIM_VALID_AU -0.1 Message has a valid DKIM or DK signature from author's domain DKIM_VALID_EF -0.1 Message has a valid DKIM or DK signature from envelope-from domain DMARC_PASS -0.1 DMARC pass policy FREEMAIL_ENVFROM_END_DIGIT 1 Envelope-from freemail username ends in digit FREEMAIL_FROM 0.001 Sender email is commonly abused enduser mail provider GB_FREEMAIL_NUM 0.75 Freemail spammy address RCVD_IN_DNSWL_NONE -0.0001 Sender listed at https://www.dnswl.org/, no trust SPF_HELO_NONE 0.001 SPF: HELO does not publish an SPF Record SPF_PASS -0.001 SPF: sender matches SPF record X-MailFrom: xy2462381442@gmail.com X-Mailman-Rule-Hits: nonmember-moderation X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; emergency; member-moderation Message-ID-Hash: YVGDWHUDDLR7YIBUNEGZULCCC2ZEMHIK X-Message-ID-Hash: YVGDWHUDDLR7YIBUNEGZULCCC2ZEMHIK X-Mailman-Approved-At: Wed, 26 Aug 2026 11:28:46 +0200 CC: pve-devel@lists.proxmox.com X-Mailman-Version: 3.3.10 Precedence: list List-Id: Proxmox VE development discussion List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: Hi Dominik, Thanks for looking at it. The CLA is signed and on record =E2=80=94 sent to office@proxmox.com and acknowledged on 2026-08-24. I left it off the cover letter on the assumption that pve-devel is the technical list; happy to be corrected. > I have a question: how is this handled on real hardware, since I > guess that works just fine there? (The cards in question > do have multiple functions...) It does work on real hardware, and the function count is not what decides it. The kernel side is pci_enable_atomic_ops_to_root() in drivers/pci/pci.c: it requires the device to be a PCIe endpoint, the root port's DEVCAP2 to advertise the requested completion widths, and every bridge on the path to route AtomicOps without blocking egress. It never reads the device's function number and never asks whether the slot is multifunction. A card with a GPU and an HDMI audio function gets atomics on bare metal exactly as a single-function card would. The restriction is QEMU's, and it is deliberate. It also has already been challenged upstream, for exactly the reason you give. In February 2026 AMD posted "vfio/pci: Add multifunction atomic ops support", which removed the multifunction guard and computed the intersection of the functions' capabilities instead. Their motivation was yours: "we have come up on more than one occasion where the topology of the bare metal was mimicked by VM's configuration ... from UX standpoint, the correct way is that user shouldn't think about it". Alex Williamson declined it: Back to the point, atomic ops routing is complicated, QEMU currently kicks anything beyond the trivial case back to the VM administrator. If the VM administrator doesn't want to think about it, analyze the host topology, create a compatible VM topology, and manually set appropriate atomic ops bits, then the burden probably needs to go in the direction of VM builders and management tools rather than pushed down into QEMU. QEMU doesn't have the visibility to determine host routing and is forced to work with the topology that's been specified. and closed with "I'm not convinced it's QEMU's job, or that QEMU is even capable of serving the intended goal here." https://lore.kernel.org/qemu-devel/8b3e30e6-3c3e-49ab-b9db-8296aaf819d1@a= pp.fastmail.com/ His reasoning is worth reading in full, because it is not the reason the code comment gives =E2=80=94 he says so himself: the restriction is less ab= out picking a common capability set than about device-to-device AtomicOps, which QEMU cannot reason about because it cannot see the host's routing, and because a guest multifunction package need not correspond to one on the host at all. > If yes, I'd rather have this reported as a bug on the QEMU side, > rather have a behavior documented that might change with any release. It has been reported, by the vendor, and declined six months ago. I can reopen it if you want, but I do not think re-litigating it is a good use of anyone's time, and on the release-stability point the guard is unchanged from v8.1.0 through v11.1.0. That is also why I sent this to pve-devel rather than only upstream. The decision QEMU is declining to make is which capability to advertise when a slot's functions disagree, and Proxmox is what composes the slot. > Also, the way the note is phrased makes it sound like this is something > everyone needs, while it's probably only relevant for some use cases. You are right, and your framing is better than mine. It is a narrow case: single-GPU inference is unaffected, and what breaks is multi-GPU collectives, which is where AtomicOps get used. I will send a v2 with the note scoped that way and the version stated, along the lines of: .Multi-GPU passthrough and PCIe AtomicOps [NOTE] =3D=3D=3D=3D Workloads that use PCIe AtomicOps =E2=80=94 in practice multi-GPU collectives such as ROCm's RCCL =E2=80=94 need the guest's virtual root= port to advertise completer support. QEMU (as of 11.1) adds it only for a single-function device, so passing every function of a card leaves it unadvertised, and the AMD driver logs `PCIE atomic ops is not supported`. If you need it, pass the function explicitly, for example ``hostpci0: 00:02.0,pcie=3Don`' on a `q35` machine. The host root port above the device must also support AtomicOp completion, which `lspci -vv` shows as `AtomicOpsCap: 32bit+ 64bit+`. =3D=3D=3D=3D > From my experience, in most situations, you want to pass the card > through as it is on the host, e.g. with functions. Then patch 2 should go. It changes the GPU example to a single function, and on your reading of what most users want that makes the default worse for them to fix a case the note already covers. I will drop it from v2 rather than argue for it. Verified on PVE 9.2.4 with QEMU 11.0.2, two RX 7900 XT in one q35 guest: with both cards passed as 0000:0b:00 and 0000:44:00 the guest root ports report AtomicOpsCap: 32bit- 64bit- and a two-rank RCCL all_reduce fails; appending .0 to both and changing nothing else gives 32bit+ 64bit+ and the same collective completes. Reverting reproduces the failure. Best regards Yao