From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from gate001.proxmox.com (gate001.proxmox.com [IPv6:2a0f:8001:1:32::40]) by lore.proxmox.com (Postfix) with ESMTPS id 772931FF09C for ; Mon, 05 Oct 2026 13:47:33 +0200 (CEST) Received: from gate001.proxmox.com (localhost.localdomain [127.0.0.1]) by gate001.proxmox.com (Proxmox) with ESMTP id C31A621512; Mon, 05 Oct 2026 13:47:32 +0200 (CEST) From: =?UTF-8?q?Fabian=20Gr=C3=BCnbichler?= To: pve-devel@lists.proxmox.com Subject: [PATCH pve-kernel] pick proposed fix for bnxt regression in 7.0.14-20+ Date: Mon, 5 Oct 2026 13:47:11 +0200 Message-ID: <20261005114716.1338299-1-f.gruenbichler@proxmox.com> X-Mailer: git-send-email 2.47.3 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit X-Bm-Milter-Handled: 55990f41-d878-4baa-be0a-ee34c49e34d2 X-Bm-Transport-Timestamp: 1791200837039 X-SPAM-LEVEL: Spam detection results: 0 AWL 0.738 Adjusted score from AWL reputation of From: address DMARC_MISSING 0.1 Missing DMARC policy KAM_DMARC_STATUS 0.01 Test Rule for DKIM or SPF Failure with Strict Alignment (newer systems) RCVD_IN_DNSWL_MED -2.3 Sender listed at https://www.dnswl.org/, medium trust SPF_HELO_NONE 0.001 SPF: HELO does not publish an SPF Record SPF_PASS -0.001 SPF: sender matches SPF record Message-ID-Hash: ZJ2Y2IBSWYFKCHDXJSLVDYDE3W535S5S X-Message-ID-Hash: ZJ2Y2IBSWYFKCHDXJSLVDYDE3W535S5S X-MailFrom: f.gruenbichler@proxmox.com X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; emergency; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header X-Mailman-Version: 3.3.10 Precedence: list List-Id: Proxmox VE development discussion List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: picked from upstream lore, fixes a widely reported regression with bnxt NICs Signed-off-by: Fabian Grünbichler --- positive feedback in the forum thread: https://forum.proxmox.com/threads/proxmox-ve-9-2-21-bcm57412-bnxt_en-network-failure-after-upgrade-to-kernel-7-0-14-20-pve.186749/page-4 the test build there was identical to the patch here, modulo a proper commit message ...mapping-length-for-padded-small-pack.patch | 114 ++++++++++++++++++ 1 file changed, 114 insertions(+) create mode 100644 patches/kernel/0210-bnxt_en-fix-DMA-mapping-length-for-padded-small-pack.patch diff --git a/patches/kernel/0210-bnxt_en-fix-DMA-mapping-length-for-padded-small-pack.patch b/patches/kernel/0210-bnxt_en-fix-DMA-mapping-length-for-padded-small-pack.patch new file mode 100644 index 0000000..dc25856 --- /dev/null +++ b/patches/kernel/0210-bnxt_en-fix-DMA-mapping-length-for-padded-small-pack.patch @@ -0,0 +1,114 @@ +From 0000000000000000000000000000000000000000 Mon Sep 17 00:00:00 2001 +From: Eric Dumazet +Date: Mon, 5 Oct 2026 04:38:12 +0200 +Subject: [PATCH] bnxt_en: fix DMA mapping length for padded small packets + +Stefan Fleischmann reported Intel IOMMU DMA Read faults on BCM57412 +NetXtreme-E NICs when transmitting packets on VLAN/macvlan interfaces: + + DMAR: [DMA Read NO_PASID] Request device [18:00.0] fault addr 0xfc499000 + [fault reason 0x06] PTE Read access is not set + bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0xb4 0x41a} len: 0 due to firmware status: 0x2000001 + ... + NETDEV WATCHDOG: eno1np0 (bnxt_en): transmit queue 0 timed out + +The fault address (0xfc499000) is on an exact 4KB page boundary, +pointing to a DMA read buffer overrun. + +In bnxt_start_xmit(), packets smaller than BNXT_MIN_PKT_SIZE (52 bytes), +such as 42-byte untagged ARP frames, are padded: + + if (length < BNXT_MIN_PKT_SIZE) { + pad = BNXT_MIN_PKT_SIZE - length; + if (skb_pad(skb, pad)) + goto tx_kick_pending; + length = BNXT_MIN_PKT_SIZE; + } + + mapping = dma_map_single(&pdev->dev, skb->data, len, DMA_TO_DEVICE); + ... + dma_unmap_len_set(tx_buf, len, len); + +However, 'len' was initialized earlier to skb_headlen(skb) (e.g. 42 bytes) +and is left unadjusted after padding. Consequently, dma_map_single() and +dma_unmap_len_set() map and track only 42 bytes. + +Later, the hardware TX buffer descriptor is programmed with the padded length: + + txbd->tx_bd_len_flags_type = + cpu_to_le32(((len + pad) << TX_BD_LEN_SHIFT) | flags | + TX_BD_FLAGS_PACKET_END); + +The NIC DMA engine is thus instructed to read 52 bytes from a region where +only 42 bytes were DMA-mapped. If skb->data ends near the boundary of a 4KB +page (within 'pad' bytes of the next page), the hardware DMA read overruns +into the unmapped adjacent page, triggering an IOMMU fault. + +This issue was exposed after commit 447cbe95ebb9 ("vlan: fix skb_under_panic +and races when toggling HW VLAN offload") because reserving extra VLAN +headroom rounded LL_RESERVED_SPACE from 48 up to 64 bytes, shifting skb->data +offsets and potentially causing small frames to land right against page +boundaries. + +Fix this by using skb_put_padto(skb, BNXT_MIN_PKT_SIZE) in the normal_tx +path. This ensures skb->len and skb_headlen(skb) reflect the padded size so +that dma_map_single() maps the full buffer and the descriptor length is +consistent. This also removes the temporary 'pad' variable and masking logic. + +Fixes: c0c050c58d84 ("bnxt_en: New Broadcom ethernet driver.") +Cc: stable@vger.kernel.org +Reported-by: Stefan Fleischmann +Closes: https://lore.kernel.org/netdev/20261004122616.56714cbd@nargothrond/ +Signed-off-by: Eric Dumazet +Reviewed-by: Michael Chan +Link: https://lore.kernel.org/all/20261005023812.130639-1-edumazet@kernel.org +Signed-off-by: Fabian Grünbichler +--- + drivers/net/ethernet/broadcom/bnxt/bnxt.c | 19 +++++++------------ + 1 file changed, 7 insertions(+), 12 deletions(-) + +diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt.c b/drivers/net/ethernet/broadcom/bnxt/bnxt.c +index 60a47e8785d143c1340e9e3b48b6f6f1f3183356..835eaf019b7e032e4f716c0fed1831feba11bccb 100644 +--- a/drivers/net/ethernet/broadcom/bnxt/bnxt.c ++++ b/drivers/net/ethernet/broadcom/bnxt/bnxt.c +@@ -474,7 +474,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev) + struct netdev_queue *txq; + int i; + dma_addr_t mapping; +- unsigned int length, pad = 0; ++ unsigned int length; + u32 len, free_size, vlan_tag_flags, cfa_action, flags; + struct bnxt_ptp_cfg *ptp = bp->ptp_cfg; + struct pci_dev *pdev = bp->pdev; +@@ -639,14 +639,12 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev) + } + + normal_tx: +- if (length < BNXT_MIN_PKT_SIZE) { +- pad = BNXT_MIN_PKT_SIZE - length; +- if (skb_pad(skb, pad)) +- /* SKB already freed. */ +- goto tx_kick_pending; +- length = BNXT_MIN_PKT_SIZE; ++ if (skb_put_padto(skb, BNXT_MIN_PKT_SIZE)) { ++ /* SKB already freed. */ ++ goto tx_kick_pending; + } +- ++ length = skb->len; ++ len = skb_headlen(skb); + mapping = dma_map_single(&pdev->dev, skb->data, len, DMA_TO_DEVICE); + + if (unlikely(dma_mapping_error(&pdev->dev, mapping))) +@@ -729,10 +727,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev) + txbd->tx_bd_len_flags_type = cpu_to_le32(flags); + } + +- flags &= ~TX_BD_LEN; +- txbd->tx_bd_len_flags_type = +- cpu_to_le32(((len + pad) << TX_BD_LEN_SHIFT) | flags | +- TX_BD_FLAGS_PACKET_END); ++ txbd->tx_bd_len_flags_type |= cpu_to_le32(TX_BD_FLAGS_PACKET_END); + + netdev_tx_sent_queue(txq, skb->len); + -- 2.47.3