public inbox for pve-devel@lists.proxmox.com
 help / color / mirror / Atom feed
* [PATCH pve-kernel] pick proposed fix for bnxt regression in 7.0.14-20+
@ 2026-10-05 11:47 Fabian Grünbichler
  2026-10-07  7:38 ` applied: " Fabian Grünbichler
  0 siblings, 1 reply; 2+ messages in thread
From: Fabian Grünbichler @ 2026-10-05 11:47 UTC (permalink / raw)
  To: pve-devel

picked from upstream lore, fixes a widely reported regression with bnxt NICs

Signed-off-by: Fabian Grünbichler <f.gruenbichler@proxmox.com>
---
positive feedback in the forum thread:

https://forum.proxmox.com/threads/proxmox-ve-9-2-21-bcm57412-bnxt_en-network-failure-after-upgrade-to-kernel-7-0-14-20-pve.186749/page-4

the test build there was identical to the patch here, modulo a proper commit
message

 ...mapping-length-for-padded-small-pack.patch | 114 ++++++++++++++++++
 1 file changed, 114 insertions(+)
 create mode 100644 patches/kernel/0210-bnxt_en-fix-DMA-mapping-length-for-padded-small-pack.patch

diff --git a/patches/kernel/0210-bnxt_en-fix-DMA-mapping-length-for-padded-small-pack.patch b/patches/kernel/0210-bnxt_en-fix-DMA-mapping-length-for-padded-small-pack.patch
new file mode 100644
index 0000000..dc25856
--- /dev/null
+++ b/patches/kernel/0210-bnxt_en-fix-DMA-mapping-length-for-padded-small-pack.patch
@@ -0,0 +1,114 @@
+From 0000000000000000000000000000000000000000 Mon Sep 17 00:00:00 2001
+From: Eric Dumazet <edumazet@kernel.org>
+Date: Mon, 5 Oct 2026 04:38:12 +0200
+Subject: [PATCH] bnxt_en: fix DMA mapping length for padded small packets
+
+Stefan Fleischmann reported Intel IOMMU DMA Read faults on BCM57412
+NetXtreme-E NICs when transmitting packets on VLAN/macvlan interfaces:
+
+  DMAR: [DMA Read NO_PASID] Request device [18:00.0] fault addr 0xfc499000
+        [fault reason 0x06] PTE Read access is not set
+  bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0xb4 0x41a} len: 0 due to firmware status: 0x2000001
+  ...
+  NETDEV WATCHDOG: eno1np0 (bnxt_en): transmit queue 0 timed out
+
+The fault address (0xfc499000) is on an exact 4KB page boundary,
+pointing to a DMA read buffer overrun.
+
+In bnxt_start_xmit(), packets smaller than BNXT_MIN_PKT_SIZE (52 bytes),
+such as 42-byte untagged ARP frames, are padded:
+
+    if (length < BNXT_MIN_PKT_SIZE) {
+        pad = BNXT_MIN_PKT_SIZE - length;
+        if (skb_pad(skb, pad))
+            goto tx_kick_pending;
+        length = BNXT_MIN_PKT_SIZE;
+    }
+
+    mapping = dma_map_single(&pdev->dev, skb->data, len, DMA_TO_DEVICE);
+    ...
+    dma_unmap_len_set(tx_buf, len, len);
+
+However, 'len' was initialized earlier to skb_headlen(skb) (e.g. 42 bytes)
+and is left unadjusted after padding. Consequently, dma_map_single() and
+dma_unmap_len_set() map and track only 42 bytes.
+
+Later, the hardware TX buffer descriptor is programmed with the padded length:
+
+    txbd->tx_bd_len_flags_type =
+        cpu_to_le32(((len + pad) << TX_BD_LEN_SHIFT) | flags |
+                    TX_BD_FLAGS_PACKET_END);
+
+The NIC DMA engine is thus instructed to read 52 bytes from a region where
+only 42 bytes were DMA-mapped. If skb->data ends near the boundary of a 4KB
+page (within 'pad' bytes of the next page), the hardware DMA read overruns
+into the unmapped adjacent page, triggering an IOMMU fault.
+
+This issue was exposed after commit 447cbe95ebb9 ("vlan: fix skb_under_panic
+and races when toggling HW VLAN offload") because reserving extra VLAN
+headroom rounded LL_RESERVED_SPACE from 48 up to 64 bytes, shifting skb->data
+offsets and potentially causing small frames to land right against page
+boundaries.
+
+Fix this by using skb_put_padto(skb, BNXT_MIN_PKT_SIZE) in the normal_tx
+path. This ensures skb->len and skb_headlen(skb) reflect the padded size so
+that dma_map_single() maps the full buffer and the descriptor length is
+consistent. This also removes the temporary 'pad' variable and masking logic.
+
+Fixes: c0c050c58d84 ("bnxt_en: New Broadcom ethernet driver.")
+Cc: stable@vger.kernel.org
+Reported-by: Stefan Fleischmann <sfle@kth.se>
+Closes: https://lore.kernel.org/netdev/20261004122616.56714cbd@nargothrond/
+Signed-off-by: Eric Dumazet <edumazet@google.com>
+Reviewed-by: Michael Chan <michael.chan@broadcom.com>
+Link: https://lore.kernel.org/all/20261005023812.130639-1-edumazet@kernel.org
+Signed-off-by: Fabian Grünbichler <f.gruenbichler@proxmox.com>
+---
+ drivers/net/ethernet/broadcom/bnxt/bnxt.c | 19 +++++++------------
+ 1 file changed, 7 insertions(+), 12 deletions(-)
+
+diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt.c b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
+index 60a47e8785d143c1340e9e3b48b6f6f1f3183356..835eaf019b7e032e4f716c0fed1831feba11bccb 100644
+--- a/drivers/net/ethernet/broadcom/bnxt/bnxt.c
++++ b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
+@@ -474,7 +474,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev)
+ 	struct netdev_queue *txq;
+ 	int i;
+ 	dma_addr_t mapping;
+-	unsigned int length, pad = 0;
++	unsigned int length;
+ 	u32 len, free_size, vlan_tag_flags, cfa_action, flags;
+ 	struct bnxt_ptp_cfg *ptp = bp->ptp_cfg;
+ 	struct pci_dev *pdev = bp->pdev;
+@@ -639,14 +639,12 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev)
+ 	}
+ 
+ normal_tx:
+-	if (length < BNXT_MIN_PKT_SIZE) {
+-		pad = BNXT_MIN_PKT_SIZE - length;
+-		if (skb_pad(skb, pad))
+-			/* SKB already freed. */
+-			goto tx_kick_pending;
+-		length = BNXT_MIN_PKT_SIZE;
++	if (skb_put_padto(skb, BNXT_MIN_PKT_SIZE)) {
++		/* SKB already freed. */
++		goto tx_kick_pending;
+ 	}
+-
++	length = skb->len;
++	len = skb_headlen(skb);
+ 	mapping = dma_map_single(&pdev->dev, skb->data, len, DMA_TO_DEVICE);
+ 
+ 	if (unlikely(dma_mapping_error(&pdev->dev, mapping)))
+@@ -729,10 +727,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev)
+ 		txbd->tx_bd_len_flags_type = cpu_to_le32(flags);
+ 	}
+ 
+-	flags &= ~TX_BD_LEN;
+-	txbd->tx_bd_len_flags_type =
+-		cpu_to_le32(((len + pad) << TX_BD_LEN_SHIFT) | flags |
+-			    TX_BD_FLAGS_PACKET_END);
++	txbd->tx_bd_len_flags_type |= cpu_to_le32(TX_BD_FLAGS_PACKET_END);
+ 
+ 	netdev_tx_sent_queue(txq, skb->len);
+ 
-- 
2.47.3





^ permalink raw reply related	[flat|nested] 2+ messages in thread

* applied: [PATCH pve-kernel] pick proposed fix for bnxt regression in 7.0.14-20+
  2026-10-05 11:47 [PATCH pve-kernel] pick proposed fix for bnxt regression in 7.0.14-20+ Fabian Grünbichler
@ 2026-10-07  7:38 ` Fabian Grünbichler
  0 siblings, 0 replies; 2+ messages in thread
From: Fabian Grünbichler @ 2026-10-07  7:38 UTC (permalink / raw)
  To: pve-devel

got applied, though there is a v3 that we probably want to replace it
with..

On October 5, 2026 1:47 pm, Fabian Grünbichler wrote:
> picked from upstream lore, fixes a widely reported regression with bnxt NICs
> 
> Signed-off-by: Fabian Grünbichler <f.gruenbichler@proxmox.com>
> ---
> positive feedback in the forum thread:
> 
> https://forum.proxmox.com/threads/proxmox-ve-9-2-21-bcm57412-bnxt_en-network-failure-after-upgrade-to-kernel-7-0-14-20-pve.186749/page-4
> 
> the test build there was identical to the patch here, modulo a proper commit
> message
> 
>  ...mapping-length-for-padded-small-pack.patch | 114 ++++++++++++++++++
>  1 file changed, 114 insertions(+)
>  create mode 100644 patches/kernel/0210-bnxt_en-fix-DMA-mapping-length-for-padded-small-pack.patch
> 
> diff --git a/patches/kernel/0210-bnxt_en-fix-DMA-mapping-length-for-padded-small-pack.patch b/patches/kernel/0210-bnxt_en-fix-DMA-mapping-length-for-padded-small-pack.patch
> new file mode 100644
> index 0000000..dc25856
> --- /dev/null
> +++ b/patches/kernel/0210-bnxt_en-fix-DMA-mapping-length-for-padded-small-pack.patch
> @@ -0,0 +1,114 @@
> +From 0000000000000000000000000000000000000000 Mon Sep 17 00:00:00 2001
> +From: Eric Dumazet <edumazet@kernel.org>
> +Date: Mon, 5 Oct 2026 04:38:12 +0200
> +Subject: [PATCH] bnxt_en: fix DMA mapping length for padded small packets
> +
> +Stefan Fleischmann reported Intel IOMMU DMA Read faults on BCM57412
> +NetXtreme-E NICs when transmitting packets on VLAN/macvlan interfaces:
> +
> +  DMAR: [DMA Read NO_PASID] Request device [18:00.0] fault addr 0xfc499000
> +        [fault reason 0x06] PTE Read access is not set
> +  bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0xb4 0x41a} len: 0 due to firmware status: 0x2000001
> +  ...
> +  NETDEV WATCHDOG: eno1np0 (bnxt_en): transmit queue 0 timed out
> +
> +The fault address (0xfc499000) is on an exact 4KB page boundary,
> +pointing to a DMA read buffer overrun.
> +
> +In bnxt_start_xmit(), packets smaller than BNXT_MIN_PKT_SIZE (52 bytes),
> +such as 42-byte untagged ARP frames, are padded:
> +
> +    if (length < BNXT_MIN_PKT_SIZE) {
> +        pad = BNXT_MIN_PKT_SIZE - length;
> +        if (skb_pad(skb, pad))
> +            goto tx_kick_pending;
> +        length = BNXT_MIN_PKT_SIZE;
> +    }
> +
> +    mapping = dma_map_single(&pdev->dev, skb->data, len, DMA_TO_DEVICE);
> +    ...
> +    dma_unmap_len_set(tx_buf, len, len);
> +
> +However, 'len' was initialized earlier to skb_headlen(skb) (e.g. 42 bytes)
> +and is left unadjusted after padding. Consequently, dma_map_single() and
> +dma_unmap_len_set() map and track only 42 bytes.
> +
> +Later, the hardware TX buffer descriptor is programmed with the padded length:
> +
> +    txbd->tx_bd_len_flags_type =
> +        cpu_to_le32(((len + pad) << TX_BD_LEN_SHIFT) | flags |
> +                    TX_BD_FLAGS_PACKET_END);
> +
> +The NIC DMA engine is thus instructed to read 52 bytes from a region where
> +only 42 bytes were DMA-mapped. If skb->data ends near the boundary of a 4KB
> +page (within 'pad' bytes of the next page), the hardware DMA read overruns
> +into the unmapped adjacent page, triggering an IOMMU fault.
> +
> +This issue was exposed after commit 447cbe95ebb9 ("vlan: fix skb_under_panic
> +and races when toggling HW VLAN offload") because reserving extra VLAN
> +headroom rounded LL_RESERVED_SPACE from 48 up to 64 bytes, shifting skb->data
> +offsets and potentially causing small frames to land right against page
> +boundaries.
> +
> +Fix this by using skb_put_padto(skb, BNXT_MIN_PKT_SIZE) in the normal_tx
> +path. This ensures skb->len and skb_headlen(skb) reflect the padded size so
> +that dma_map_single() maps the full buffer and the descriptor length is
> +consistent. This also removes the temporary 'pad' variable and masking logic.
> +
> +Fixes: c0c050c58d84 ("bnxt_en: New Broadcom ethernet driver.")
> +Cc: stable@vger.kernel.org
> +Reported-by: Stefan Fleischmann <sfle@kth.se>
> +Closes: https://lore.kernel.org/netdev/20261004122616.56714cbd@nargothrond/
> +Signed-off-by: Eric Dumazet <edumazet@google.com>
> +Reviewed-by: Michael Chan <michael.chan@broadcom.com>
> +Link: https://lore.kernel.org/all/20261005023812.130639-1-edumazet@kernel.org
> +Signed-off-by: Fabian Grünbichler <f.gruenbichler@proxmox.com>
> +---
> + drivers/net/ethernet/broadcom/bnxt/bnxt.c | 19 +++++++------------
> + 1 file changed, 7 insertions(+), 12 deletions(-)
> +
> +diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt.c b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
> +index 60a47e8785d143c1340e9e3b48b6f6f1f3183356..835eaf019b7e032e4f716c0fed1831feba11bccb 100644
> +--- a/drivers/net/ethernet/broadcom/bnxt/bnxt.c
> ++++ b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
> +@@ -474,7 +474,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev)
> + 	struct netdev_queue *txq;
> + 	int i;
> + 	dma_addr_t mapping;
> +-	unsigned int length, pad = 0;
> ++	unsigned int length;
> + 	u32 len, free_size, vlan_tag_flags, cfa_action, flags;
> + 	struct bnxt_ptp_cfg *ptp = bp->ptp_cfg;
> + 	struct pci_dev *pdev = bp->pdev;
> +@@ -639,14 +639,12 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev)
> + 	}
> + 
> + normal_tx:
> +-	if (length < BNXT_MIN_PKT_SIZE) {
> +-		pad = BNXT_MIN_PKT_SIZE - length;
> +-		if (skb_pad(skb, pad))
> +-			/* SKB already freed. */
> +-			goto tx_kick_pending;
> +-		length = BNXT_MIN_PKT_SIZE;
> ++	if (skb_put_padto(skb, BNXT_MIN_PKT_SIZE)) {
> ++		/* SKB already freed. */
> ++		goto tx_kick_pending;
> + 	}
> +-
> ++	length = skb->len;
> ++	len = skb_headlen(skb);
> + 	mapping = dma_map_single(&pdev->dev, skb->data, len, DMA_TO_DEVICE);
> + 
> + 	if (unlikely(dma_mapping_error(&pdev->dev, mapping)))
> +@@ -729,10 +727,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev)
> + 		txbd->tx_bd_len_flags_type = cpu_to_le32(flags);
> + 	}
> + 
> +-	flags &= ~TX_BD_LEN;
> +-	txbd->tx_bd_len_flags_type =
> +-		cpu_to_le32(((len + pad) << TX_BD_LEN_SHIFT) | flags |
> +-			    TX_BD_FLAGS_PACKET_END);
> ++	txbd->tx_bd_len_flags_type |= cpu_to_le32(TX_BD_FLAGS_PACKET_END);
> + 
> + 	netdev_tx_sent_queue(txq, skb->len);
> + 
> -- 
> 2.47.3
> 
> 
> 
> 
> 
> 




^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-10-07  7:38 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-05 11:47 [PATCH pve-kernel] pick proposed fix for bnxt regression in 7.0.14-20+ Fabian Grünbichler
2026-10-07  7:38 ` applied: " Fabian Grünbichler

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Service provided by Proxmox Server Solutions GmbH | Privacy | Legal