public inbox for pve-devel@lists.proxmox.com
 help / color / mirror / Atom feed
From: Tobias Fiebig <tobias@internet.wien>
To: Hannes Laimer <h.laimer@proxmox.com>, pve-devel@lists.proxmox.com
Subject: Re: SDN with BGP Fabric and IPv6 Only Underlay for VXLAN
Date: Wed, 07 Oct 2026 11:46:32 +0200	[thread overview]
Message-ID: <8de6cd729969eb4e318d4e76aa1137ddf47c7954.camel@internet.wien> (raw)
In-Reply-To: <acb5aba0-aa88-4a60-9d59-5880a0f196c4@proxmox.com>

Hello Hannes,


> yes, a patch for this specifically already exists [1], and with [2]
> also v6 fabrics should work.

Thanks, looking forward to it being merged.

In the context of the L3 fabric, i stumbled over two additional things.

1. BGP Pseudo interface MTU

If the L3 fabric is via links running on jumbos, and there is no
additional L2 network between the proxmox nodes, VM migration breaks as
soon as the BGP dummy interface for the loopback address in the fabric
is added on nodes, as that interface is created with an MTU of 1500 and
used as the default source address for contacting other nodes, i.e.,
while the other node originates packets for an MTU of, e.g., 9000,
those reach the initiating node, but cannot be forwarded to the (dummy)
interface where the selected source address resides. I am also not sure
why there are not PTB ICMPv6 packets originated that should fix this.
But on migration, copying over the memory state just hangs.

This can be fixed by either adjusting the MTU of the dummy interface to
match that of the actual links, or by creating a dedicated shared L2,
e.g., via another VXLAN, that is than addressed and explicitly
configured as the migration network.


2. Pure L3 Fabric + RFC8950 / IPv4 with IPv6 Nexthops

Beyond that, there is also the overall network concept. Not sure
whether this would be viable feature; But dumping it in here for
reference.

What I am essentially trying to build is a network (not only the PVE
components) that runs mostly IPv6 only, and has IPv4 as /32 routed to
GUA, handled via (i)BGP.

A description of the concept can be found here:
https://ripe88.ripe.net/presentations/15-ripe_88_v6_wg.pdf

The general advantages would be:
- No need for bridge devices
- No need for any L2 within the whole network; This also means that,
e.g., LACP is no longer needed on switches; Instead the L3 fabric is
BGP based (numbered or unnumbered); Also, with ECMP, traffic actually
balances etc.
- Most importantly: No ARP ;-)

Things I noticed that stand in the way there:

- Technically, nodes would have to be able to create a VM interface
that is not added to a bridge; Instead, it would just get DHCP/SLAAC
with a prefix selected from a configured covering prefix; The more
specific (automatically) selected for the VM would then just be
announced in the BGP fabric, and upon VM migration, the prefix would be
moved. The same could be done with v4, considering the VM OS supports
v6 NH for v4 default routes.
- There is no way to add communities via configurable routemaps in
proxmox atm; This makes it a bit more difficult to properly configure
filters when interacting with hosts speaking BGP outside of the proxmox
cluster; Same holds for import filters based on matching communities.
- The imported routes webinterface in the fabric view is a bit
overwhelmed when the nodes get a (v6) fulltable (questionable concept,
but... it can make sense ;-)) Pagination there would probably be
helpful (also, in general, for larger networks).

Happy to hear your thoughts; If this is the wrong place for discussing
this, please also let me know.

With best regards,
Tobias

-- 
My working day may not be your working day. Please do not feel obliged 
to reply to my email outside of your normal working hours.
-----------------------------------------------------------------
Univ.Prof. Dr.-Ing. Tobias Fiebig, E191-06 - Internet Infrastructures
TU Wien, Fakultät für Informatik, Treitl Str. 3, 1040 Wien
Room DE01 15 mobile phone: +31 616 80 98 99 mail: tobias@internet.wien



  reply	other threads:[~2026-10-07  9:46 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-06 14:46 SDN with BGP Fabric and IPv6 Only Underlay for VXLAN Tobias Fiebig
2026-10-06 15:19 ` Tobias Fiebig
2026-10-07  6:46   ` Hannes Laimer
2026-10-07  9:46     ` Tobias Fiebig [this message]
2026-10-07 10:37       ` Hannes Laimer

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=8de6cd729969eb4e318d4e76aa1137ddf47c7954.camel@internet.wien \
    --to=tobias@internet.wien \
    --cc=h.laimer@proxmox.com \
    --cc=pve-devel@lists.proxmox.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Service provided by Proxmox Server Solutions GmbH | Privacy | Legal