public inbox for pve-devel@lists.proxmox.com
 help / color / mirror / Atom feed
From: Hannes Laimer <h.laimer@proxmox.com>
To: Tobias Fiebig <tobias@internet.wien>, pve-devel@lists.proxmox.com
Subject: Re: SDN with BGP Fabric and IPv6 Only Underlay for VXLAN
Date: Wed, 7 Oct 2026 12:37:59 +0200	[thread overview]
Message-ID: <a07dfbaf-bdb3-4db3-a52a-7ff749cded57@proxmox.com> (raw)
In-Reply-To: <8de6cd729969eb4e318d4e76aa1137ddf47c7954.camel@internet.wien>

On 2026-10-07 11:46, Tobias Fiebig wrote:
> Hello Hannes,
> 
> 
>> yes, a patch for this specifically already exists [1], and with [2]
>> also v6 fabrics should work.
> 
> Thanks, looking forward to it being merged.
> 
> In the context of the L3 fabric, i stumbled over two additional things.
> 
> 1. BGP Pseudo interface MTU
> 
> If the L3 fabric is via links running on jumbos, and there is no
> additional L2 network between the proxmox nodes, VM migration breaks as
> soon as the BGP dummy interface for the loopback address in the fabric
> is added on nodes, as that interface is created with an MTU of 1500 and
> used as the default source address for contacting other nodes, i.e.,
> while the other node originates packets for an MTU of, e.g., 9000,
> those reach the initiating node, but cannot be forwarded to the (dummy)
> interface where the selected source address resides. I am also not sure
> why there are not PTB ICMPv6 packets originated that should fix this.
> But on migration, copying over the memory state just hangs.
> 

i think a switch somewhere along the way would not emit those..

> This can be fixed by either adjusting the MTU of the dummy interface to
> match that of the actual links, or by creating a dedicated shared L2,
> e.g., via another VXLAN, that is than addressed and explicitly
> configured as the migration network.
> 

hmm, the dummy doesn't transport any traffic, it just terminates it, MTU
should not matter there.. (did not test this)

does
`ping -M do -s 8952 {dummy ip}`
work between all hosts?

> 
> 2. Pure L3 Fabric + RFC8950 / IPv4 with IPv6 Nexthops
> 

yes, funnily enough I do have a poc laying around for basically this(L3
routed zone with /32 routes for each guest), but there are some things
that need to be addressed first. most importantly, we have to have a
stable/reliable way to know the ip a guest is using, and for that we
need IPAM to be more general and not only DHCP scoped(for which i also
have a poc flying around somewhere)

the main bottleneck is me getting around to work on this, and of course
also someone getting around to review it :) technically there isn't
really something blocking this for us

> Beyond that, there is also the overall network concept. Not sure
> whether this would be viable feature; But dumping it in here for
> reference.
> 
> What I am essentially trying to build is a network (not only the PVE
> components) that runs mostly IPv6 only, and has IPv4 as /32 routed to
> GUA, handled via (i)BGP.
> 
> A description of the concept can be found here:
> https://ripe88.ripe.net/presentations/15-ripe_88_v6_wg.pdf
> 
> The general advantages would be:
> - No need for bridge devices
> - No need for any L2 within the whole network; This also means that,
> e.g., LACP is no longer needed on switches; Instead the L3 fabric is
> BGP based (numbered or unnumbered); Also, with ECMP, traffic actually
> balances etc.
> - Most importantly: No ARP ;-)
> 
> Things I noticed that stand in the way there:
> 
> - Technically, nodes would have to be able to create a VM interface
> that is not added to a bridge; Instead, it would just get DHCP/SLAAC
> with a prefix selected from a configured covering prefix; The more
> specific (automatically) selected for the VM would then just be
> announced in the BGP fabric, and upon VM migration, the prefix would be
> moved. The same could be done with v4, considering the VM OS supports
> v6 NH for v4 default routes.

the main difference to what you propose is that it would be SRv6 based.
technically encapsulation would not be needed, but having it in place
would allow for some nice additions in the future, that would otherwise
be not really possible. also the local bridge would stay, but
technically that could also be dropped, just that currently vnet means
bridge basically, so im not sure about that..

> - There is no way to add communities via configurable routemaps in
> proxmox atm; This makes it a bit more difficult to properly configure
> filters when interacting with hosts speaking BGP outside of the proxmox
> cluster; Same holds for import filters based on matching communities.
> - The imported routes webinterface in the fabric view is a bit
> overwhelmed when the nodes get a (v6) fulltable (questionable concept,
> but... it can make sense ;-)) Pagination there would probably be
> helpful (also, in general, for larger networks).
> 

your last two points sound like a reasonable addition, and we track
things like this at [1] as an `enhancement`, you can just file them
there with your rationale and use-case

> Happy to hear your thoughts; If this is the wrong place for discussing
> this, please also let me know.
> 

it's fine, pve-user[2] might be more fitting, since pve-devel is mostly
for stuff that involves code

> With best regards,
> Tobias
> 

[1] https://bugzilla.proxmox.com/
[2] https://lists.proxmox.com/postorius/lists/pve-user.lists.proxmox.com/





      reply	other threads:[~2026-10-07 10:38 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-06 14:46 SDN with BGP Fabric and IPv6 Only Underlay for VXLAN Tobias Fiebig
2026-10-06 15:19 ` Tobias Fiebig
2026-10-07  6:46   ` Hannes Laimer
2026-10-07  9:46     ` Tobias Fiebig
2026-10-07 10:37       ` Hannes Laimer [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=a07dfbaf-bdb3-4db3-a52a-7ff749cded57@proxmox.com \
    --to=h.laimer@proxmox.com \
    --cc=pve-devel@lists.proxmox.com \
    --cc=tobias@internet.wien \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Service provided by Proxmox Server Solutions GmbH | Privacy | Legal