From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from gate001.proxmox.com (gate001.proxmox.com [45.144.208.40]) by lore.proxmox.com (Postfix) with ESMTPS id 2B1271FF0AB for ; Wed, 07 Oct 2026 12:38:13 +0200 (CEST) Received: from gate001.proxmox.com (localhost.localdomain [127.0.0.1]) by gate001.proxmox.com (Proxmox) with ESMTP id EE51A2125D; Wed, 07 Oct 2026 12:38:12 +0200 (CEST) Message-ID: Date: Wed, 7 Oct 2026 12:37:59 +0200 MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: SDN with BGP Fabric and IPv6 Only Underlay for VXLAN To: Tobias Fiebig , pve-devel@lists.proxmox.com References: <6a01449fcaee735648dbd8424d64c737a0428c9f.camel@internet.wien> <870c6dac7eb6683febc6bedaf679c500315216bd.camel@internet.wien> <8de6cd729969eb4e318d4e76aa1137ddf47c7954.camel@internet.wien> From: Hannes Laimer Content-Language: en-US In-Reply-To: <8de6cd729969eb4e318d4e76aa1137ddf47c7954.camel@internet.wien> Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 7bit X-Bm-Milter-Handled: 55990f41-d878-4baa-be0a-ee34c49e34d2 X-Bm-Transport-Timestamp: 1791369479756 X-SPAM-LEVEL: Spam detection results: 0 AWL 0.508 Adjusted score from AWL reputation of From: address DMARC_MISSING 0.1 Missing DMARC policy KAM_DMARC_STATUS 0.01 Test Rule for DKIM or SPF Failure with Strict Alignment (newer systems) RCVD_IN_DNSWL_MED -2.3 Sender listed at https://www.dnswl.org/, medium trust SPF_HELO_NONE 0.001 SPF: HELO does not publish an SPF Record SPF_PASS -0.001 SPF: sender matches SPF record Message-ID-Hash: F6QPJAAPE2JWFOJCK72Q3ZUS4CH5XBWC X-Message-ID-Hash: F6QPJAAPE2JWFOJCK72Q3ZUS4CH5XBWC X-MailFrom: h.laimer@proxmox.com X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; emergency; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header X-Mailman-Version: 3.3.10 Precedence: list List-Id: Proxmox VE development discussion List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: On 2026-10-07 11:46, Tobias Fiebig wrote: > Hello Hannes, > > >> yes, a patch for this specifically already exists [1], and with [2] >> also v6 fabrics should work. > > Thanks, looking forward to it being merged. > > In the context of the L3 fabric, i stumbled over two additional things. > > 1. BGP Pseudo interface MTU > > If the L3 fabric is via links running on jumbos, and there is no > additional L2 network between the proxmox nodes, VM migration breaks as > soon as the BGP dummy interface for the loopback address in the fabric > is added on nodes, as that interface is created with an MTU of 1500 and > used as the default source address for contacting other nodes, i.e., > while the other node originates packets for an MTU of, e.g., 9000, > those reach the initiating node, but cannot be forwarded to the (dummy) > interface where the selected source address resides. I am also not sure > why there are not PTB ICMPv6 packets originated that should fix this. > But on migration, copying over the memory state just hangs. > i think a switch somewhere along the way would not emit those.. > This can be fixed by either adjusting the MTU of the dummy interface to > match that of the actual links, or by creating a dedicated shared L2, > e.g., via another VXLAN, that is than addressed and explicitly > configured as the migration network. > hmm, the dummy doesn't transport any traffic, it just terminates it, MTU should not matter there.. (did not test this) does `ping -M do -s 8952 {dummy ip}` work between all hosts? > > 2. Pure L3 Fabric + RFC8950 / IPv4 with IPv6 Nexthops > yes, funnily enough I do have a poc laying around for basically this(L3 routed zone with /32 routes for each guest), but there are some things that need to be addressed first. most importantly, we have to have a stable/reliable way to know the ip a guest is using, and for that we need IPAM to be more general and not only DHCP scoped(for which i also have a poc flying around somewhere) the main bottleneck is me getting around to work on this, and of course also someone getting around to review it :) technically there isn't really something blocking this for us > Beyond that, there is also the overall network concept. Not sure > whether this would be viable feature; But dumping it in here for > reference. > > What I am essentially trying to build is a network (not only the PVE > components) that runs mostly IPv6 only, and has IPv4 as /32 routed to > GUA, handled via (i)BGP. > > A description of the concept can be found here: > https://ripe88.ripe.net/presentations/15-ripe_88_v6_wg.pdf > > The general advantages would be: > - No need for bridge devices > - No need for any L2 within the whole network; This also means that, > e.g., LACP is no longer needed on switches; Instead the L3 fabric is > BGP based (numbered or unnumbered); Also, with ECMP, traffic actually > balances etc. > - Most importantly: No ARP ;-) > > Things I noticed that stand in the way there: > > - Technically, nodes would have to be able to create a VM interface > that is not added to a bridge; Instead, it would just get DHCP/SLAAC > with a prefix selected from a configured covering prefix; The more > specific (automatically) selected for the VM would then just be > announced in the BGP fabric, and upon VM migration, the prefix would be > moved. The same could be done with v4, considering the VM OS supports > v6 NH for v4 default routes. the main difference to what you propose is that it would be SRv6 based. technically encapsulation would not be needed, but having it in place would allow for some nice additions in the future, that would otherwise be not really possible. also the local bridge would stay, but technically that could also be dropped, just that currently vnet means bridge basically, so im not sure about that.. > - There is no way to add communities via configurable routemaps in > proxmox atm; This makes it a bit more difficult to properly configure > filters when interacting with hosts speaking BGP outside of the proxmox > cluster; Same holds for import filters based on matching communities. > - The imported routes webinterface in the fabric view is a bit > overwhelmed when the nodes get a (v6) fulltable (questionable concept, > but... it can make sense ;-)) Pagination there would probably be > helpful (also, in general, for larger networks). > your last two points sound like a reasonable addition, and we track things like this at [1] as an `enhancement`, you can just file them there with your rationale and use-case > Happy to hear your thoughts; If this is the wrong place for discussing > this, please also let me know. > it's fine, pve-user[2] might be more fitting, since pve-devel is mostly for stuff that involves code > With best regards, > Tobias > [1] https://bugzilla.proxmox.com/ [2] https://lists.proxmox.com/postorius/lists/pve-user.lists.proxmox.com/