public inbox for pve-devel@lists.proxmox.com
 help / color / mirror / Atom feed
From: "DERUMIER, Alexandre" <alexandre.derumier@groupe-cyllene.com>
To: "pve-devel@lists.proxmox.com" <pve-devel@lists.proxmox.com>,
	"s.hanreich@proxmox.com" <s.hanreich@proxmox.com>
Subject: Re: [PATCH] sdn: evpn: allow vlan-aware vnets
Date: Wed, 9 Sep 2026 04:52:44 +0000	[thread overview]
Message-ID: <06a5831f0c8de47dc97614be9395940682117d3f.camel@groupe-cyllene.com> (raw)
In-Reply-To: <d073da2f-dc42-42af-8c13-4c360b7c882b@proxmox.com>

Hi,

they are also another way (that work with vxlan, but I'm not sure with
evpn),

https://docs.nvidia.com/networking-ethernet-software/cumulus-linux-515/Network-Virtualization/VXLAN-Devices/
https://blog.vyos.io/evpn-vxlan-enhancements-introducing-single-vxlan-device-support
https://github.com/FRRouting/frr/pull/12364

it's mapping vlan tags from a vlan aware bridge to vnis, and transport
them in a single vxlan interface.


It was pretty new and buggy 5year ago, so I never tried to implemented
it,
maybe it could be interesting to look at it.


Le mardi 08 septembre 2026 à 11:45 +0200, Stefan Hanreich a écrit :
> Thanks for your contribution!
> 
> We were discussing internally if this approach is the right one.
> Currently,
> with the EVPN zone it is possible to create an L3 EVI and via the
> VXLAN
> zone it is possible to create a standalone L2 EVI, but this is
> currently a
> bit of a hack.
> 
> We're strongly considering improving upon this workaround and making
> it the officially supported and documented way of creating standalone
> L2
> EVIs, as well as integrating this nicer into our stack to make it
> more
> discoverable. Does utilizing a VLAN-aware VXLAN zone with an EVPN
> controller
> and then attaching a route map to the EVPN controller work for your
> use case
> as well? We know of quite a few people that are using this setup and
> are
> considering it when extending the SDN stack. That's also part of why
> we want
> to go down that route.
> 
> Main problems I currently see are that it is not possible to create a
> VXLAN
> zone without any peers. Learning VTEP IPs via type-3 routes works
> perfectly
> fine though - it is just necessary to enter any peer IP address
> (which is
> awkward to say the least). Also, the handling of route distribution
> happens
> implicitly instead of explicitly, so adding the option of assigning
> EVPN
> controllers to VXLAN zones explicitly would be preferred over the
> status
> quo - but we'd have to find a way that makes this backwards
> compatible.
> 
> Do you think this is a tenable solution for your use-case? Do you see
> any
> other issues, particularly with your specific setup?
> 
> 
> On 9/4/26 11:27 AM, Thomas Glanzmann wrote:
> > EVPN vnets rejected the vlanaware flag outright, so a guest could
> > not be
> > attached to an EVPN vnet with a VLAN tag -- tap_plug bails out with
> > "vm
> > vlans are not allowed on vnet <vnet>" because the vnet bridge has
> > VLAN
> > filtering disabled.
> > 
> > Lift the restriction and generate a VLAN-aware vnet bridge (vids 2-
> > 4094)
> > just like the VXLAN zone does, so the guest VLAN tags are carried
> > transparently over the vnet's VNI: a guest on an untagged port
> > sends
> > 802.1q frames that get VXLAN encapsulated as-is, and a guest on a
> > port
> > with a VLAN tag receives them with the tag popped.> Such a vnet is
> > a VLAN bundle service, and FRR does not implement the
> > VLAN-aware variant of it: it originates type-2 (MAC/IP) routes with
> > an
> > Ethernet Tag ID of 0, so the VLAN a MAC was learned on is not
> > carried in
> > the route. FRR does advertise those MACs -- it does not restrict
> > itself
> > to the VNI's access VLAN -- but a receiver has no way to tell which
> > VLAN
> > they belong to.
> > 
> > Between FRR peers that is harmless. The receiver installs the route
> > into
> > the VNI's access VLAN where it is inert, and the entry that
> > actually
> > forwards the traffic comes from data plane learning:
> > 
> >   bc:24:11:c7:71:9c dev vxlan_e1 vlan 1 extern_learn master e1
> >   bc:24:11:c7:71:9c dev vxlan_e1 vlan 5 master e1
> > 
> > Against a third-party VTEP it is not harmless, and it breaks
> > differently
> > depending on how the peer maps the Ethernet Tag. Measured in a
> > four-VTEP
> > fabric (VNI 1, customer VLAN 5 carried tagged inside it) with the
> > route-map deny below removed:
> > 
> >   - Cisco NX-OS 9.3, vlan-based VNI fed by a QinQ hairpin: imports
> > the
> >     tag-0 route as a control-plane MAC in the VNI's bridge domain,
> > 
> >       C 1000 bc24.11f1.9236 dynamic nve1(10.101.0.11)
> > 
> >     which bypasses the hairpin that restores the inner VLAN tag.
> >     Broadcast survives -- an ARP request is flooded through the
> > hairpin
> >     and arrives correctly tagged -- but the unicast that follows is
> >     dropped and never reaches the guest. 100% loss to every FRR-
> > hosted
> >     guest, in both directions.
> > 
> >   - MikroTik RouterOS 7.24.1: files the tag-0 route under the VXLAN
> >     port's bridge-pvid instead of the VLAN carrying the traffic, as
> > an
> >     EXTERNAL bridge host entry, and then blackholes that MAC
> > outright --
> >     including for traffic in the VLAN it really belongs to.
> > Withdrawing
> >     the route does not clear the entry, the VXLAN interface has to
> > be
> >     bounced.
> > 
> >   - BIRD 3.3.2: unaffected. Its EVPN protocol has an explicit tag
> > to VID
> >     mapping (vni 1 / vid 5 / tag 0), so the route lands in the
> > correct
> >     VLAN and forwards normally.
> > 
> > So keep MAC learning enabled and ARP/ND suppression disabled on the
> > VXLAN
> > interface of such a vnet, and additionally deny the type-2 routes
> > for its
> > VNI in the outgoing VTEP route-map, so no peer is fed MAC routes
> > that
> > carry no usable VLAN. BUM traffic is unaffected, it is still head-
> > end
> > replicated via the type-3 routes installed by the controller, and
> > the
> > tagged VLANs are covered by data plane learning on both sides.
> > Subnets stay mutually exclusive with vlanaware, which is already
> > enforced by the vnet and subnet plugins, so this does not affect
> > the
> > L3/gateway setup of an EVPN zone.
> > 
> > Signed-off-by: Thomas Glanzmann <thomas@glanzmann.de>
> > ---
> >  .../PVE/Network/SDN/Controllers/EvpnPlugin.pm | 33
> > ++++++++++++++++++++++
> >  src/PVE/Network/SDN/Zones/EvpnPlugin.pm       | 21 ++++++++++----
> >  2 files changed, 49 insertions(+), 5 deletions(-)
> > 
> > diff --git a/src/PVE/Network/SDN/Controllers/EvpnPlugin.pm
> > b/src/PVE/Network/SDN/Controllers/EvpnPlugin.pm
> > --- a/src/PVE/Network/SDN/Controllers/EvpnPlugin.pm
> > +++ b/src/PVE/Network/SDN/Controllers/EvpnPlugin.pm
> > @@ -586,6 +586,39 @@
> >  sub generate_vnet_frr_config {
> >      my ($class, $plugin_config, $controller, $zone, $zoneid,
> > $vnetid, $config) = @_;
> >  
> > +    # A vlan-aware vnet carries the guest VLAN tags inside a
> > single VNI, but FRR
> > +    # originates type-2 (MAC/IP) routes with an Ethernet Tag ID of
> > 0, so the VLAN
> > +    # the MAC was learned on is not carried in the route. A
> > receiver installs such
> > +    # a route into the VNI's access VLAN, which is the wrong VLAN
> > for every tagged
> > +    # MAC. Between FRR peers the entry is inert (the tagged VLANs
> > are forwarded via
> > +    # data plane learning), but third-party VTEPs break on it in
> > different ways:
> > +    # Cisco NX-OS imports it as a control-plane MAC in the VNI's
> > bridge domain,
> > +    # which bypasses the QinQ hairpin that restores the inner tag
> > -- flooded
> > +    # traffic still works, the unicast that follows it is dropped.
> > RouterOS files
> > +    # it under the VXLAN port's bridge-pvid and blackholes that
> > MAC outright.
> > +    # BIRD is unaffected, it maps the Ethernet Tag to a VLAN
> > explicitly.
> > +    # Suppress the type-2 routes for these vnets, BUM traffic is
> > unaffected as it
> > +    # is still head-end replicated via the type-3 routes.
> > +    if ($plugin_config->{vlanaware}) {
> > +        my $route_map_out = 'MAP_VTEP_OUT';
> > +        $route_map_out .= "_$controller->{'peer-group-name'}"
> > +            if $controller->{'peer-group-name'};
> > +
> > +        # seq is renumbered by array order in
> > PVE::Network::SDN::Frr, so unshift
> > +        # to be evaluated before the permit entry that terminates
> > the route-map
> > +        unshift(
> > +            @{ $config->{frr}->{routemaps}->{$route_map_out} },
> > +            {
> > +                seq => 1,
> > +                action => 'deny',
> > +                matches => [
> > +                    { key => 'evpn route-type', value => 'macip'
> > },
> > +                    { key => 'evpn vni', value => $plugin_config-
> > >{tag} },
> > +                ],
> > +            },
> > +        );
> > +    }
> > +
> >      my $exitnodes = $zone->{'exitnodes'};
> >      my $exitnodes_local_routing = $zone->{'exitnodes-local-
> > routing'};
> >  
> > diff --git a/src/PVE/Network/SDN/Zones/EvpnPlugin.pm
> > b/src/PVE/Network/SDN/Zones/EvpnPlugin.pm
> > index 0e79707..e597120 100644
> > --- a/src/PVE/Network/SDN/Zones/EvpnPlugin.pm
> > +++ b/src/PVE/Network/SDN/Zones/EvpnPlugin.pm
> > @@ -218,15 +218,24 @@ sub generate_sdn_config {
> >      }
> >      $mtu = $plugin_config->{mtu} if $plugin_config->{mtu};
> >  
> > +    # a vlan-aware vnet transports the guest VLAN tags
> > transparently over a single
> > +    # VNI. The EVPN control plane only knows about the VNI's
> > access VLAN, so it can
> > +    # neither advertise nor install MAC entries for the other
> > VLANs. Fall back to
> > +    # data plane learning (and no ARP/ND suppression) for those
> > vnets, BUM traffic
> > +    # is still replicated via the type-3 routes installed by the
> > controller.
> > +    my $vlanaware = $vnet->{vlanaware};
> > +
> >      #vxlan interface
> >      my $vxlan_iface = "vxlan_$vnetid";
> >      my @iface_config = ();
> >      push @iface_config, "vxlan-id $tag";
> >      push @iface_config, "vxlan-local-tunnelip $ifaceip" if
> > $ifaceip;
> >      push @iface_config, "vxlan-port $vxlanport" if $vxlanport;
> > -    push @iface_config, "bridge-learning off";
> > -    push @iface_config, "bridge-arp-nd-suppress on"
> > -        if !$plugin_config->{'disable-arp-nd-suppression'};
> > +    if (!$vlanaware) {
> > +        push @iface_config, "bridge-learning off";
> > +        push @iface_config, "bridge-arp-nd-suppress on"
> > +            if !$plugin_config->{'disable-arp-nd-suppression'};
> > +    }
> >  
> >      push @iface_config, "mtu $mtu" if $mtu;
> >      push(@{ $config->{$vxlan_iface} }, @iface_config) if !$config-
> > >{$vxlan_iface};
> > @@ -299,6 +308,10 @@ sub generate_sdn_config {
> >      push @iface_config, "bridge_ports $vxlan_iface";
> >      push @iface_config, "bridge_stp off";
> >      push @iface_config, "bridge_fd 0";
> > +    if ($vlanaware) {
> > +        push @iface_config, "bridge-vlan-aware yes";
> > +        push @iface_config, "bridge-vids 2-4094";
> > +    }
> >      push @iface_config, "mtu $mtu" if $mtu;
> >      push @iface_config, "alias $alias" if $alias;
> >      push @iface_config, "ip-forward on" if $enable_forward_v4;
> > @@ -423,8 +436,6 @@ sub vnet_update_hook {
> >  
> >      raise_param_exc({ tag => "missing vxlan tag" }) if
> > !defined($tag);
> >      raise_param_exc({ tag => "vxlan tag max value is 16777216" })
> > if $tag > 16777216;
> > -    raise_param_exc({ 'vlan-aware' => "vlan-aware option can't be
> > enabled with evpn" })
> > -        if $vnet->{vlanaware};
> >  
> >      # verify that tag is not already defined globally (vxlan-id
> > are unique)
> >      foreach my $id (keys %{ $vnet_cfg->{ids} }) {
> 
> 
> 

  parent reply	other threads:[~2026-09-09  4:53 UTC|newest]

Thread overview: 7+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-03 21:49 [PATCH] sdn: evpn: allow vlan-aware vnets Thomas Glanzmann
2026-09-08  9:45 ` Stefan Hanreich
2026-09-08 20:56   ` Thomas Glanzmann
2026-09-09  8:40     ` Stefan Hanreich
2026-09-09  9:22       ` Gabriel Goller
2026-09-09  4:52   ` DERUMIER, Alexandre [this message]
2026-09-09  8:06     ` Stefan Hanreich

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=06a5831f0c8de47dc97614be9395940682117d3f.camel@groupe-cyllene.com \
    --to=alexandre.derumier@groupe-cyllene.com \
    --cc=pve-devel@lists.proxmox.com \
    --cc=s.hanreich@proxmox.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Service provided by Proxmox Server Solutions GmbH | Privacy | Legal