public inbox for pve-devel@lists.proxmox.com
 help / color / mirror / Atom feed
From: Stefan Hanreich <s.hanreich@proxmox.com>
To: pve-devel@lists.proxmox.com
Subject: Re: [PATCH] sdn: evpn: allow vlan-aware vnets
Date: Tue, 8 Sep 2026 11:45:46 +0200	[thread overview]
Message-ID: <d073da2f-dc42-42af-8c13-4c360b7c882b@proxmox.com> (raw)
In-Reply-To: <apnrUDqiauznKjnF@glanzmann.de>

Thanks for your contribution!

We were discussing internally if this approach is the right one. Currently,
with the EVPN zone it is possible to create an L3 EVI and via the VXLAN
zone it is possible to create a standalone L2 EVI, but this is currently a
bit of a hack.

We're strongly considering improving upon this workaround and making
it the officially supported and documented way of creating standalone L2
EVIs, as well as integrating this nicer into our stack to make it more
discoverable. Does utilizing a VLAN-aware VXLAN zone with an EVPN controller
and then attaching a route map to the EVPN controller work for your use case
as well? We know of quite a few people that are using this setup and are
considering it when extending the SDN stack. That's also part of why we want
to go down that route.

Main problems I currently see are that it is not possible to create a VXLAN
zone without any peers. Learning VTEP IPs via type-3 routes works perfectly
fine though - it is just necessary to enter any peer IP address (which is
awkward to say the least). Also, the handling of route distribution happens
implicitly instead of explicitly, so adding the option of assigning EVPN
controllers to VXLAN zones explicitly would be preferred over the status
quo - but we'd have to find a way that makes this backwards compatible.

Do you think this is a tenable solution for your use-case? Do you see any
other issues, particularly with your specific setup?


On 9/4/26 11:27 AM, Thomas Glanzmann wrote:
> EVPN vnets rejected the vlanaware flag outright, so a guest could not be
> attached to an EVPN vnet with a VLAN tag -- tap_plug bails out with "vm
> vlans are not allowed on vnet <vnet>" because the vnet bridge has VLAN
> filtering disabled.
> 
> Lift the restriction and generate a VLAN-aware vnet bridge (vids 2-4094)
> just like the VXLAN zone does, so the guest VLAN tags are carried
> transparently over the vnet's VNI: a guest on an untagged port sends
> 802.1q frames that get VXLAN encapsulated as-is, and a guest on a port
> with a VLAN tag receives them with the tag popped.> Such a vnet is a VLAN bundle service, and FRR does not implement the
> VLAN-aware variant of it: it originates type-2 (MAC/IP) routes with an
> Ethernet Tag ID of 0, so the VLAN a MAC was learned on is not carried in
> the route. FRR does advertise those MACs -- it does not restrict itself
> to the VNI's access VLAN -- but a receiver has no way to tell which VLAN
> they belong to.
> 
> Between FRR peers that is harmless. The receiver installs the route into
> the VNI's access VLAN where it is inert, and the entry that actually
> forwards the traffic comes from data plane learning:
> 
>   bc:24:11:c7:71:9c dev vxlan_e1 vlan 1 extern_learn master e1
>   bc:24:11:c7:71:9c dev vxlan_e1 vlan 5 master e1
> 
> Against a third-party VTEP it is not harmless, and it breaks differently
> depending on how the peer maps the Ethernet Tag. Measured in a four-VTEP
> fabric (VNI 1, customer VLAN 5 carried tagged inside it) with the
> route-map deny below removed:
> 
>   - Cisco NX-OS 9.3, vlan-based VNI fed by a QinQ hairpin: imports the
>     tag-0 route as a control-plane MAC in the VNI's bridge domain,
> 
>       C 1000 bc24.11f1.9236 dynamic nve1(10.101.0.11)
> 
>     which bypasses the hairpin that restores the inner VLAN tag.
>     Broadcast survives -- an ARP request is flooded through the hairpin
>     and arrives correctly tagged -- but the unicast that follows is
>     dropped and never reaches the guest. 100% loss to every FRR-hosted
>     guest, in both directions.
> 
>   - MikroTik RouterOS 7.24.1: files the tag-0 route under the VXLAN
>     port's bridge-pvid instead of the VLAN carrying the traffic, as an
>     EXTERNAL bridge host entry, and then blackholes that MAC outright --
>     including for traffic in the VLAN it really belongs to. Withdrawing
>     the route does not clear the entry, the VXLAN interface has to be
>     bounced.
> 
>   - BIRD 3.3.2: unaffected. Its EVPN protocol has an explicit tag to VID
>     mapping (vni 1 / vid 5 / tag 0), so the route lands in the correct
>     VLAN and forwards normally.
> 
> So keep MAC learning enabled and ARP/ND suppression disabled on the VXLAN
> interface of such a vnet, and additionally deny the type-2 routes for its
> VNI in the outgoing VTEP route-map, so no peer is fed MAC routes that
> carry no usable VLAN. BUM traffic is unaffected, it is still head-end
> replicated via the type-3 routes installed by the controller, and the
> tagged VLANs are covered by data plane learning on both sides.
> Subnets stay mutually exclusive with vlanaware, which is already
> enforced by the vnet and subnet plugins, so this does not affect the
> L3/gateway setup of an EVPN zone.
> 
> Signed-off-by: Thomas Glanzmann <thomas@glanzmann.de>
> ---
>  .../PVE/Network/SDN/Controllers/EvpnPlugin.pm | 33 ++++++++++++++++++++++
>  src/PVE/Network/SDN/Zones/EvpnPlugin.pm       | 21 ++++++++++----
>  2 files changed, 49 insertions(+), 5 deletions(-)
> 
> diff --git a/src/PVE/Network/SDN/Controllers/EvpnPlugin.pm b/src/PVE/Network/SDN/Controllers/EvpnPlugin.pm
> --- a/src/PVE/Network/SDN/Controllers/EvpnPlugin.pm
> +++ b/src/PVE/Network/SDN/Controllers/EvpnPlugin.pm
> @@ -586,6 +586,39 @@
>  sub generate_vnet_frr_config {
>      my ($class, $plugin_config, $controller, $zone, $zoneid, $vnetid, $config) = @_;
>  
> +    # A vlan-aware vnet carries the guest VLAN tags inside a single VNI, but FRR
> +    # originates type-2 (MAC/IP) routes with an Ethernet Tag ID of 0, so the VLAN
> +    # the MAC was learned on is not carried in the route. A receiver installs such
> +    # a route into the VNI's access VLAN, which is the wrong VLAN for every tagged
> +    # MAC. Between FRR peers the entry is inert (the tagged VLANs are forwarded via
> +    # data plane learning), but third-party VTEPs break on it in different ways:
> +    # Cisco NX-OS imports it as a control-plane MAC in the VNI's bridge domain,
> +    # which bypasses the QinQ hairpin that restores the inner tag -- flooded
> +    # traffic still works, the unicast that follows it is dropped. RouterOS files
> +    # it under the VXLAN port's bridge-pvid and blackholes that MAC outright.
> +    # BIRD is unaffected, it maps the Ethernet Tag to a VLAN explicitly.
> +    # Suppress the type-2 routes for these vnets, BUM traffic is unaffected as it
> +    # is still head-end replicated via the type-3 routes.
> +    if ($plugin_config->{vlanaware}) {
> +        my $route_map_out = 'MAP_VTEP_OUT';
> +        $route_map_out .= "_$controller->{'peer-group-name'}"
> +            if $controller->{'peer-group-name'};
> +
> +        # seq is renumbered by array order in PVE::Network::SDN::Frr, so unshift
> +        # to be evaluated before the permit entry that terminates the route-map
> +        unshift(
> +            @{ $config->{frr}->{routemaps}->{$route_map_out} },
> +            {
> +                seq => 1,
> +                action => 'deny',
> +                matches => [
> +                    { key => 'evpn route-type', value => 'macip' },
> +                    { key => 'evpn vni', value => $plugin_config->{tag} },
> +                ],
> +            },
> +        );
> +    }
> +
>      my $exitnodes = $zone->{'exitnodes'};
>      my $exitnodes_local_routing = $zone->{'exitnodes-local-routing'};
>  
> diff --git a/src/PVE/Network/SDN/Zones/EvpnPlugin.pm b/src/PVE/Network/SDN/Zones/EvpnPlugin.pm
> index 0e79707..e597120 100644
> --- a/src/PVE/Network/SDN/Zones/EvpnPlugin.pm
> +++ b/src/PVE/Network/SDN/Zones/EvpnPlugin.pm
> @@ -218,15 +218,24 @@ sub generate_sdn_config {
>      }
>      $mtu = $plugin_config->{mtu} if $plugin_config->{mtu};
>  
> +    # a vlan-aware vnet transports the guest VLAN tags transparently over a single
> +    # VNI. The EVPN control plane only knows about the VNI's access VLAN, so it can
> +    # neither advertise nor install MAC entries for the other VLANs. Fall back to
> +    # data plane learning (and no ARP/ND suppression) for those vnets, BUM traffic
> +    # is still replicated via the type-3 routes installed by the controller.
> +    my $vlanaware = $vnet->{vlanaware};
> +
>      #vxlan interface
>      my $vxlan_iface = "vxlan_$vnetid";
>      my @iface_config = ();
>      push @iface_config, "vxlan-id $tag";
>      push @iface_config, "vxlan-local-tunnelip $ifaceip" if $ifaceip;
>      push @iface_config, "vxlan-port $vxlanport" if $vxlanport;
> -    push @iface_config, "bridge-learning off";
> -    push @iface_config, "bridge-arp-nd-suppress on"
> -        if !$plugin_config->{'disable-arp-nd-suppression'};
> +    if (!$vlanaware) {
> +        push @iface_config, "bridge-learning off";
> +        push @iface_config, "bridge-arp-nd-suppress on"
> +            if !$plugin_config->{'disable-arp-nd-suppression'};
> +    }
>  
>      push @iface_config, "mtu $mtu" if $mtu;
>      push(@{ $config->{$vxlan_iface} }, @iface_config) if !$config->{$vxlan_iface};
> @@ -299,6 +308,10 @@ sub generate_sdn_config {
>      push @iface_config, "bridge_ports $vxlan_iface";
>      push @iface_config, "bridge_stp off";
>      push @iface_config, "bridge_fd 0";
> +    if ($vlanaware) {
> +        push @iface_config, "bridge-vlan-aware yes";
> +        push @iface_config, "bridge-vids 2-4094";
> +    }
>      push @iface_config, "mtu $mtu" if $mtu;
>      push @iface_config, "alias $alias" if $alias;
>      push @iface_config, "ip-forward on" if $enable_forward_v4;
> @@ -423,8 +436,6 @@ sub vnet_update_hook {
>  
>      raise_param_exc({ tag => "missing vxlan tag" }) if !defined($tag);
>      raise_param_exc({ tag => "vxlan tag max value is 16777216" }) if $tag > 16777216;
> -    raise_param_exc({ 'vlan-aware' => "vlan-aware option can't be enabled with evpn" })
> -        if $vnet->{vlanaware};
>  
>      # verify that tag is not already defined globally (vxlan-id are unique)
>      foreach my $id (keys %{ $vnet_cfg->{ids} }) {





      reply	other threads:[~2026-09-08  9:45 UTC|newest]

Thread overview: 2+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-03 21:49 [PATCH] sdn: evpn: allow vlan-aware vnets Thomas Glanzmann
2026-09-08  9:45 ` Stefan Hanreich [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=d073da2f-dc42-42af-8c13-4c360b7c882b@proxmox.com \
    --to=s.hanreich@proxmox.com \
    --cc=pve-devel@lists.proxmox.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Service provided by Proxmox Server Solutions GmbH | Privacy | Legal