Skip to content

Draining an IS-IS router: three vendors, three LSPs, and a bug in FRR

The same IS-IS drain sends three different LSPs on SR OS, Junos and FRR. With several exits, not every drain works, and FRR breaks RFC 5305 on the maximum metric.

12 min read
bug containerlab frr interoperability isis juniper multi-vendor nokia
Photo by Alexey Taktarov on Unsplash

Draining an IS-IS router: three vendors, three LSPs, and a bug in FRR

In the previous post, one LSP with both the attached bit and the overload bit gave a default route on Junos and FRR, and none on SR OS. I left two questions open: does every vendor keep the attached bit when overload is set, and what happens when the access router has several exits, which is the real use of a drain? And the first post tested only one sender, SR OS.

So I ran all of it: three vendors as the drained router, three vendors as the access router, one exit and several exits. The short version:

  • The same overload configuration sends three different LSPs. SR OS and FRR keep the attached bit with the overload bit. Junos clears it. With the maximum metric, only FRR keeps it.
  • With several exits, a drain works for every access router only if the drained router clears ATT. Junos does it by default. SR OS needs the maximum metric or suppress-attached-bit, and FRR needs no attached-bit send.
  • The maximum metric is not one value. SR OS and Junos advertise 16777214 and clear the attached bit. FRR keeps the attached bit and advertises 16777215, the value that RFC 5305 excludes from the SPF. Junos applies that rule and loses the drained FRR router completely. SR OS and FRR access routers keep their default route through it.
  • FRR's own SPF uses links at 16777215. That is a bug, or at least a gap. It is reported, a fix is proposed, and it still exists in the current master.

The lab, now with three senders

The first lab had one drained router (SR OS). This one has three, and every vendor is also an access router. Three independent pods share one level 2 core. In each pod, one L1/L2 border router of a different vendor serves three L1 routers (SR OS, Junos, FRR). That is 13 nodes in containerlab.

flowchart TB
  l2core["<b>l2core</b><br/>FRR<br/>L2 only, area 49.0100"]

  subgraph pod1["pod srsim, area 49.0011"]
    sb["<b>srsim-border</b><br/>SR-SIM, L1/L2"]
    s1["l1srsim"]
    s2["l1junos"]
    s3["l1frr"]
  end
  subgraph pod2["pod frr, area 49.0012"]
    fb["<b>frr-border</b><br/>FRR, L1/L2"]
    f1["l1srsim"]
    f2["l1junos"]
    f3["l1frr"]
  end
  subgraph pod3["pod junos, area 49.0013"]
    jb["<b>junos-border</b><br/>Junos, L1/L2"]
    j1["l1srsim"]
    j2["l1junos"]
    j3["l1frr"]
  end

  l2core ---|L2| sb
  l2core ---|L2| fb
  l2core ---|L2| jb
  sb ---|L1| s1 & s2 & s3
  fb ---|L1| f1 & f2 & f3
  jb ---|L1| j1 & j2 & j3

  classDef border fill:#fde7c8,stroke:#c77700,color:#000
  classDef other fill:#e5e7eb,stroke:#6b7280,color:#000
  class sb,fb,jb border
  class l2core other

Each border has an L2 adjacency in another area, so each one sets the attached bit. The L1 routers have no other way out: a default route on them can only come from that bit.

The lab: labs/isis_interop_drain.

What a drained border sends

Border Overload bit Maximum metric
SR OS (SR-SIM 26.3.R1) ATT + OL no flag, links at 16777214
FRR 10.2.5 ATT + OL ATT, links and prefixes at 16777215
Junos 23.4R2-S6.9 OL only no flag, links at 16777214

Junos clears ATT, as its documentation says. SR OS and FRR keep the attached bit. In maximum metric mode, SR OS drops it too (as it did on the vSIM), and FRR is the only one to keep it.

The attached bit decides whether the access routers have a default route through the border at all. A drained router that clears it gives the access routers no reason to use it. The interesting case is the router that sends ATT while it is drained.

Several exits: who still sends traffic to the drained router?

With one border, "no default route" is a bad outcome, not a drain. The real use has two exits or more. So the second lab has three borders (SR OS, FRR, Junos) in one area, and three L1 routers (SR OS, Junos, FRR). Each L1 router links to all three borders with the metric 10.

%%{init: {"flowchart": {"rankSpacing": 90}}}%%
flowchart TB
  l2core["<b>l2core</b><br/>FRR<br/>L2 only, area 49.0100"]

  subgraph area["area 49.0011"]
    bs["<b>border-srsim</b><br/>SR-SIM, L1/L2"]
    bf["<b>border-frr</b><br/>FRR, L1/L2"]
    bj["<b>border-junos</b><br/>Junos, L1/L2"]
    ls["l1-srsim"]
    lj["l1-junos"]
    lf["l1-frr"]
  end

  l2core ---|L2| bs
  l2core ---|L2| bf
  l2core ---|L2| bj
  bs --- ls & lj & lf
  bf --- ls & lj & lf
  bj --- ls & lj & lf

  classDef srsim fill:#fde7c8,stroke:#ea580c,stroke-width:3px,color:#000
  classDef frr fill:#fde7c8,stroke:#2563eb,stroke-width:3px,color:#000
  classDef junos fill:#fde7c8,stroke:#16a34a,stroke-width:3px,color:#000
  classDef other fill:#e5e7eb,stroke:#6b7280,color:#000
  class bs srsim
  class bf frr
  class bj junos
  class l2core other
  linkStyle 3,4,5 stroke:#ea580c,stroke-width:2px
  linkStyle 6,7,8 stroke:#2563eb,stroke-width:2px
  linkStyle 9,10,11 stroke:#16a34a,stroke-width:2px

Before the drain, every L1 router has three equal-cost next hops for 0.0.0.0/0, one per border. The drain of one border then shows in the next-hop list: if the drained router leaves the list, the drain works for that access router. (I set ecmp 3 on the SR OS access router so that it can install three next hops.)

Drained border Mode In its LSP l1 SR OS l1 Junos l1 FRR
SR OS overload bit ATT + OL leaves the list stays in the list stays in the list
SR OS max-metric no flag, links at 16777214 leaves leaves leaves
SR OS overload bit + suppress-attached-bit OL leaves leaves leaves
FRR overload bit ATT + OL leaves stays stays
FRR advertise-high-metrics ATT, links at 16777215 stays leaves stays
FRR overload bit + no attached-bit send OL leaves leaves leaves
FRR advertise-high-metrics + no attached-bit send links at 16777215 leaves leaves leaves
Junos overload bit OL leaves leaves leaves
Junos advertise-high-metrics no flag, links at 16777214 leaves leaves leaves

The cells in bold are the problem: the router was drained, and the access router still sends inter-area traffic to it, next to the other borders. The two rows with no attached-bit send come from the same lab without the Junos border, so with two exits (SR OS and FRR) and the same three L1 routers.

Three things come out of this table:

  1. The claim of the first post is confirmed. With the overload bit from SR OS or FRR, the Junos and FRR access routers keep the overloaded router as a next hop. Only the SR OS access router moves away.
  2. Junos is the safest sender, in both modes, because it clears ATT. SR OS is safe with the maximum metric or with suppress-attached-bit.
  3. FRR drains for all three only without ATT. Its overload bit works for SR OS access routers only, and advertise-high-metrics for Junos access routers only. With no attached-bit send, both modes drain for every access router.

The FRR advertise-high-metrics row is the odd one: only the Junos access router leaves. SR OS and Junos clear ATT in their maximum-metric mode, so the border is no longer a default-route candidate, whatever the metric. FRR keeps ATT, and the cost from an L1 router to the border is the L1 router's own link metric (10), which does not change.

So the border can only leave if it becomes unreachable. FRR advertises its links toward the L1 routers at 16777215, a value that RFC 5305 excludes from the SPF. Junos applies this to the links of the border, the two-way connectivity check fails, and the border is unreachable. SR OS and FRR still use those links, so they keep the border.

L1 routerSR OS, Junos or FRR border-frradvertise-high-metricsATT set set by the L1 router 10 16777215 set by border-frr Access router The 16777215 entry and the two-way check Drained border Junos does not count: check fails, border unreachable leaves SR OS counts: border reachable at 10, ATT set stays FRR counts: border reachable at 10, ATT set stays

The maximum metric is not one value

SR OS and Junos advertise 16777214 (2^24 - 2) on every link in maximum metric mode. FRR advertises 16777215 (2^24 - 1), on the links and on its own prefixes:

frr-border.00-00     *    149   0x0000000a  0xabf3    1125    1/0/0
  Extended Reachability: 0000.0002.0002.00 (Metric: 16777215)
  Extended Reachability: 0000.0002.0003.00 (Metric: 16777215)
  Extended Reachability: 0000.0002.0004.00 (Metric: 16777215)
  Extended IP Reachability: 10.2.0.0/30 (Metric: 16777215)

One unit of difference matters, because the RFCs give the two values two different meanings. RFC 5305, section 3:

If a link is advertised with the maximum link metric (2^24 - 1), this link MUST NOT be considered during the normal SPF computation.

And RFC 5443, section 2, which is about LDP IGP synchronization and also needs a link that stays usable as a last resort:

In the case of ISIS, the maximum metric value is 2^24-2 (0xFFFFFE). Indeed, if a link is configured with 2^24-1 (the maximum link metric per [RFC5305]), then this link is not advertised in the topology. It is important to keep the link in the topology to allow IP traffic to use the link as a last resort in case of massive failure.

So 2^24 - 2 means: avoid this link, but keep it as a last resort. And 2^24 - 1 means: this link is not part of the SPF. That is exactly the choice that a metric-based drain needs (RFC 3277, section 4: with metrics, "transit paths may still be calculated through the router").

Who obeys the RFC?

To see who follows RFC 5305, I set the value 16777215 on one link and looked at the default route of the L1 router. A link has two directions, so there are two cases:

  • Forward: the L1 router advertises 16777215 for its own link to the border. It uses this metric to reach the border.
  • Reverse: the border advertises 16777215 for its link to the L1 router. The L1 router reads it in the LSP of the border. It does not add to the cost of reaching the border, but it matters for the two-way connectivity check.
Case Metric SR OS L1 Junos L1 FRR L1
Forward 16777215 no default route no default route default route, metric 16777215
Reverse (SR OS border) 16777215 default route no default route default route
Either 16777214 default route default route default route

With 16777214, all three behave the same. With 16777215:

  • Junos follows the RFC in both directions. The link is out of the SPF, and in the reverse case the two-way connectivity check fails, so the border is unreachable. This is why a drained FRR border disappears for Junos.
  • SR OS follows it for its own links, not for the entry of its neighbor. RFC 5305 does not say if the two-way check is part of "the normal SPF computation", so this is a different reading, not a clear fault. In my opinion, Junos has the better reading. ISO/IEC 10589, section 7.2.8.2, puts the check in the decision process: it "shall not utilise a link between two Intermediate Systems unless both ISs report the link", and "Reporting the link indicates that it has a defined value for at least the default routeing metric". And RFC 5443 says that a link at 2^24 - 1 "is not advertised in the topology" (in section 2, this means: not part of the topology that the SPF uses), so its entry should not prove a two-way adjacency either.
  • FRR does not follow it at all. A link with 16777215 is used like any other link, with a huge metric.

Back to the drained FRR border. The value 16777215 changes the result on Junos only: Junos applies the RFC to the links of the border and loses the router. SR OS and FRR access routers keep the default route through it, because FRR keeps ATT in this mode and their own link metric to the border does not change. With 16777214, all three access vendors would keep the default route through a border that sets ATT, as the reverse test shows.

The bug in FRR

I reduced it to two FRR containers on a Docker network, with one IS-IS adjacency and wide metrics. On a 10.2.5 image and on the current master image (10.8.0-dev, version string git20261002), r1 has a route to the loopback of r2 with the metric 20:

r1# show ip route 10.255.0.2/32
  Known via "isis", distance 115, metric 20, best

Then the metric of the link of r1, in the CLI that accepts (0-16777215):

interface eth0
 isis metric level-1 16777215
Extended Reachability: 0000.0000.0002.00 (Metric: 16777215)
Known via "isis", distance 115, metric 16777225, best

The link is used: 10 + 16777215. According to RFC 5305 the route should be gone. The same happens for a link that a neighbor advertises. With a third router r3 behind r2, and 16777215 on the link of r2 toward r3, r1 still reaches r3 through it, with the metric 16777235. And advertise-high-metrics on r2 writes 16777215 on its links and on its loopback prefix:

Extended Reachability: 0000.0000.0001.00 (Metric: 16777215)
r1FRR r2FRR r3FRR set by r1 10 set by r2 16777215 loopback, metric 10 RFC 5305, section 3 a link at 2^24 - 1 "MUST NOT be considered during the normal SPF computation" route from r1 to r3 none FRR 10.2.5 and master the link is used like any other link route from r1 to r3 10 + 16777215 + 10 = 16777235

In the source, the SPF adds the metric of each IS reachability entry to the path cost without any check (isis_spf_process_lsp()), and isis_area_advertise_high_metrics_set() sets MAX_WIDE_LINK_METRIC (0x00FFFFFF). FRR already has the right constant, ISIS_WIDE_METRIC_INFINITY (0xFFFFFE), for its LDP-IGP synchronization. And the FRR documentation states the value ("for wide metrics, 16777215"), and the topotest of the feature expects it.

The clear deviation is the SPF. The two should be discussed together, because fixing the SPF alone changes what advertise-high-metrics does on FRR peers. And the right value alone does not make advertise-high-metrics drain the default route unless no attached-bit send is set: FRR keeps the attached bit in this mode, where SR OS and Junos clear it.

Who runs FRR

You may run FRR without calling it FRR. Several products use it as their routing engine:

Product Where FRR runs
VyOS (router) all routing
Netgate TNSR (router) dynamic routing
Proxmox VE 9 (virtualization) SDN fabrics
NVIDIA Cumulus Linux (switch) all routing
Palo Alto Networks PAN-OS (firewall) frr 8.4_dpd in the open-source listing of PAN-OS 12.1 (listed since PAN-OS 10.1)

Where the bug stands

I found no earlier report. The two related items are the PR that added the feature in 2023 and a 2024 PR about the 0xFFFFFE flag of LFA. I opened issue 23517 with a reproduction on two and three FRR routers, the multi-vendor results, the RFC quotes and a suggested fix: ignore the value 0xFFFFFF in the SPF, use 0xFFFFFE in advertise-high-metrics, and update the topotest and the documentation.

On 2026-10-04, a contributor confirmed the issue and opened PR 23527: the SPF ignores links at 16777215, and advertise-high-metrics is unchanged. A maintainer approved the PR on 2026-10-06. At the time of writing, it is not merged yet, and the issue is still open.

What to do on your network

The commands that drained for every access vendor in the labs with several exits, per sender. A drain that clears ATT removes the default route of the access routers that use the drained router, so it assumes that they have another exit, as in that lab.

Drained router What drains for every access vendor What does not
Junos set protocols isis overload, with or without advertise-high-metrics (it clears ATT) nothing found
SR OS overload max-metric true, or overload with suppress-attached-bit true plain overload (SR OS access routers only)
FRR set-overload-bit or advertise-high-metrics, with no attached-bit send the same commands without it: set-overload-bit (SR OS access routers only), advertise-high-metrics (Junos access routers only)

If you cannot change the sender, the receiver can ignore the attached bit (ignore-attached-bit on SR OS and Junos, attached-bit receive ignore on FRR), as in the first post. It ignores the bit of every neighbor, not only the drained ones, so it needs a static default route or another exit.

Lessons

  1. A drain is a pair of vendors, not a command. The sender decides the flags and the metrics in the LSP, the receiver decides what to do with them. Test every pair that you run: with the overload bit, a Junos border drains for every access router, but an FRR border keeps inter-area traffic from Junos and FRR access routers.
  2. A similar configuration is not a similar behavior. Each vendor has an overload command and a maximum-metric mode, but not with the same defaults: in overload, SR OS and FRR keep ATT and Junos clears it; in maximum-metric mode, SR OS and Junos clear ATT, FRR keeps it and uses another metric value. In a multi-vendor network, check what each command really sends, not only its name.
  3. Look at the LSP, not at the command. show router isis database and show isis database tell you at a glance if ATT and OL are there. In maximum metric mode there is no OL flag: look at the metric of the neighbors (16777214, or 16777215 on FRR).
  4. Close to the same value is not the same value. One unit of metric, 2^24 - 2 against 2^24 - 1, is the difference between "very costly" and "not in the topology", and only the RFC text says so.
  5. A lab finds in an afternoon what an incident finds at night. Twenty containers in two labs, a script, and the question "what does the other vendor do with this?" were enough.

The lab and the scripts are in labs/isis_interop_drain. If you run it on other releases, or with another vendor that speaks IS-IS, I would like to know what you see.