Draining an IS-IS router: three vendors, three LSPs, and a bug in FRR
The same IS-IS drain sends three different LSPs on SR OS, Junos and FRR. With several exits, not every drain works, and FRR breaks RFC 5305 on the maximum metric.
Photo by Alexey Taktarov on UnsplashDraining an IS-IS router: three vendors, three LSPs, and a bug in FRR¶
In the previous post, one LSP with both the attached bit and the overload bit gave a default route on Junos and FRR, and none on SR OS. I left two questions open: does every vendor keep the attached bit when overload is set, and what happens when the access router has several exits, which is the real use of a drain? And the first post tested only one sender, SR OS.
So I ran all of it: three vendors as the drained router, three vendors as the access router, one exit and several exits. The short version:
- The same overload configuration sends three different LSPs. SR OS and FRR keep the attached bit with the overload bit. Junos clears it. With the maximum metric, only FRR keeps it.
- With several exits, a drain works for every access router only if the drained router clears
ATT. Junos does it by default. SR OS needs the maximum metric orsuppress-attached-bit, and FRR needsno attached-bit send. - The maximum metric is not one value. SR OS and Junos advertise
16777214and clear the attached bit. FRR keeps the attached bit and advertises16777215, the value that RFC 5305 excludes from the SPF. Junos applies that rule and loses the drained FRR router completely. SR OS and FRR access routers keep their default route through it. - FRR's own SPF uses links at
16777215. That is a bug, or at least a gap. It is reported, a fix is proposed, and it still exists in the currentmaster.
The lab, now with three senders¶
The first lab had one drained router (SR OS). This one has three, and every vendor is also an access router. Three independent pods share one level 2 core. In each pod, one L1/L2 border router of a different vendor serves three L1 routers (SR OS, Junos, FRR). That is 13 nodes in containerlab.
flowchart TB
l2core["<b>l2core</b><br/>FRR<br/>L2 only, area 49.0100"]
subgraph pod1["pod srsim, area 49.0011"]
sb["<b>srsim-border</b><br/>SR-SIM, L1/L2"]
s1["l1srsim"]
s2["l1junos"]
s3["l1frr"]
end
subgraph pod2["pod frr, area 49.0012"]
fb["<b>frr-border</b><br/>FRR, L1/L2"]
f1["l1srsim"]
f2["l1junos"]
f3["l1frr"]
end
subgraph pod3["pod junos, area 49.0013"]
jb["<b>junos-border</b><br/>Junos, L1/L2"]
j1["l1srsim"]
j2["l1junos"]
j3["l1frr"]
end
l2core ---|L2| sb
l2core ---|L2| fb
l2core ---|L2| jb
sb ---|L1| s1 & s2 & s3
fb ---|L1| f1 & f2 & f3
jb ---|L1| j1 & j2 & j3
classDef border fill:#fde7c8,stroke:#c77700,color:#000
classDef other fill:#e5e7eb,stroke:#6b7280,color:#000
class sb,fb,jb border
class l2core other
Each border has an L2 adjacency in another area, so each one sets the attached bit. The L1 routers have no other way out: a default route on them can only come from that bit.
The lab: labs/isis_interop_drain.
What a drained border sends¶
| Border | Overload bit | Maximum metric |
|---|---|---|
| SR OS (SR-SIM 26.3.R1) | ATT + OL |
no flag, links at 16777214 |
| FRR 10.2.5 | ATT + OL |
ATT, links and prefixes at 16777215 |
| Junos 23.4R2-S6.9 | OL only |
no flag, links at 16777214 |
Junos clears ATT, as its documentation says. SR OS and FRR keep the attached bit. In maximum metric mode, SR OS drops it too (as it did on the vSIM), and FRR is the only one to keep it.
The attached bit decides whether the access routers have a default route through the border at all. A drained router that clears it gives the access routers no reason to use it. The interesting case is the router that sends ATT while it is drained.
Several exits: who still sends traffic to the drained router?¶
With one border, "no default route" is a bad outcome, not a drain. The real use has two exits or more. So the second lab has three borders (SR OS, FRR, Junos) in one area, and three L1 routers (SR OS, Junos, FRR). Each L1 router links to all three borders with the metric 10.
%%{init: {"flowchart": {"rankSpacing": 90}}}%%
flowchart TB
l2core["<b>l2core</b><br/>FRR<br/>L2 only, area 49.0100"]
subgraph area["area 49.0011"]
bs["<b>border-srsim</b><br/>SR-SIM, L1/L2"]
bf["<b>border-frr</b><br/>FRR, L1/L2"]
bj["<b>border-junos</b><br/>Junos, L1/L2"]
ls["l1-srsim"]
lj["l1-junos"]
lf["l1-frr"]
end
l2core ---|L2| bs
l2core ---|L2| bf
l2core ---|L2| bj
bs --- ls & lj & lf
bf --- ls & lj & lf
bj --- ls & lj & lf
classDef srsim fill:#fde7c8,stroke:#ea580c,stroke-width:3px,color:#000
classDef frr fill:#fde7c8,stroke:#2563eb,stroke-width:3px,color:#000
classDef junos fill:#fde7c8,stroke:#16a34a,stroke-width:3px,color:#000
classDef other fill:#e5e7eb,stroke:#6b7280,color:#000
class bs srsim
class bf frr
class bj junos
class l2core other
linkStyle 3,4,5 stroke:#ea580c,stroke-width:2px
linkStyle 6,7,8 stroke:#2563eb,stroke-width:2px
linkStyle 9,10,11 stroke:#16a34a,stroke-width:2px
Before the drain, every L1 router has three equal-cost next hops for 0.0.0.0/0, one per border. The drain of one border then shows in the next-hop list: if the drained router leaves the list, the drain works for that access router. (I set ecmp 3 on the SR OS access router so that it can install three next hops.)
| Drained border | Mode | In its LSP | l1 SR OS | l1 Junos | l1 FRR |
|---|---|---|---|---|---|
| SR OS | overload bit | ATT + OL |
leaves the list | stays in the list | stays in the list |
| SR OS | max-metric | no flag, links at 16777214 |
leaves | leaves | leaves |
| SR OS | overload bit + suppress-attached-bit |
OL |
leaves | leaves | leaves |
| FRR | overload bit | ATT + OL |
leaves | stays | stays |
| FRR | advertise-high-metrics |
ATT, links at 16777215 |
stays | leaves | stays |
| FRR | overload bit + no attached-bit send |
OL |
leaves | leaves | leaves |
| FRR | advertise-high-metrics + no attached-bit send |
links at 16777215 |
leaves | leaves | leaves |
| Junos | overload bit | OL |
leaves | leaves | leaves |
| Junos | advertise-high-metrics |
no flag, links at 16777214 |
leaves | leaves | leaves |
The cells in bold are the problem: the router was drained, and the access router still sends inter-area traffic to it, next to the other borders. The two rows with no attached-bit send come from the same lab without the Junos border, so with two exits (SR OS and FRR) and the same three L1 routers.
Three things come out of this table:
- The claim of the first post is confirmed. With the overload bit from SR OS or FRR, the Junos and FRR access routers keep the overloaded router as a next hop. Only the SR OS access router moves away.
- Junos is the safest sender, in both modes, because it clears
ATT. SR OS is safe with the maximum metric or withsuppress-attached-bit. - FRR drains for all three only without
ATT. Its overload bit works for SR OS access routers only, andadvertise-high-metricsfor Junos access routers only. Withno attached-bit send, both modes drain for every access router.
The FRR advertise-high-metrics row is the odd one: only the Junos access router leaves. SR OS and Junos clear ATT in their maximum-metric mode, so the border is no longer a default-route candidate, whatever the metric. FRR keeps ATT, and the cost from an L1 router to the border is the L1 router's own link metric (10), which does not change.
So the border can only leave if it becomes unreachable. FRR advertises its links toward the L1 routers at 16777215, a value that RFC 5305 excludes from the SPF. Junos applies this to the links of the border, the two-way connectivity check fails, and the border is unreachable. SR OS and FRR still use those links, so they keep the border.
The maximum metric is not one value¶
SR OS and Junos advertise 16777214 (2^24 - 2) on every link in maximum metric mode. FRR advertises 16777215 (2^24 - 1), on the links and on its own prefixes:
frr-border.00-00 * 149 0x0000000a 0xabf3 1125 1/0/0
Extended Reachability: 0000.0002.0002.00 (Metric: 16777215)
Extended Reachability: 0000.0002.0003.00 (Metric: 16777215)
Extended Reachability: 0000.0002.0004.00 (Metric: 16777215)
Extended IP Reachability: 10.2.0.0/30 (Metric: 16777215)
One unit of difference matters, because the RFCs give the two values two different meanings. RFC 5305, section 3:
If a link is advertised with the maximum link metric (2^24 - 1), this link MUST NOT be considered during the normal SPF computation.
And RFC 5443, section 2, which is about LDP IGP synchronization and also needs a link that stays usable as a last resort:
In the case of ISIS, the maximum metric value is 2^24-2 (0xFFFFFE). Indeed, if a link is configured with 2^24-1 (the maximum link metric per [RFC5305]), then this link is not advertised in the topology. It is important to keep the link in the topology to allow IP traffic to use the link as a last resort in case of massive failure.
So 2^24 - 2 means: avoid this link, but keep it as a last resort. And 2^24 - 1 means: this link is not part of the SPF. That is exactly the choice that a metric-based drain needs (RFC 3277, section 4: with metrics, "transit paths may still be calculated through the router").
Who obeys the RFC?¶
To see who follows RFC 5305, I set the value 16777215 on one link and looked at the default route of the L1 router. A link has two directions, so there are two cases:
- Forward: the L1 router advertises
16777215for its own link to the border. It uses this metric to reach the border. - Reverse: the border advertises
16777215for its link to the L1 router. The L1 router reads it in the LSP of the border. It does not add to the cost of reaching the border, but it matters for the two-way connectivity check.
| Case | Metric | SR OS L1 | Junos L1 | FRR L1 |
|---|---|---|---|---|
| Forward | 16777215 |
no default route | no default route | default route, metric 16777215 |
| Reverse (SR OS border) | 16777215 |
default route | no default route | default route |
| Either | 16777214 |
default route | default route | default route |
With 16777214, all three behave the same. With 16777215:
- Junos follows the RFC in both directions. The link is out of the SPF, and in the reverse case the two-way connectivity check fails, so the border is unreachable. This is why a drained FRR border disappears for Junos.
- SR OS follows it for its own links, not for the entry of its neighbor. RFC 5305 does not say if the two-way check is part of "the normal SPF computation", so this is a different reading, not a clear fault. In my opinion, Junos has the better reading. ISO/IEC 10589, section 7.2.8.2, puts the check in the decision process: it "shall not utilise a link between two Intermediate Systems unless both ISs report the link", and "Reporting the link indicates that it has a defined value for at least the default routeing metric". And RFC 5443 says that a link at
2^24 - 1"is not advertised in the topology" (in section 2, this means: not part of the topology that the SPF uses), so its entry should not prove a two-way adjacency either. - FRR does not follow it at all. A link with
16777215is used like any other link, with a huge metric.
Back to the drained FRR border. The value 16777215 changes the result on Junos only: Junos applies the RFC to the links of the border and loses the router. SR OS and FRR access routers keep the default route through it, because FRR keeps ATT in this mode and their own link metric to the border does not change. With 16777214, all three access vendors would keep the default route through a border that sets ATT, as the reverse test shows.
The bug in FRR¶
I reduced it to two FRR containers on a Docker network, with one IS-IS adjacency and wide metrics. On a 10.2.5 image and on the current master image (10.8.0-dev, version string git20261002), r1 has a route to the loopback of r2 with the metric 20:
Then the metric of the link of r1, in the CLI that accepts (0-16777215):
Extended Reachability: 0000.0000.0002.00 (Metric: 16777215)
Known via "isis", distance 115, metric 16777225, best
The link is used: 10 + 16777215. According to RFC 5305 the route should be gone. The same happens for a link that a neighbor advertises. With a third router r3 behind r2, and 16777215 on the link of r2 toward r3, r1 still reaches r3 through it, with the metric 16777235. And advertise-high-metrics on r2 writes 16777215 on its links and on its loopback prefix:
In the source, the SPF adds the metric of each IS reachability entry to the path cost without any check (isis_spf_process_lsp()), and isis_area_advertise_high_metrics_set() sets MAX_WIDE_LINK_METRIC (0x00FFFFFF). FRR already has the right constant, ISIS_WIDE_METRIC_INFINITY (0xFFFFFE), for its LDP-IGP synchronization. And the FRR documentation states the value ("for wide metrics, 16777215"), and the topotest of the feature expects it.
The clear deviation is the SPF. The two should be discussed together, because fixing the SPF alone changes what advertise-high-metrics does on FRR peers. And the right value alone does not make advertise-high-metrics drain the default route unless no attached-bit send is set: FRR keeps the attached bit in this mode, where SR OS and Junos clear it.
Who runs FRR¶
You may run FRR without calling it FRR. Several products use it as their routing engine:
| Product | Where FRR runs |
|---|---|
| VyOS (router) | all routing |
| Netgate TNSR (router) | dynamic routing |
| Proxmox VE 9 (virtualization) | SDN fabrics |
| NVIDIA Cumulus Linux (switch) | all routing |
| Palo Alto Networks PAN-OS (firewall) | frr 8.4_dpd in the open-source listing of PAN-OS 12.1 (listed since PAN-OS 10.1) |
Where the bug stands¶
I found no earlier report. The two related items are the PR that added the feature in 2023 and a 2024 PR about the 0xFFFFFE flag of LFA. I opened issue 23517 with a reproduction on two and three FRR routers, the multi-vendor results, the RFC quotes and a suggested fix: ignore the value 0xFFFFFF in the SPF, use 0xFFFFFE in advertise-high-metrics, and update the topotest and the documentation.
On 2026-10-04, a contributor confirmed the issue and opened PR 23527: the SPF ignores links at 16777215, and advertise-high-metrics is unchanged. A maintainer approved the PR on 2026-10-06. At the time of writing, it is not merged yet, and the issue is still open.
What to do on your network¶
The commands that drained for every access vendor in the labs with several exits, per sender. A drain that clears ATT removes the default route of the access routers that use the drained router, so it assumes that they have another exit, as in that lab.
| Drained router | What drains for every access vendor | What does not |
|---|---|---|
| Junos | set protocols isis overload, with or without advertise-high-metrics (it clears ATT) |
nothing found |
| SR OS | overload max-metric true, or overload with suppress-attached-bit true |
plain overload (SR OS access routers only) |
| FRR | set-overload-bit or advertise-high-metrics, with no attached-bit send |
the same commands without it: set-overload-bit (SR OS access routers only), advertise-high-metrics (Junos access routers only) |
If you cannot change the sender, the receiver can ignore the attached bit (ignore-attached-bit on SR OS and Junos, attached-bit receive ignore on FRR), as in the first post. It ignores the bit of every neighbor, not only the drained ones, so it needs a static default route or another exit.
Lessons¶
- A drain is a pair of vendors, not a command. The sender decides the flags and the metrics in the LSP, the receiver decides what to do with them. Test every pair that you run: with the overload bit, a Junos border drains for every access router, but an FRR border keeps inter-area traffic from Junos and FRR access routers.
- A similar configuration is not a similar behavior. Each vendor has an overload command and a maximum-metric mode, but not with the same defaults: in overload, SR OS and FRR keep
ATTand Junos clears it; in maximum-metric mode, SR OS and Junos clearATT, FRR keeps it and uses another metric value. In a multi-vendor network, check what each command really sends, not only its name. - Look at the LSP, not at the command.
show router isis databaseandshow isis databasetell you at a glance ifATTandOLare there. In maximum metric mode there is noOLflag: look at the metric of the neighbors (16777214, or16777215on FRR). - Close to the same value is not the same value. One unit of metric,
2^24 - 2against2^24 - 1, is the difference between "very costly" and "not in the topology", and only the RFC text says so. - A lab finds in an afternoon what an incident finds at night. Twenty containers in two labs, a script, and the question "what does the other vendor do with this?" were enough.
The lab and the scripts are in labs/isis_interop_drain. If you run it on other releases, or with another vendor that speaks IS-IS, I would like to know what you see.