No egress shaping: a saturated uplink collapses downstream throughput #156

Closed
opened 2026-08-21 12:11:48 -04:00 by mysticalsoap · 2 comments
Owner

The connection is heavily asymmetric and nothing shapes the upload, so any sustained outbound transfer starves downstream ACKs and collapses the whole link for every other consumer.

Measured 2026-08-21 from 30d of cadvisor counters on gluetun:

Direction Observed
Down ~927 Mbps peak (115.89 MB/s)
Up ~43.5 Mbps hard ceiling (5.44 MB/s peak, p99 5.0–5.37)

The upload distribution sits flat on its own maximum rather than tailing off, which is a hard cap and not a workload artifact. Both directions cross the same tunnel, so the VPN isn't the constraint — the WAN link is simply that asymmetric.

Impact when the uplink is full: downstream drops to 84–175 KB/s, a ~99.98% degradation on a gigabit downlink. That's textbook bufferbloat — ACKs queue behind the upload. It took out CI repeatedly (mysticalsoap/aqomui#115) but affects everything: pulls, backups, streaming, any interactive use.

This is currently latent rather than fixed. The immediate trigger was public-tracker seeding, resolved by a qbit_manage share-limit change, but nothing prevents a restic run, a private-tracker seed spike, or any other sustained upload from doing it again.

Why the host doesn't already absorb it

enp5s0 has fq_codel on its hardware queues, but it drains at 1 Gbps into the LAN switch, so its queue never builds and fq_codel never engages. The buffer that actually fills is in the ISP modem, on the ~43 Mbps WAN egress.

Fix

Make the host the bottleneck instead of the modem, by shaping egress slightly below the real uplink rate so the queue forms somewhere with a competent AQM:

tc qdisc replace dev enp5s0 root cake bandwidth 38Mbit

Roughly 90–95% of measured line rate; needs tuning against a real bufferbloat test. This only governs traffic originating from this host — which is where essentially all the upload comes from, so it addresses the actual offender without touching the router.

To make it durable it needs a home: a systemd unit or a NetworkManager dispatcher script in dotfiles, applied on link-up. Router-side CAKE/SQM would additionally cover other devices on the LAN and is worth considering if the router supports it, but it isn't required to fix this.

Verification is blocked on a monitoring gap

There are no host WAN interface metrics at all. node-exporter runs in a container netns, so its device label only has eth0/lo — enp5s0 is absent and no node_network_* series exist for it. The figures above had to be inferred from a container's counters.

Without that, there's no way to confirm a shaper is working, no alert when the uplink pins, and no way to notice the next occurrence. Worth splitting into its own issue if it's more than a one-line fix.

The connection is heavily asymmetric and nothing shapes the upload, so any sustained outbound transfer starves downstream ACKs and collapses the whole link for every other consumer. Measured 2026-08-21 from 30d of cadvisor counters on `gluetun`: | Direction | Observed | |---|---| | Down | ~927 Mbps peak (115.89 MB/s) | | Up | **~43.5 Mbps hard ceiling** (5.44 MB/s peak, p99 5.0–5.37) | The upload distribution sits flat on its own maximum rather than tailing off, which is a hard cap and not a workload artifact. Both directions cross the same tunnel, so the VPN isn't the constraint — the WAN link is simply that asymmetric. **Impact when the uplink is full:** downstream drops to 84–175 KB/s, a ~99.98% degradation on a gigabit downlink. That's textbook bufferbloat — ACKs queue behind the upload. It took out CI repeatedly (`mysticalsoap/aqomui#115`) but affects everything: pulls, backups, streaming, any interactive use. This is currently latent rather than fixed. The immediate trigger was public-tracker seeding, resolved by a qbit_manage share-limit change, but nothing prevents a restic run, a private-tracker seed spike, or any other sustained upload from doing it again. ## Why the host doesn't already absorb it `enp5s0` has fq_codel on its hardware queues, but it drains at 1 Gbps into the LAN switch, so its queue never builds and fq_codel never engages. The buffer that actually fills is in the ISP modem, on the ~43 Mbps WAN egress. ## Fix Make the host the bottleneck instead of the modem, by shaping egress slightly below the real uplink rate so the queue forms somewhere with a competent AQM: ``` tc qdisc replace dev enp5s0 root cake bandwidth 38Mbit ``` Roughly 90–95% of measured line rate; needs tuning against a real bufferbloat test. This only governs traffic originating from this host — which is where essentially all the upload comes from, so it addresses the actual offender without touching the router. To make it durable it needs a home: a systemd unit or a NetworkManager dispatcher script in `dotfiles`, applied on link-up. Router-side CAKE/SQM would additionally cover other devices on the LAN and is worth considering if the router supports it, but it isn't required to fix this. ## Verification is blocked on a monitoring gap There are no host WAN interface metrics at all. node-exporter runs in a container netns, so its `device` label only has `eth0`/`lo` — `enp5s0` is absent and no `node_network_*` series exist for it. The figures above had to be inferred from a container's counters. Without that, there's no way to confirm a shaper is working, no alert when the uplink pins, and no way to notice the next occurrence. Worth splitting into its own issue if it's more than a one-line fix.
Author
Owner

Deprioritise rather than cap

A flat bandwidth limit gives up capacity permanently. CAKE's diffserv4 gives the same protection without the waste — its Bulk tin has a 6.25% floor, not a ceiling, so torrents take the whole uplink when it's idle and collapse to ~6% the moment a Jellyfin stream wants it. No manual switching.

tc qdisc replace dev enp5s0 root cake bandwidth 38Mbit diffserv4 nat wash
iptables -t mangle -A FORWARD -s <gluetun> -j DSCP --set-dscp 0x01

0x01 is LE (RFC 8622) rather than CS1 — CAKE maps both to Bulk on this kernel, but CS1 is documented as inconsistent (prioritised over CS0 on some switches, deprioritised on wifi). wash clears the mark after classification so nothing nonstandard reaches the VPN provider.

This is the standard recipe, not an invention. From CAKE's own technical wiki: BitTorrent "should be directed to the Bulk class", and "the only way known to 'fix' bittorrent is to classify it somewhat, somehow, as 'background'." The same page recommends diffserv4, and notes "de-prioritization seems a good idea, prioritization not so much" — which is what this does. Nothing gets promoted, so nothing can be starved by mistake.

µTP does not make this redundant. qBittorrent enables it by default and LEDBAT targets 100ms added delay specifically so torrents yield (libtorrent, BEP 29). But it yields reliably only against TCP on the same bottleneck — peers connecting over plain TCP don't back off at all, and qBittorrent's own forum is blunt that it is not a magic bullet. The shaper is the backstop for what µTP structurally can't cover.

Correction to the original text

gluetun runs WireGuard, not OpenVPN (VPN_TYPE=wireguard, kernel implementation, peer 163.5.171.2:51820). That makes classification easier: all torrent traffic is a single outer UDP flow to one fixed endpoint, so CAKE's flow hashing already gives it one flow's share rather than letting hundreds of connections gang up. Meaningful improvement lands from the shaper alone, before any marking.

Blocked on mysticalsoap/aqomui#116

While aqomui is connected, container traffic is captured by its /1 route override and double-tunnelled — measured 111MB through gluetun reappearing as 118MB through tun_aqomui. Everything on enp5s0 is then one indistinguishable outer flow, and no DSCP marking inside it is visible to a shaper.

The shaper still helps whenever the host VPN is down, but classification only works once #116 is resolved.

## Deprioritise rather than cap A flat `bandwidth` limit gives up capacity permanently. CAKE's `diffserv4` gives the same protection without the waste — its **Bulk** tin has a 6.25% *floor*, not a ceiling, so torrents take the whole uplink when it's idle and collapse to ~6% the moment a Jellyfin stream wants it. No manual switching. ``` tc qdisc replace dev enp5s0 root cake bandwidth 38Mbit diffserv4 nat wash iptables -t mangle -A FORWARD -s <gluetun> -j DSCP --set-dscp 0x01 ``` `0x01` is LE (RFC 8622) rather than CS1 — CAKE maps both to Bulk on this kernel, but CS1 is documented as inconsistent (prioritised over CS0 on some switches, deprioritised on wifi). `wash` clears the mark after classification so nothing nonstandard reaches the VPN provider. This is the standard recipe, not an invention. From CAKE's own [technical wiki](https://www.bufferbloat.net/projects/codel/wiki/CakeTechnical/): BitTorrent "should be directed to the Bulk class", and *"the only way known to 'fix' bittorrent is to classify it somewhat, somehow, as 'background'."* The same page recommends `diffserv4`, and notes **"de-prioritization seems a good idea, prioritization not so much"** — which is what this does. Nothing gets promoted, so nothing can be starved by mistake. **µTP does not make this redundant.** qBittorrent enables it by default and LEDBAT targets 100ms added delay specifically so torrents yield ([libtorrent](https://www.libtorrent.org/utp.html), [BEP 29](https://www.bittorrent.org/beps/bep_0029.html)). But it yields reliably only against TCP on the same bottleneck — peers connecting over plain TCP don't back off at all, and qBittorrent's own forum is blunt that it is [not a magic bullet](https://forum.qbittorrent.org/viewtopic.php?t=3835). The shaper is the backstop for what µTP structurally can't cover. ## Correction to the original text gluetun runs **WireGuard**, not OpenVPN (`VPN_TYPE=wireguard`, kernel implementation, peer `163.5.171.2:51820`). That makes classification easier: all torrent traffic is a *single* outer UDP flow to one fixed endpoint, so CAKE's flow hashing already gives it one flow's share rather than letting hundreds of connections gang up. Meaningful improvement lands from the shaper alone, before any marking. ## Blocked on mysticalsoap/aqomui#116 While aqomui is connected, container traffic is captured by its `/1` route override and double-tunnelled — measured 111MB through gluetun reappearing as 118MB through `tun_aqomui`. Everything on `enp5s0` is then one indistinguishable outer flow, and no DSCP marking inside it is visible to a shaper. The shaper still helps whenever the host VPN is down, but classification only works once #116 is resolved.
Author
Owner

Blocker cleared, and a gap in the shaper design

aqomui#116 is closed. Network bypass entries (source subnet plus optional proto/port) shipped in aqomui/bypass.py. With a gluetun-subnet udp 51820 entry, gluetun's outer WireGuard flow leaves enp5s0 directly while aqomui is up, so a DSCP mark on it reaches the host shaper. Classification is no longer blocked; it is blocked on that entry being present in the aqomui config, which should be confirmed before the marking step.

The one-liner caps LAN traffic too. enp5s0 carries LAN and WAN on the same wire. A root cake bandwidth 38Mbit shapes every byte leaving the NIC, including Jellyfin streams to a client on 192.168.0.0/24 and LAN-side backups. "Only governs traffic originating from this host" is true but that includes local playback, which would be held under a single 4K bitrate.

Shape only WAN-bound packets instead:

  • iptables -t mangle -A POSTROUTING -o enp5s0 ! -d 192.168.0.0/16 -j MARK --set-mark 0x30 (plus 10/8 and 172.16/12 if anything on the LAN uses them)
  • HTB root on enp5s0: default class unshaped at line rate, one child class at 38Mbit with cake diffserv4 nat wash as its qdisc
  • tc filter ... handle 0x30 fw classid <wan class> steers marked packets into the CAKE class

CAKE's own rate limiter is unused in this layout; HTB sets the ceiling and CAKE does AQM and tins underneath it. Still needs tuning against a bufferbloat test on an idle link.

Persistence home. NetworkManager owns enp5s0 (Wired connection 1), so a dispatcher script under ~/dotfiles/system/ applied on up for that device. The mark rule has to survive the docker and aqomui firewall teardowns; aqomui#116's comment records a manual mangle rule vanishing within a day.

Order. #176 first (no way to verify a shaper without host NIC counters), then the HTB+CAKE shaper alone, then DSCP marking of the gluetun subnet as a separate step.

## Blocker cleared, and a gap in the shaper design **aqomui#116 is closed.** Network bypass entries (source subnet plus optional proto/port) shipped in `aqomui/bypass.py`. With a gluetun-subnet `udp 51820` entry, gluetun's outer WireGuard flow leaves enp5s0 directly while aqomui is up, so a DSCP mark on it reaches the host shaper. Classification is no longer blocked; it is blocked on that entry being present in the aqomui config, which should be confirmed before the marking step. **The one-liner caps LAN traffic too.** enp5s0 carries LAN and WAN on the same wire. A root `cake bandwidth 38Mbit` shapes every byte leaving the NIC, including Jellyfin streams to a client on 192.168.0.0/24 and LAN-side backups. "Only governs traffic originating from this host" is true but that includes local playback, which would be held under a single 4K bitrate. Shape only WAN-bound packets instead: - `iptables -t mangle -A POSTROUTING -o enp5s0 ! -d 192.168.0.0/16 -j MARK --set-mark 0x30` (plus 10/8 and 172.16/12 if anything on the LAN uses them) - HTB root on enp5s0: default class unshaped at line rate, one child class at 38Mbit with `cake diffserv4 nat wash` as its qdisc - `tc filter ... handle 0x30 fw classid <wan class>` steers marked packets into the CAKE class CAKE's own rate limiter is unused in this layout; HTB sets the ceiling and CAKE does AQM and tins underneath it. Still needs tuning against a bufferbloat test on an idle link. **Persistence home.** NetworkManager owns enp5s0 (`Wired connection 1`), so a dispatcher script under `~/dotfiles/system/` applied on `up` for that device. The mark rule has to survive the docker and aqomui firewall teardowns; aqomui#116's comment records a manual mangle rule vanishing within a day. **Order.** #176 first (no way to verify a shaper without host NIC counters), then the HTB+CAKE shaper alone, then DSCP marking of the gluetun subnet as a separate step.
Sign in to join this conversation.
No milestone
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
mysticalsoap/docker#156
No description provided.