alt-DNS at a local address intermittently unreachable on the tunnel link #118

Open
opened 2026-08-22 12:59:51 -04:00 by mysticalsoap · 4 comments
Owner

Observed 2026-08-21 (live). With alt_dns1 pointing at this host's own resolver (192.168.0.205) applied to the tunnel link with ~., resolved's query traffic to the primary was observed taking 192.168.0.205 via 10.96.0.1 dev tun_aqomui — into the VPN, toward an exit that cannot route an RFC1918 destination. The primary therefore always failed, selection fell to the public fallback and stayed there (compounded by the watchdog bug), and split-horizon resolution was dead host-wide while any tunnel was up.

Observed 2026-08-22 (live, same config). Not reproducing: tunnel up, ~. on tun_aqomui, CurrentDNSServer: 192.168.0.205, split-horizon names answered correctly in ~5ms. Plain routing agrees it should work — the destination is the host's own address, so the local table (or the connected /24 for a true LAN resolver) outranks any tunnel default; only an oif-bound socket would be forced into the tunnel.

So the failure is conditional, not by-construction, and the trigger is unknown. Candidates worth ruling in/out: whether resolved binds link-scoped query sockets to the link's ifindex in some versions/paths and not others, a race at link setup before routes settle, or a difference between the code path that applied the DNS (aqomui set_dns vs OpenVPN dns-updown).

The #34 split-horizon recipe (alt-DNS = local resolver, public fallback second) depends on this path being reliable — when it degrades, the failure is silent public answers, indistinguishable from the services being down. Needs a reproduction before 0.9.1, or at minimum a documented diagnostic (resolvectl status <tun> | grep 'Current DNS').

**Observed 2026-08-21 (live).** With `alt_dns1` pointing at this host's own resolver (`192.168.0.205`) applied to the tunnel link with `~.`, resolved's query traffic to the primary was observed taking `192.168.0.205 via 10.96.0.1 dev tun_aqomui` — into the VPN, toward an exit that cannot route an RFC1918 destination. The primary therefore always failed, selection fell to the public fallback and stayed there (compounded by the watchdog bug), and split-horizon resolution was dead host-wide while any tunnel was up. **Observed 2026-08-22 (live, same config).** Not reproducing: tunnel up, `~.` on `tun_aqomui`, `CurrentDNSServer: 192.168.0.205`, split-horizon names answered correctly in ~5ms. Plain routing agrees it should work — the destination is the host's own address, so the `local` table (or the connected /24 for a true LAN resolver) outranks any tunnel default; only an oif-bound socket would be forced into the tunnel. **So the failure is conditional, not by-construction, and the trigger is unknown.** Candidates worth ruling in/out: whether resolved binds link-scoped query sockets to the link's ifindex in some versions/paths and not others, a race at link setup before routes settle, or a difference between the code path that applied the DNS (aqomui `set_dns` vs OpenVPN dns-updown). The #34 split-horizon recipe (alt-DNS = local resolver, public fallback second) depends on this path being reliable — when it degrades, the failure is silent public answers, indistinguishable from the services being down. Needs a reproduction before 0.9.1, or at minimum a documented diagnostic (`resolvectl status <tun> | grep 'Current DNS'`).
Author
Owner

Reproduced hard on 2026-08-24, all evening, with a strong new signal. Every tunnel-up shows the same ladder in resolved's journal: ~7s after aqomui sets the tun link's DNS, 'Using degraded feature set UDP instead of UDP+EDNS0 for DNS server 192.168.0.205', then failover to 9.9.9.9; the #95 watchdog's direct probe (unbound socket, LAN route) succeeds, re-pins, and resolved fails again — seven flips between 20:08 and 20:17, each with a full flush-caches. Same pattern at 17:50–18:00 and 22:47. The 'grace period over, resuming full feature set' entries at 18:01:22 and 20:17:40 are followed by an immediate re-degrade, so the path to .205 is consistently dead from resolved's sockets while the tunnel is up, consistently fine from the watchdog's.

That asymmetry is the bound-socket candidate from the issue: resolved binding link-scoped queries to tun_aqomui's ifindex would force RFC1918-bound packets into the tunnel (no route can rescue an oif-bound socket), while the watchdog probes unbound. Next tunnel session, two captures settle it: 'tcpdump -ni tun_aqomui udp port 53' vs 'tcpdump -ni enp5s0 udp port 53 and host 192.168.0.205' — if resolved's queries to .205 appear on tun and never on enp5s0, the mechanism is confirmed and a per-link LAN-address primary can't work by construction.

User-visible cost meanwhile: the flap loop caused multi-second resolution stalls plus a cold cache every ~90s — surfaced as Discord dropping in and out through the tunnel. The watchdog's contribution (no flap damping, unconditional flush) split out as its own issue.

🤖 Generated with Claude Code

Reproduced hard on 2026-08-24, all evening, with a strong new signal. Every tunnel-up shows the same ladder in resolved's journal: ~7s after aqomui sets the tun link's DNS, 'Using degraded feature set UDP instead of UDP+EDNS0 for DNS server 192.168.0.205', then failover to 9.9.9.9; the #95 watchdog's direct probe (unbound socket, LAN route) succeeds, re-pins, and resolved fails again — seven flips between 20:08 and 20:17, each with a full flush-caches. Same pattern at 17:50–18:00 and 22:47. The 'grace period over, resuming full feature set' entries at 18:01:22 and 20:17:40 are followed by an immediate re-degrade, so the path to .205 is consistently dead from resolved's sockets while the tunnel is up, consistently fine from the watchdog's. That asymmetry is the bound-socket candidate from the issue: resolved binding link-scoped queries to tun_aqomui's ifindex would force RFC1918-bound packets into the tunnel (no route can rescue an oif-bound socket), while the watchdog probes unbound. Next tunnel session, two captures settle it: 'tcpdump -ni tun_aqomui udp port 53' vs 'tcpdump -ni enp5s0 udp port 53 and host 192.168.0.205' — if resolved's queries to .205 appear on tun and never on enp5s0, the mechanism is confirmed and a per-link LAN-address primary can't work by construction. User-visible cost meanwhile: the flap loop caused multi-second resolution stalls plus a cold cache every ~90s — surfaced as Discord dropping in and out through the tunnel. The watchdog's contribution (no flap damping, unconditional flush) split out as its own issue. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Author
Owner

Comparative finding from the Proton desktop app (its custom-DNS setting ran this exact split-horizon config 'flawlessly', per live use): reading the installed app source (proton-vpn-gtk-app 4.16.5, both the WG and OpenVPN NM backends), custom DNS puts the server on the tunnel connection as dns=[192.168.0.205] with dns-priority=-1500 and ignore-auto-dns — alone, with no fallback server, and no explicit search domain in the custom branch. So the app also served .205 from the tunnel link, and it worked.

Two consequences. First, the intermittent-unreachability trigger can't be 'per-link LAN-address DNS never works' — the app era plus the 8-22 observation both show it working. Second, the pair-vs-single difference explains the felt contrast: with one server, resolved's sticky failover has nowhere to go — a blip is a failed query and a same-server retry, self-healing and invisible. Our [.205, 9.9.9.9] pair is what converts a blip into strand→re-pin→flush→flap.

Also relevant: the app's flawless run was WireGuard; aqomui runs OpenVPN+DCO, which has prior fwmark-inheritance history on this machine — protocol is back on the candidate list for the intermittent loss. The tcpdump discriminator from the previous comment still decides it.

Host resolver topology, confirmed while investigating: /etc/resolv.conf is NM-written pointing at 127.0.0.53 and nsswitch has resolve [!UNAVAIL=return] — every resolver path on the host goes through systemd-resolved, so per-link behavior governs all applications.

🤖 Generated with Claude Code

Comparative finding from the Proton desktop app (its custom-DNS setting ran this exact split-horizon config 'flawlessly', per live use): reading the installed app source (proton-vpn-gtk-app 4.16.5, both the WG and OpenVPN NM backends), custom DNS puts the server on the tunnel connection as dns=[192.168.0.205] with dns-priority=-1500 and ignore-auto-dns — **alone, with no fallback server**, and no explicit search domain in the custom branch. So the app also served .205 from the tunnel link, and it worked. Two consequences. First, the intermittent-unreachability trigger can't be 'per-link LAN-address DNS never works' — the app era plus the 8-22 observation both show it working. Second, the pair-vs-single difference explains the felt contrast: with one server, resolved's sticky failover has nowhere to go — a blip is a failed query and a same-server retry, self-healing and invisible. Our [.205, 9.9.9.9] pair is what converts a blip into strand→re-pin→flush→flap. Also relevant: the app's flawless run was WireGuard; aqomui runs OpenVPN+DCO, which has prior fwmark-inheritance history on this machine — protocol is back on the candidate list for the intermittent loss. The tcpdump discriminator from the previous comment still decides it. Host resolver topology, confirmed while investigating: /etc/resolv.conf is NM-written pointing at 127.0.0.53 and nsswitch has resolve [!UNAVAIL=return] — every resolver path on the host goes through systemd-resolved, so per-link behavior governs all applications. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Author
Owner

Root cause found — packet capture (2026-08-26 11:52–12:10, -i any, the 11:53–12:02 flap storm included) plus the resolved journal settle it, and it is not per-link socket binding.

Zero packets to 192.168.0.205 ever entered tun_aqomui. resolved's queries to the primary always take the local docker path (DNAT to the AdGuard container on br-proxy) and get sub-millisecond answers — for names AdGuard can answer locally. The failing class is names needing recursion: AdGuard's own upstream traffic rides the tunnel (host default route; its upstream flows leave masqueraded out tun_aqomui), and the capture shows the tunnel intermittently blackholing UDP — resolved's direct 9.9.9.9 queries out tun with no replies at 11:59:22–12:00:05, matching the journal's 9.9.9.9 UDP→TCP downgrade at 12:00:39. In those windows the primary 'fails' for real-world names (letterboxd.com retried 12s; Discord's voice endpoint c-iad08-ea1b96a9.discord.media answered after 7s) while still acing split-horizon names (wildcard rewrite, no upstream) — which is exactly the class the #95 watchdog probes, so the watchdog kept re-pinning + flushing against a genuinely broken resolver path. Discord's 'in and out' was the same window twice over: voice UDP on the lossy tunnel plus DNS stalls on reconnect.

Also explained: the 'degraded feature set ~5s after every tunnel-up' signature — the default-route flip breaks AdGuard's established upstream connections (conntrack mismatch), a short re-establishment blip that the [primary, fallback] pair escalates into a sticky failover. And the degrades appear on the bare LAN link too (journal 08-25 13:13:41, 21:46:55 — no tunnel), consistent with upstream blips, not link routing.

Remaining question, now well-defined: what causes the tunnel's UDP loss windows — uplink saturation (shared 43 Mbps up with the gluetun/qbit stack), the Proton server, or DCO. Next: correlate the 11:53–12:02 window (and 08-24 evening) against host egress in Grafana.

Consequences for the fixes: #167 (single server, no fallback) remains the right mitigation — it removes the failover/flap amplifier. The watchdog's probe design is invalidated either way: probing a locally-rewritten name can never detect the actual failure mode (relevant to #166 if a pair remains anywhere).

🤖 Generated with Claude Code

Root cause found — packet capture (2026-08-26 11:52–12:10, -i any, the 11:53–12:02 flap storm included) plus the resolved journal settle it, and it is not per-link socket binding. **Zero packets to 192.168.0.205 ever entered tun_aqomui.** resolved's queries to the primary always take the local docker path (DNAT to the AdGuard container on br-proxy) and get sub-millisecond answers — for names AdGuard can answer locally. The failing class is names needing recursion: **AdGuard's own upstream traffic rides the tunnel** (host default route; its upstream flows leave masqueraded out tun_aqomui), and the capture shows the tunnel intermittently blackholing UDP — resolved's direct 9.9.9.9 queries out tun with no replies at 11:59:22–12:00:05, matching the journal's 9.9.9.9 UDP→TCP downgrade at 12:00:39. In those windows the primary 'fails' for real-world names (letterboxd.com retried 12s; Discord's voice endpoint c-iad08-ea1b96a9.discord.media answered after 7s) while still acing split-horizon names (wildcard rewrite, no upstream) — which is exactly the class the #95 watchdog probes, so the watchdog kept re-pinning + flushing against a genuinely broken resolver path. Discord's 'in and out' was the same window twice over: voice UDP on the lossy tunnel plus DNS stalls on reconnect. Also explained: the 'degraded feature set ~5s after every tunnel-up' signature — the default-route flip breaks AdGuard's established upstream connections (conntrack mismatch), a short re-establishment blip that the [primary, fallback] pair escalates into a sticky failover. And the degrades appear on the bare LAN link too (journal 08-25 13:13:41, 21:46:55 — no tunnel), consistent with upstream blips, not link routing. Remaining question, now well-defined: **what causes the tunnel's UDP loss windows** — uplink saturation (shared 43 Mbps up with the gluetun/qbit stack), the Proton server, or DCO. Next: correlate the 11:53–12:02 window (and 08-24 evening) against host egress in Grafana. Consequences for the fixes: #167 (single server, no fallback) remains the right mitigation — it removes the failover/flap amplifier. The watchdog's probe design is invalidated either way: probing a locally-rewritten name can never detect the actual failure mode (relevant to #166 if a pair remains anywhere). 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Author
Owner

Trigger found for the post-reboot loss windows, on the wire: with the tunnel disconnected, gluetun's WireGuard endpoint flow leaves via enp5s0; connected, the same flow leaves via tun_aqomui as the tunnel address — the container VPN and its full qbittorrent load double-tunnel through the OpenVPN connection. The bypass_networks exemption that should prevent this wasn't installed: the machine rebooted 2026-08-25 morning and nothing rebuilds the bypass after a boot (#168) — last 'Creating bypass' was Aug 24 22:46, so every session investigated on Aug 25–26 ran with the exemption (and the whole bypass) absent. The 15 Mbps seeding bursts riding inside the tunnel line up with the flap onsets, and gluetun's own collapses were its WireGuard dying inside the congested tunnel.

Still open: the Aug 24 evening sessions predate the reboot and had a freshly built bypass (rebuilds logged 17:17, 17:18, 20:06), yet flapped too — either the exemption was ineffective then for a reason not yet seen, or a second, milder trigger exists (the .205 EDNS0 degrades appearing with no tunnel at all suggest background upstream blips that the fallback pair amplifies, #167). After #168 is fixed (interim: re-apply options to rebuild), a session with the exemption verified live (aqomui_bypass_net chain present) plus the ping pair (tunnel vs -I enp5s0) during any remaining stutter closes that half.

🤖 Generated with Claude Code

Trigger found for the post-reboot loss windows, on the wire: with the tunnel disconnected, gluetun's WireGuard endpoint flow leaves via enp5s0; connected, the same flow leaves via tun_aqomui as the tunnel address — the container VPN and its full qbittorrent load double-tunnel through the OpenVPN connection. The bypass_networks exemption that should prevent this wasn't installed: the machine rebooted 2026-08-25 morning and nothing rebuilds the bypass after a boot (#168) — last 'Creating bypass' was Aug 24 22:46, so every session investigated on Aug 25–26 ran with the exemption (and the whole bypass) absent. The 15 Mbps seeding bursts riding inside the tunnel line up with the flap onsets, and gluetun's own collapses were its WireGuard dying inside the congested tunnel. Still open: the Aug 24 evening sessions predate the reboot and had a freshly built bypass (rebuilds logged 17:17, 17:18, 20:06), yet flapped too — either the exemption was ineffective then for a reason not yet seen, or a second, milder trigger exists (the .205 EDNS0 degrades appearing with no tunnel at all suggest background upstream blips that the fallback pair amplifies, #167). After #168 is fixed (interim: re-apply options to rebuild), a session with the exemption verified live (aqomui_bypass_net chain present) plus the ping pair (tunnel vs -I enp5s0) during any remaining stutter closes that half. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Sign in to join this conversation.
No milestone
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
mysticalsoap/aqomui#118
No description provided.