change: wake NetObserver over netlink instead of polling (#122) #195

Merged
mysticalsoap merged 1 commit from netlink-observer into trunk 2026-08-28 18:08:49 -04:00
Owner

Problem

NetObserver (#111) polls default_routes() every 2 seconds: up to 2s of latency on a network transition, plus a wakeup every 2s forever. The kernel already multicasts these events over netlink. Closes #122.

Fix

work() now subscribes a raw stdlib AF_NETLINK socket to RTMGRP_LINK | RTMGRP_IPV4_ROUTE | RTMGRP_IPV6_ROUTE and blocks in select() until the kernel reports a change, instead of sleeping 2s between reads.

Design points from the issue, and how they landed:

  • Dependency choice: stdlib socket, but without hand-parsing rtattr — the messages are used purely as a wake-up call and immediately discarded; state is still re-derived from default_routes() on every read (#89 derive-at-point-of-use). Parsing (or pyroute2) would only pay for replacing the ip -j route subprocess calls, which the issue listed as a bonus, not the goal — deliberately not taken here.
  • Up means usable: preserved for free — a link-up event just triggers a read that still finds no default route, so nothing is reported until a route event backs it.
  • ENOBUFS: inherently safe — every wakeup re-reads the routing table rather than trusting the event stream, so a queue overflow is caught, logged nothing, and costs nothing. drain() swallows the error and the next read is authoritative.

The two-identical-reads debounce is unchanged: while a changed read awaits its confirming second read (which no further event would announce), the wait keeps a 2s timeout; once settled, it blocks indefinitely. poll_once() and all transition logic are untouched.

Verification

  • Full suite passes (492 tests), including new coverage for subscribe() group mask, wait() timeout selection (settled → block forever, unconfirmed change → 2s confirm window), message draining, and ENOBUFS survival.
  • ruff check --select E9,F63,F7,F82 . and python -W error::SyntaxWarning -m compileall clean.
  • Real-socket smoke test: NetObserver.subscribe() binds unprivileged with groups 0x441 and stays quiet with no network changes.
  • Not yet live-verified on the running service — needs packaging/arch/build-branch-and-install.sh plus a real link flip before merging.

🤖 Generated with Claude Code

## Problem NetObserver (#111) polls `default_routes()` every 2 seconds: up to 2s of latency on a network transition, plus a wakeup every 2s forever. The kernel already multicasts these events over netlink. Closes #122. ## Fix `work()` now subscribes a raw stdlib `AF_NETLINK` socket to `RTMGRP_LINK | RTMGRP_IPV4_ROUTE | RTMGRP_IPV6_ROUTE` and blocks in `select()` until the kernel reports a change, instead of sleeping 2s between reads. Design points from the issue, and how they landed: - **Dependency choice**: stdlib socket, but *without* hand-parsing rtattr — the messages are used purely as a wake-up call and immediately discarded; state is still re-derived from `default_routes()` on every read (#89 derive-at-point-of-use). Parsing (or pyroute2) would only pay for replacing the `ip -j route` subprocess calls, which the issue listed as a bonus, not the goal — deliberately not taken here. - **Up means usable**: preserved for free — a link-up event just triggers a read that still finds no default route, so nothing is reported until a route event backs it. - **ENOBUFS**: inherently safe — every wakeup re-reads the routing table rather than trusting the event stream, so a queue overflow is caught, logged nothing, and costs nothing. `drain()` swallows the error and the next read is authoritative. The two-identical-reads debounce is unchanged: while a changed read awaits its confirming second read (which no further event would announce), the wait keeps a 2s timeout; once settled, it blocks indefinitely. `poll_once()` and all transition logic are untouched. ## Verification - Full suite passes (492 tests), including new coverage for `subscribe()` group mask, `wait()` timeout selection (settled → block forever, unconfirmed change → 2s confirm window), message draining, and ENOBUFS survival. - `ruff check --select E9,F63,F7,F82 .` and `python -W error::SyntaxWarning -m compileall` clean. - Real-socket smoke test: `NetObserver.subscribe()` binds unprivileged with groups `0x441` and stays quiet with no network changes. - Not yet live-verified on the running service — needs `packaging/arch/build-branch-and-install.sh` plus a real link flip before merging. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
change: wake NetObserver over netlink instead of polling (#122)
All checks were successful
ci / test (pull_request) Successful in 27s
a844059ff9
The kernel multicasts every link and route change; subscribing to
RTMGRP_LINK/RTMGRP_IPV4_ROUTE/RTMGRP_IPV6_ROUTE replaces the 2-second
sleep loop with reacting to the change itself.

The messages are treated purely as a wake-up call and never parsed:
state is still re-derived from the routing table on every read (#89),
so no rtattr hand-parsing and no pyroute2 dependency, and an ENOBUFS
queue overflow costs nothing because the next read does not trust the
event stream to begin with. Parsing would only pay for replacing the
ip -j route subprocess calls, which #122 left as a bonus.

The wait blocks indefinitely once settled, but a changed read's
confirming second read has no event of its own to announce it -- that
case keeps a 2-second timeout, which is also what preserves the
two-identical-reads debounce unchanged.

A link-up without a default route still reports down: the read finds
no route until a route event backs it, which is the up-means-usable
rule #122 required kept.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
mysticalsoap force-pushed netlink-observer from a844059ff9
All checks were successful
ci / test (pull_request) Successful in 27s
to 823a86e337
Some checks failed
ci / test (pull_request) Successful in 27s
ci / test (push) Failing after 14s
2026-08-28 16:02:16 -04:00
Compare
Author
Owner

Live baseline on the installed build caught a regression before any link-flip testing: service-startup route churn produced a down-report and up-report 0.5s apart ("Lost network connection" → "Network: new connection detected"), each confirmed by two reads only milliseconds apart. Any unrelated netlink wakeup (startup churn, veth noise) could supply the confirming read immediately, collapsing the two-identical-reads debounce — and re-opening the staggered-v4/v6 double-teardown it exists to prevent.

Amended into the commit: while a change is pending, wait() now holds the full 2-second window, draining events without re-reading early. Settled state still blocks indefinitely. New tests cover the held window (events must not shorten it) alongside the existing timeout-selection cases. 493 tests pass. Needs a rebuild before live verification continues.

Live baseline on the installed build caught a regression before any link-flip testing: service-startup route churn produced a down-report and up-report 0.5s apart ("Lost network connection" → "Network: new connection detected"), each confirmed by two reads only milliseconds apart. Any unrelated netlink wakeup (startup churn, veth noise) could supply the confirming read immediately, collapsing the two-identical-reads debounce — and re-opening the staggered-v4/v6 double-teardown it exists to prevent. Amended into the commit: while a change is pending, wait() now holds the full 2-second window, draining events without re-reading early. Settled state still blocks indefinitely. New tests cover the held window (events must not shorten it) alongside the existing timeout-selection cases. 493 tests pass. Needs a rebuild before live verification continues.
Author
Owner

Correction to the previous comment: the startup down/up log pair was misattributed. "Lost network connection" at startup comes from the GUI's sync_net_state reading the service's initial state 0 before the observer's first confirmed report lands (~2s after start) — pre-existing behavior, present on trunk, not a debounce collapse. The held-window fix stands on the mechanism itself (the earlier wait() re-read on any event during the confirm window, shortening it to milliseconds under netlink noise), now locked by test_events_do_not_shorten_the_confirm_window rather than by those log lines.

Live verification on the installed build:

  • Link down at 16:45:46.145 → "connection lost - closing tunnels" at 16:45:48.202: 2.06s (event wakeup + full 2s confirm window). The poller was 2–4s.
  • Link up → report gated by carrier/DHCP (~5s) plus the 2s confirm; tunnel auto-reconnected cleanly.
  • Null events stay silent: adding/deleting a non-default route, a WireGuard connect (OWN_DEVICES filter, whose route adds now wake the observer), and docker container veth churn all produced zero reports.
  • Polling is gone: strace -f -e trace=execve on the idle service for 15s shows zero ip route execs (the poller spawned a pair every 2s).
  • No GuardedThread "network watching stopped" failure across two service restarts.
Correction to the previous comment: the startup down/up log pair was misattributed. "Lost network connection" at startup comes from the GUI's sync_net_state reading the service's initial state 0 before the observer's first confirmed report lands (~2s after start) — pre-existing behavior, present on trunk, not a debounce collapse. The held-window fix stands on the mechanism itself (the earlier wait() re-read on any event during the confirm window, shortening it to milliseconds under netlink noise), now locked by test_events_do_not_shorten_the_confirm_window rather than by those log lines. Live verification on the installed build: - Link down at 16:45:46.145 → "connection lost - closing tunnels" at 16:45:48.202: 2.06s (event wakeup + full 2s confirm window). The poller was 2–4s. - Link up → report gated by carrier/DHCP (~5s) plus the 2s confirm; tunnel auto-reconnected cleanly. - Null events stay silent: adding/deleting a non-default route, a WireGuard connect (OWN_DEVICES filter, whose route adds now wake the observer), and docker container veth churn all produced zero reports. - Polling is gone: strace -f -e trace=execve on the idle service for 15s shows zero ip route execs (the poller spawned a pair every 2s). - No GuardedThread "network watching stopped" failure across two service restarts.
mysticalsoap deleted branch netlink-observer 2026-08-28 18:08:49 -04:00
Sign in to join this conversation.
No description provided.