change: wake NetObserver over netlink instead of polling (#122) #195
Loading…
Reference in a new issue
No description provided.
Delete branch "netlink-observer"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Problem
NetObserver (#111) polls
default_routes()every 2 seconds: up to 2s of latency on a network transition, plus a wakeup every 2s forever. The kernel already multicasts these events over netlink. Closes #122.Fix
work()now subscribes a raw stdlibAF_NETLINKsocket toRTMGRP_LINK | RTMGRP_IPV4_ROUTE | RTMGRP_IPV6_ROUTEand blocks inselect()until the kernel reports a change, instead of sleeping 2s between reads.Design points from the issue, and how they landed:
default_routes()on every read (#89 derive-at-point-of-use). Parsing (or pyroute2) would only pay for replacing theip -j routesubprocess calls, which the issue listed as a bonus, not the goal — deliberately not taken here.drain()swallows the error and the next read is authoritative.The two-identical-reads debounce is unchanged: while a changed read awaits its confirming second read (which no further event would announce), the wait keeps a 2s timeout; once settled, it blocks indefinitely.
poll_once()and all transition logic are untouched.Verification
subscribe()group mask,wait()timeout selection (settled → block forever, unconfirmed change → 2s confirm window), message draining, and ENOBUFS survival.ruff check --select E9,F63,F7,F82 .andpython -W error::SyntaxWarning -m compileallclean.NetObserver.subscribe()binds unprivileged with groups0x441and stays quiet with no network changes.packaging/arch/build-branch-and-install.shplus a real link flip before merging.🤖 Generated with Claude Code
a844059ff9823a86e337Live baseline on the installed build caught a regression before any link-flip testing: service-startup route churn produced a down-report and up-report 0.5s apart ("Lost network connection" → "Network: new connection detected"), each confirmed by two reads only milliseconds apart. Any unrelated netlink wakeup (startup churn, veth noise) could supply the confirming read immediately, collapsing the two-identical-reads debounce — and re-opening the staggered-v4/v6 double-teardown it exists to prevent.
Amended into the commit: while a change is pending, wait() now holds the full 2-second window, draining events without re-reading early. Settled state still blocks indefinitely. New tests cover the held window (events must not shorten it) alongside the existing timeout-selection cases. 493 tests pass. Needs a rebuild before live verification continues.
Correction to the previous comment: the startup down/up log pair was misattributed. "Lost network connection" at startup comes from the GUI's sync_net_state reading the service's initial state 0 before the observer's first confirmed report lands (~2s after start) — pre-existing behavior, present on trunk, not a debounce collapse. The held-window fix stands on the mechanism itself (the earlier wait() re-read on any event during the confirm window, shortening it to milliseconds under netlink noise), now locked by test_events_do_not_shorten_the_confirm_window rather than by those log lines.
Live verification on the installed build: