fix: damp the DNS watchdog when re-pins never stick #170
Loading…
Reference in a new issue
No description provided.
Delete branch "watchdog-flap-damping"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Problem
The #95 watchdog re-pins the primary whenever its direct probe succeeds while resolved's selection sits on the fallback — and every re-pin ends in a host-wide
resolvectl flush-caches. When the primary answers probes but resolved's own queries to it die (the #118 condition), the rescue becomes a permanent loop: re-pin + full flush every interval, undone within seconds. Observed 2026-08-24: seven re-pins between 20:08 and 20:17, Discord visibly in-and-out on every cold cache.Fix
watch_dns_primarynow remembers that it just re-pinned and counts a strike when selection is back off the primary by the next check.Verification
TestWatchDnsPrimaryDampingdrives the loop with a fake stop event that records every wait: cadence goes[1,1,1,1]then alternates[20,1]; the warning fires exactly once; re-pins continue while backed off (damped, not abandoned); a sticking re-pin resets to the fast cadence.TestWatchDnsPrimarycases unchanged and passing.tests/test_mgmt.py(12) run separately; ruff E9,F63,F7,F82 andcompileallclean.Closes #166
🤖 Generated with Claude Code