fix: damp the DNS watchdog when re-pins never stick #170

Merged
mysticalsoap merged 1 commit from watchdog-flap-damping into trunk 2026-08-26 14:52:14 -04:00
Owner

Problem

The #95 watchdog re-pins the primary whenever its direct probe succeeds while resolved's selection sits on the fallback — and every re-pin ends in a host-wide resolvectl flush-caches. When the primary answers probes but resolved's own queries to it die (the #118 condition), the rescue becomes a permanent loop: re-pin + full flush every interval, undone within seconds. Observed 2026-08-24: seven re-pins between 20:08 and 20:17, Discord visibly in-and-out on every cold cache.

Fix

watch_dns_primary now remembers that it just re-pinned and counts a strike when selection is back off the primary by the next check.

  • After 3 consecutive re-pins fail to hold: warn once ("it answers probes but not resolved's own queries") and slow the retry cadence to 20 × interval (5 min at the default 15 s).
  • Each slow-cadence re-pin is followed by one short interval to see whether it stuck, so real recovery is noticed promptly.
  • A re-pin that holds — or selection returning on its own — resets the counter and cadence, with an info line when it ends a backed-off stretch.
  • The flush stays (it is what drops the fallback's stale public answers on a successful rescue), but damping bounds it to one per slow tick instead of one per interval.

Verification

  • New TestWatchDnsPrimaryDamping drives the loop with a fake stop event that records every wait: cadence goes [1,1,1,1] then alternates [20,1]; the warning fires exactly once; re-pins continue while backed off (damped, not abandoned); a sticking re-pin resets to the fast cadence.
  • Existing TestWatchDnsPrimary cases unchanged and passing.
  • Full suite: 378 passed + tests/test_mgmt.py (12) run separately; ruff E9,F63,F7,F82 and compileall clean.

Closes #166

🤖 Generated with Claude Code

**Problem** The #95 watchdog re-pins the primary whenever its direct probe succeeds while resolved's selection sits on the fallback — and every re-pin ends in a host-wide `resolvectl flush-caches`. When the primary answers probes but resolved's own queries to it die (the #118 condition), the rescue becomes a permanent loop: re-pin + full flush every interval, undone within seconds. Observed 2026-08-24: seven re-pins between 20:08 and 20:17, Discord visibly in-and-out on every cold cache. **Fix** `watch_dns_primary` now remembers that it just re-pinned and counts a strike when selection is back off the primary by the next check. - After 3 consecutive re-pins fail to hold: warn once ("it answers probes but not resolved's own queries") and slow the retry cadence to 20 × interval (5 min at the default 15 s). - Each slow-cadence re-pin is followed by one short interval to see whether it stuck, so real recovery is noticed promptly. - A re-pin that holds — or selection returning on its own — resets the counter and cadence, with an info line when it ends a backed-off stretch. - The flush stays (it is what drops the fallback's stale public answers on a successful rescue), but damping bounds it to one per slow tick instead of one per interval. **Verification** - New `TestWatchDnsPrimaryDamping` drives the loop with a fake stop event that records every wait: cadence goes `[1,1,1,1]` then alternates `[20,1]`; the warning fires exactly once; re-pins continue while backed off (damped, not abandoned); a sticking re-pin resets to the fast cadence. - Existing `TestWatchDnsPrimary` cases unchanged and passing. - Full suite: 378 passed + `tests/test_mgmt.py` (12) run separately; ruff E9,F63,F7,F82 and `compileall` clean. Closes #166 🤖 Generated with [Claude Code](https://claude.com/claude-code)
fix: damp the DNS watchdog when re-pins never stick
Some checks failed
ci / test (pull_request) Successful in 25s
ci / test (push) Failing after 5s
489885742c
The #95 watchdog re-pins the primary whenever its direct probe succeeds
while resolved's selection sits on the fallback, and every re-pin ends
in a full cache flush. When the primary answers probes but resolved's
own queries to it die (the #118 condition), that rescue became a
permanent loop: a re-pin plus a host-wide flush every interval, each
undone within seconds, observed 2026-08-24 as seven re-pins in ten
minutes with Discord visibly stalling on every cold cache.

watch_dns_primary now remembers that it just re-pinned and counts a
strike when selection is back off the primary by the next check. After
strikes_max (3) consecutive re-pins fail to hold, it warns once and
slows the retry cadence to backoff * interval (20 * 15s = 5 min); each
slow-cadence re-pin is followed by one short interval to see whether it
stuck. A re-pin that holds, or the selection returning on its own,
resets the cadence and the counter. The flush itself stays - it is
still what drops the fallback's stale public answers on a successful
rescue - but the damping bounds it to one per slow tick instead of one
per interval.

Closes #166

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
mysticalsoap deleted branch watchdog-flap-damping 2026-08-26 14:52:14 -04:00
Sign in to join this conversation.
No description provided.