Bring the bypass up after a reboot #169

Merged
mysticalsoap merged 1 commit from bypass-reboot into trunk 2026-08-26 14:10:46 -04:00
Owner

Problem

After a machine reboot, nothing rebuilds the bypass: reconcile_bypass needs a leftover cgroup (a fresh boot has none), service_recovered only fires when the service restarts under a running gui, and plain gui startup never called bypass(). Confirmed live: zero rebuilds for 40 hours after the Aug 25 reboot, with the configured bypass network (gluetun) double-tunneling its full seeding load through the OpenVPN connection — the mechanism behind the post-reboot share of #118's tunnel UDP-loss windows, felt as Discord dropping in and out.

Fix

Split along what each side can know:

  • Service (uid-independent half): the fwmark rules, table-11 default route and rp_filter settings move out of create_cgroup into bypass.set_routing (behavior unchanged for the cgroup path — it calls the same function). reconcile_bypass now applies routing + set_network_rules even when there is no cgroup to restore, gated on bypass=1 and an existing default route — so it lands at boot once the network is up, and re-asserts on every network-up via the existing network_changed → reconcile path. Forwarded-traffic exemptions need no cgroup owner and no longer wait for one.
  • GUI (cgroup half): bypass() is called at gui startup (after config load), not only in service_recovered; and service_recovered fires an idempotent delayed second call, covering the observed case of the first call being refused while a fresh service was still settling on the bus (the Aug 24 18:22 polkit refusal had no retry).

Verification

  • pytest green (373 passed) plus the ruff/compileall gates.
  • New coverage: reconcile with no cgroup applies routing + network rules with the derived interface/gateways; skipped when bypass is off; deferred when no default route exists (network-up retries); service_recovered fires the rebuild twice and skips it when bypass is disabled.
  • Live verification after merge: reboot, log in, connect — sudo iptables-legacy -t mangle -S | grep 51820 should show the aqomui_bypass_net rule with no manual settings-apply, and the 30-packet capture should show gluetun's WireGuard staying on enp5s0 with the tunnel up.

Closes #168

🤖 Generated with Claude Code

## Problem After a machine reboot, nothing rebuilds the bypass: `reconcile_bypass` needs a leftover cgroup (a fresh boot has none), `service_recovered` only fires when the service restarts under a running gui, and plain gui startup never called `bypass()`. Confirmed live: zero rebuilds for 40 hours after the Aug 25 reboot, with the configured bypass network (gluetun) double-tunneling its full seeding load through the OpenVPN connection — the mechanism behind the post-reboot share of #118's tunnel UDP-loss windows, felt as Discord dropping in and out. ## Fix Split along what each side can know: - **Service (uid-independent half):** the fwmark rules, table-11 default route and rp_filter settings move out of `create_cgroup` into `bypass.set_routing` (behavior unchanged for the cgroup path — it calls the same function). `reconcile_bypass` now applies routing + `set_network_rules` even when there is no cgroup to restore, gated on `bypass=1` and an existing default route — so it lands at boot once the network is up, and re-asserts on every network-up via the existing `network_changed` → reconcile path. Forwarded-traffic exemptions need no cgroup owner and no longer wait for one. - **GUI (cgroup half):** `bypass()` is called at gui startup (after config load), not only in `service_recovered`; and `service_recovered` fires an idempotent delayed second call, covering the observed case of the first call being refused while a fresh service was still settling on the bus (the Aug 24 18:22 polkit refusal had no retry). ## Verification - `pytest` green (373 passed) plus the ruff/compileall gates. - New coverage: reconcile with no cgroup applies routing + network rules with the derived interface/gateways; skipped when bypass is off; deferred when no default route exists (network-up retries); `service_recovered` fires the rebuild twice and skips it when bypass is disabled. - Live verification after merge: reboot, log in, connect — `sudo iptables-legacy -t mangle -S | grep 51820` should show the `aqomui_bypass_net` rule with no manual settings-apply, and the 30-packet capture should show gluetun's WireGuard staying on `enp5s0` with the tunnel up. Closes #168 🤖 Generated with [Claude Code](https://claude.com/claude-code)
fix: bring the bypass up after a reboot
All checks were successful
ci / test (pull_request) Successful in 26s
0883e7f551
Nothing rebuilt the bypass on a fresh boot: reconcile_bypass only
restores a bypass whose leftover cgroup survived on disk, the gui's
service_recovered only fires when the service restarts underneath a
running gui, and plain gui startup never called bypass(). Confirmed
live: rules absent for 40 hours after a reboot, and the configured
bypass network's gluetun double-tunneling its full load through the
OpenVPN connection -- the mechanism behind the post-reboot share of
the #118 UDP loss windows.

Two halves, matching what each side can know:

- The uid-independent infrastructure (fwmark rules, the table-11
  default, the forwarded-traffic marks) moves out of create_cgroup
  into bypass.set_routing, and reconcile_bypass now applies it even
  with no cgroup to restore -- at boot once a route exists, and again
  on every network-up. Forwarded traffic needs no cgroup owner and
  never had a reason to wait for one.
- The cgroup half still needs a desktop uid, so the gui calls bypass()
  at startup, not only in service_recovered -- and service_recovered
  fires a delayed second call, since the first can be refused while
  the fresh service is still settling on the bus and nothing retried.

Closes #168

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
mysticalsoap deleted branch bypass-reboot 2026-08-26 14:10:46 -04:00
Sign in to join this conversation.
No description provided.