change: route the double hop from Python, not --up hooks #211

Merged
mysticalsoap merged 3 commits from change/doublehop-routing-176 into trunk 2026-08-30 12:06:34 -04:00
Owner

Stage 1 of #176: the hop routing model moves out of shell hooks into the tunnel thread, becoming protocol-neutral - the WireGuard stages build on this.

Problem: The double hop's routing lived in hop.sh/hop_down.sh, OpenVPN --up/--down hooks. Two consequences: the model was OpenVPN-only (WireGuard hop legs have no place to plug in, #176), and the hop server's pin was installed after the hop's handshake - up-hooks run post-handshake - so the handshake it was meant to steer could still ride a dying tunnel's def1 routes when switching from an active connection (#202).

Change:

  • Both legs run --route-noexec; tunnel_up installs each leg's routes at CONNECTED. The hop steers the main server into its device; the main leg takes the def1 pair - the same routing model wireguard.py already uses.
  • The hop server's pin (pin_server_route, shared with the bypass path now) is installed before the hop leg launches, and is proto-111-tagged - a killed service's startup sweep already flushes those, so crash cleanup comes free.
  • A leg whose route install fails aborts the attempt through the same #179 paths (suffixed failure signal, wait release).
  • hop.sh/hop_down.sh deleted, setup.py follows.

Verification: 10 new tests (pin-precedes-launch, per-leg route install, route-failure aborts, teardown/failure unpin, leg command lines); full suite 520 passed. Needs live double-hop runs before merge: fresh connect, switch from an active connection (the #202 scenario), and failure teardown.

Part of #176. Closes #202.

🤖 Generated with Claude Code

Stage 1 of #176: the hop routing model moves out of shell hooks into the tunnel thread, becoming protocol-neutral - the WireGuard stages build on this. **Problem**: The double hop's routing lived in `hop.sh`/`hop_down.sh`, OpenVPN `--up`/`--down` hooks. Two consequences: the model was OpenVPN-only (WireGuard hop legs have no place to plug in, #176), and the hop server's pin was installed *after* the hop's handshake - up-hooks run post-handshake - so the handshake it was meant to steer could still ride a dying tunnel's def1 routes when switching from an active connection (#202). **Change**: - Both legs run `--route-noexec`; `tunnel_up` installs each leg's routes at CONNECTED. The hop steers the main server into its device; the main leg takes the def1 pair - the same routing model `wireguard.py` already uses. - The hop server's pin (`pin_server_route`, shared with the bypass path now) is installed **before** the hop leg launches, and is proto-111-tagged - a killed service's startup sweep already flushes those, so crash cleanup comes free. - A leg whose route install fails aborts the attempt through the same #179 paths (suffixed failure signal, wait release). - `hop.sh`/`hop_down.sh` deleted, `setup.py` follows. **Verification**: 10 new tests (pin-precedes-launch, per-leg route install, route-failure aborts, teardown/failure unpin, leg command lines); full suite 520 passed. Needs live double-hop runs before merge: fresh connect, switch from an active connection (the #202 scenario), and failure teardown. Part of #176. Closes #202. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
change: route the double hop from Python, not --up hooks
All checks were successful
ci / test (pull_request) Successful in 30s
bd70cb4d11
Both legs now run --route-noexec and tunnel_up installs each leg's
routes at CONNECTED: the hop steers the main server into its device,
the main leg takes the def1 pair - the model wireguard.py already
routes with, shared so a WireGuard leg can join it (#176).

The hop server's pin moves ahead of the launch. hop.sh installed it
from --up, which runs only after the handshake - so the handshake it
was meant to steer could still ride a dying tunnel's def1 routes when
switching from an active connection (#202). It is proto-tagged now,
which also puts it under the startup sweep a killed service relies on.

hop.sh and hop_down.sh die; the bypass pin shares the new helper.

Closes #202

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
add: log the double hop's route installs
All checks were successful
ci / test (pull_request) Successful in 27s
6ade813e45
The shell hooks announced themselves through OpenVPN's script exec
lines; their Python replacements were silent, leaving a working run
and a routing no-op indistinguishable in the log.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
fix: re-announce the device at CONNECTED
All checks were successful
ci / test (pull_request) Successful in 24s
ci / test (push) Successful in 30s
98095e657e
On a switch the dying thread's teardown (dev=None) lands on the
service's shared role state about a second after the new thread's
pre-launch announcement - the exit notification takes that long. The
gui reads that state at CONNECTED to start its tunnel monitor, got
interface None, and the monitor died on its first counter read: uptime
frozen at one second, no traffic graph, no external-ip check.

The race predates the doublehop rework; quick switch testing surfaced
it. Re-announcing at CONNECTED rides the same signal queue as the
status the gui reacts to, so the read is ordered behind fresh state.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Author
Owner

Two commits added after the first live round (log-verified: fresh doublehop, switch-while-connected, and proto-111 teardown all behaved; the switch handshake completed 50ms in while the old tunnel was still dying - the #202 path working):

  • 6ade813 - the Python route installs now log themselves ("Pinned ... via ...", "Routed ... into ..."); the shell hooks used to announce themselves through OpenVPN's script exec lines and their replacements were silent.
  • 98095e6 - fixes the frozen-uptime observation from that round: on a switch, the dying thread's dev=None teardown landed on the shared role state after the new thread's pre-launch announcement, so the gui started its monitor on interface None and it died one tick in. Pre-existing race, not a regression from this PR - quick switching just surfaced it. CONNECTED now re-announces the device ahead of the status signal.

Remaining before merge: one more run checking egress (curl should show the main server's exit, not the hop's) and that uptime now counts across a switch.

Two commits added after the first live round (log-verified: fresh doublehop, switch-while-connected, and proto-111 teardown all behaved; the switch handshake completed 50ms in while the old tunnel was still dying - the #202 path working): - `6ade813` - the Python route installs now log themselves ("Pinned ... via ...", "Routed ... into ..."); the shell hooks used to announce themselves through OpenVPN's script exec lines and their replacements were silent. - `98095e6` - fixes the frozen-uptime observation from that round: on a switch, the dying thread's `dev=None` teardown landed on the shared role state after the new thread's pre-launch announcement, so the gui started its monitor on interface None and it died one tick in. Pre-existing race, not a regression from this PR - quick switching just surfaced it. CONNECTED now re-announces the device ahead of the status signal. Remaining before merge: one more run checking egress (curl should show the main server's exit, not the hop's) and that uptime now counts across a switch.
Author
Owner

Live verification complete (2026-08-29):

  • Fresh doublehop connect and switch-while-connected both clean; the switch's hop handshake completed while the old tunnel was still tearing down - the #202 race is gone.
  • Egress checked: exit IP is the main server's and changes as expected.
  • Uptime counts across a switch after 98095e6.
  • ip route show proto 111 empty after disconnect.

Ready to merge.

Live verification complete (2026-08-29): - Fresh doublehop connect and switch-while-connected both clean; the switch's hop handshake completed while the old tunnel was still tearing down - the #202 race is gone. - Egress checked: exit IP is the main server's and changes as expected. - Uptime counts across a switch after 98095e6. - `ip route show proto 111` empty after disconnect. Ready to merge.
mysticalsoap deleted branch change/doublehop-routing-176 2026-08-30 12:06:34 -04:00
Sign in to join this conversation.
No description provided.