fix: doublehop route hooks, hang on hop failure, cli reporting, leftover devices #201

Merged
mysticalsoap merged 5 commits from fix/doublehop-179 into trunk 2026-08-29 01:14:55 -04:00
Owner

Problem

Doublehop is broken end to end on any current system, and fails by hanging instead of reporting (#179):

  1. hop.sh/hop_down.sh call /sbin/route (net-tools), which no longer exists — OpenVPN treats the failed --up as fatal and the leg dies on every attempt, GUI and cli alike.
  2. A dead hop leg never touches connect_status, so the TunnelThread spins at 1 Hz forever inside the root service; only a service restart clears it.
  3. The cli only knows the bare signals: _hop variants pass silently, and the hop bookkeeping swallowed the main tunnel's bare connection_established, so -v hung on success too.
  4. (From the follow-up comment) the failed leg's DCO device survives the process and service restarts, holding a route that blackholes later single-hop connections by subnet collision.

Fix

One commit per layer:

  • hop.sh/hop_down.sh rewritten with idempotent ip route replace/del; the hop-server pin is proto 111-tagged so the #78 startup sweep clears it after a crash, and set -e makes a half-routed leg impossible. The _gateway literal in hop_down.sh is gone (a proto-qualified del needs no gateway).
  • The hop leg's teardown flags a pre-CONNECTED death (connect_status = -1); the wait turns that into a raise, and report_failure already drops both legs' firewall rules and reports the failed attempt. A hop stuck retrying deliberately keeps the wait: each retry emits conn_attempt_failed_hop, a client kills the leg, and that lands in the same abort path.
  • The cli handles the suffixed signals (progress line on hop connect, kill-and-quit naming the leg on either failure) and drops the stale hop bookkeeping. The Qt app object moves into main() so the module imports cleanly under pytest.
  • tunnel.remove_tun_device(): each leg deletes its own device after process exit (exact TUN_DEVICES names only), and service start sweeps all three next to the #78 route sweep.

Verification

  • 500 tests pass (10 new: hop-death abort, cli signal dispatch, device sweep), ruff E9/F63/F7/F82 clean, bash -n on both scripts.
  • Live run pending (merge gate): build via packaging/arch/build-branch-and-install.sh from this branch, then aqomui-cli -v <hop> -c <server> on ProtonVPN (plan allows two concurrent connections). Worth checking while at it: failure path (bogus hop server → clean abort, no stranded thread), ip route shows the proto-111 pin gone after disconnect, and no tun_aqomui_h left behind — plus the #88 speed numbers.

Closes #179

## Problem Doublehop is broken end to end on any current system, and fails by hanging instead of reporting (#179): 1. `hop.sh`/`hop_down.sh` call `/sbin/route` (net-tools), which no longer exists — OpenVPN treats the failed `--up` as fatal and the leg dies on every attempt, GUI and cli alike. 2. A dead hop leg never touches `connect_status`, so the TunnelThread spins at 1 Hz forever inside the root service; only a service restart clears it. 3. The cli only knows the bare signals: `_hop` variants pass silently, and the hop bookkeeping swallowed the main tunnel's bare `connection_established`, so `-v` hung on success too. 4. (From the follow-up comment) the failed leg's DCO device survives the process and service restarts, holding a route that blackholes later single-hop connections by subnet collision. ## Fix One commit per layer: - `hop.sh`/`hop_down.sh` rewritten with idempotent `ip route replace`/`del`; the hop-server pin is `proto 111`-tagged so the #78 startup sweep clears it after a crash, and `set -e` makes a half-routed leg impossible. The `_gateway` literal in `hop_down.sh` is gone (a proto-qualified del needs no gateway). - The hop leg's teardown flags a pre-CONNECTED death (`connect_status = -1`); the wait turns that into a raise, and `report_failure` already drops both legs' firewall rules and reports the failed attempt. A hop stuck retrying deliberately keeps the wait: each retry emits `conn_attempt_failed_hop`, a client kills the leg, and that lands in the same abort path. - The cli handles the suffixed signals (progress line on hop connect, kill-and-quit naming the leg on either failure) and drops the stale hop bookkeeping. The Qt app object moves into `main()` so the module imports cleanly under pytest. - `tunnel.remove_tun_device()`: each leg deletes its own device after process exit (exact `TUN_DEVICES` names only), and service start sweeps all three next to the #78 route sweep. ## Verification - 500 tests pass (10 new: hop-death abort, cli signal dispatch, device sweep), ruff E9/F63/F7/F82 clean, `bash -n` on both scripts. - **Live run pending (merge gate):** build via `packaging/arch/build-branch-and-install.sh` from this branch, then `aqomui-cli -v <hop> -c <server>` on ProtonVPN (plan allows two concurrent connections). Worth checking while at it: failure path (bogus hop server → clean abort, no stranded thread), `ip route` shows the proto-111 pin gone after disconnect, and no `tun_aqomui_h` left behind — plus the #88 speed numbers. Closes #179
hop.sh and hop_down.sh called /sbin/route (net-tools), which no longer
exists on current systems, so OpenVPN's --up hook died and took the leg
with it -- doublehop has been broken on Arch for the fork's whole life
(#179). The route adds become idempotent ip route replace: a device
route into the tun needs no gateway, and the hop-server pin is tagged
proto 111 like every other pin the service may have to sweep after a
crash. hop_down.sh also passed the literal string _gateway as a
gateway; the proto-qualified del needs none.

set -e replaces silent partial failure: a leg with half its routes is
worse than a dead one, and a dead one now aborts cleanly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
connect_status was only ever written by tunnel_up, so a hop leg that
died left work()'s wait spinning at 1 Hz forever inside the root
service: the main leg never started, no terminal signal for role main
was ever emitted, and only a service restart cleared the thread (#179).

The hop leg's teardown now flags a death that happened before CONNECTED
(-1), and the wait turns that into a raise -- report_failure already
drops both legs' firewall rules and reports the failed attempt. A hop
stuck retrying still holds the wait, deliberately: the attempt is
genuinely in progress, every retry emits conn_attempt_failed_hop, and a
client reacting to it kills the leg, which lands in the same abort path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
openvpn_log_monitor only knew the bare signals, from before the roles
carried suffixes: conn_attempt_failed_hop passed silently, and the
hop_active bookkeeping -- which expected the hop to also announce
itself bare -- swallowed the main tunnel's connection_established
instead. So with -v the cli hung forever on success and failure alike,
even once the service reports both properly (#179).

A bare signal is always the main tunnel's now, so the hop bookkeeping
goes; the _hop variants get their own branches: progress line on the
hop connecting, kill-and-quit on either leg failing, naming the leg.
The Qt application object moves into main() so the module can be
imported (and tested) without claiming the process's one Qt app slot.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
add: sweep leftover tunnel devices at teardown and service start
All checks were successful
ci / test (pull_request) Successful in 28s
272721d313
A DCO device can outlive its OpenVPN -- and a service restart -- sitting
DOWN while its kernel route still claims the tunnel subnet. Hit live on
#179: the failed hop leg's tun_aqomui_h held 10.96.0.0/16, so every
later single-hop connection came up healthy while its gateway resolved
through the dead device. Connected-but-no-traffic, hours later, with
nothing in any log.

Each leg now deletes its own device after the process exits (exact
TUN_DEVICES names only -- a raw config's concrete device may be a
persistent tun somebody else owns), and the service start sweeps all
three alongside the #78 route sweep, covering legs whose thread never
ran teardown.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Author
Owner

Fifth commit: an exception inside the hop leg's own bring-up (no process, so no EOF) could still arm the hang — the leg now runs behind a hop_leg wrapper that logs the error and releases the wait. Closes the #17 unguarded-helper-thread gap for the hop. 502 tests pass.

Fifth commit: an exception inside the hop leg's own bring-up (no process, so no EOF) could still arm the hang — the leg now runs behind a hop_leg wrapper that logs the error and releases the wait. Closes the #17 unguarded-helper-thread gap for the hop. 502 tests pass.
fix: release the doublehop wait when the hop leg raises
All checks were successful
ci / test (pull_request) Successful in 24s
ci / test (push) Successful in 24s
4ce0c024db
The EOF-based abort only fires once a process existed to die. An
exception in the hop leg's own bring-up (Popen failing, an unreadable
config) escaped its plain thread with no process, no EOF, and the wait
still armed -- the #179 hang through a narrower door. The leg now runs
behind a wrapper that logs the error and releases the wait; the known
unguarded-helper-threads gap (#17) closes for the hop.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Author
Owner

Live verification complete (2026-08-29):

  • Success path: -v hop -c exit brings both legs up; cli prints the hop progress line and exits after the main tunnel connects. hop.sh -s observed at connect, hop_down.sh at teardown.
  • Egress: exit server's IP with and without the hop (as it should be — the hop is invisible at egress; its reality shows in the +7 ms, see #88).
  • Teardown: verified after -t, after Ctrl-C, and after a failed attempt — no tun devices, no proto-111 pins, no openvpn processes, table 11 intact.
  • Sad path, hit organically: launching doublehop while a single-hop connection was still up lost the hop leg's PUSH reply (handshake raced the dying tunnel's routes). The new machinery did exactly its job: conn_attempt_failed_hop → cli killed the leg and reported it by name → service logged "connection attempt aborted - RuntimeError: the hop leg died before connecting" → clean state, no hang, no service restart.

One pre-existing quirk surfaced: switching from an active connection into doublehop races the old tunnel's teardown (the hop pin only lands at --up). A single-hop switch survives via OpenVPN's retry; the hop leg dies on first failure because clients kill on conn_attempt_failed_hop. Candidate follow-up issue, not this PR's scope. From a fresh disconnect the same pair connects reliably.

Live verification complete (2026-08-29): - **Success path**: -v hop -c exit brings both legs up; cli prints the hop progress line and exits after the main tunnel connects. hop.sh -s observed at connect, hop_down.sh at teardown. - **Egress**: exit server's IP with and without the hop (as it should be — the hop is invisible at egress; its reality shows in the +7 ms, see #88). - **Teardown**: verified after -t, after Ctrl-C, and after a failed attempt — no tun devices, no proto-111 pins, no openvpn processes, table 11 intact. - **Sad path, hit organically**: launching doublehop while a single-hop connection was still up lost the hop leg's PUSH reply (handshake raced the dying tunnel's routes). The new machinery did exactly its job: conn_attempt_failed_hop → cli killed the leg and reported it by name → service logged "connection attempt aborted - RuntimeError: the hop leg died before connecting" → clean state, no hang, no service restart. One pre-existing quirk surfaced: switching from an active connection into doublehop races the old tunnel's teardown (the hop pin only lands at --up). A single-hop switch survives via OpenVPN's retry; the hop leg dies on first failure because clients kill on conn_attempt_failed_hop. Candidate follow-up issue, not this PR's scope. From a fresh disconnect the same pair connects reliably.
mysticalsoap deleted branch fix/doublehop-179 2026-08-29 01:14:55 -04:00
Sign in to join this conversation.
No description provided.