doublehop: hop.sh calls /sbin/route and a failed hop leg hangs the service #179

Closed
opened 2026-08-27 18:02:26 -04:00 by mysticalsoap · 1 comment
Owner

Found live-testing #178: -c, -t, -l work, but aqomui-cli -v <hop> -c <server> hangs forever. The log shows the hop leg comes up (TLS handshake, ASSIGN_IP on tun_aqomui_h) and then dies running the up-script — three layers, all pre-dating #178 (the #132 fix itself worked: both projections built and the service launched the hop leg).

1. hop.sh / hop_down.sh call /sbin/route (net-tools), which no longer exists on Arch.

/usr/share/aqomui/scripts/hop.sh: line 8: /sbin/route: No such file or directory
/usr/share/aqomui/scripts/hop.sh: line 9: /sbin/route: No such file or directory
OpenVPN: WARNING: Failed running command (--up/--down): could not execute external program
Exiting due to fatal error

OpenVPN treats the failed --up as fatal and exits. This breaks doublehop for the gui and cli alike on any current system — nobody noticed because doublehop hasn't been live-run since the fork. The scripts already use ip route for reading; the fix is rewriting the route add/del calls as iproute2 (ip route add <ip>/32 via ...). hop_down.sh also passes a literal gw _gateway, which needs the same look while there. Both legs' --up hooks are affected (hop.sh -f on the first leg, hop.sh -s on the second), so the second leg would have died the same way had the first survived.

2. A failed hop leg hangs the service's TunnelThread forever.

tunnel.py run(), after starting the hop thread:

while self.connect_status == 0:
    time.sleep(1)

connect_status is only ever set in tunnel_up. A hop leg that dies emits conn_attempt_failed_hop but never touches it, so the thread spins at 1 Hz forever, the main leg never starts, and no terminal signal for role "main" is ever emitted. Needs an abort path: hop failure ends the wait, cleans up (including the firewall.allow_dest_ip(hop_ip, "-I") rule inserted just before), and reports the attempt failed. A restart of aqomui-service is currently the only way to clear the stuck thread.

3. The cli is deaf to the suffixed signals.

openvpn_log_monitor matches only bare connection_established / conn_attempt_failed; the _hop variants (which the gui handles) pass it silently. Even with (2) fixed, -v needs the suffix handling to report the hop leg's fate instead of hanging.

Fixing (1) makes doublehop work at all; (2) makes failure terminate instead of hang; (3) makes the cli report it. (2) and (3) are what turned a script error into a silent hang with zero feedback.

Found live-testing #178: `-c`, `-t`, `-l` work, but `aqomui-cli -v <hop> -c <server>` hangs forever. The log shows the hop leg comes up (TLS handshake, ASSIGN_IP on tun_aqomui_h) and then dies running the up-script — three layers, all pre-dating #178 (the #132 fix itself worked: both projections built and the service launched the hop leg). **1. hop.sh / hop_down.sh call `/sbin/route` (net-tools), which no longer exists on Arch.** ``` /usr/share/aqomui/scripts/hop.sh: line 8: /sbin/route: No such file or directory /usr/share/aqomui/scripts/hop.sh: line 9: /sbin/route: No such file or directory OpenVPN: WARNING: Failed running command (--up/--down): could not execute external program Exiting due to fatal error ``` OpenVPN treats the failed `--up` as fatal and exits. This breaks doublehop for the **gui and cli alike** on any current system — nobody noticed because doublehop hasn't been live-run since the fork. The scripts already use `ip route` for reading; the fix is rewriting the `route add/del` calls as iproute2 (`ip route add <ip>/32 via ...`). `hop_down.sh` also passes a literal `gw _gateway`, which needs the same look while there. Both legs' `--up` hooks are affected (`hop.sh -f` on the first leg, `hop.sh -s` on the second), so the second leg would have died the same way had the first survived. **2. A failed hop leg hangs the service's TunnelThread forever.** tunnel.py run(), after starting the hop thread: ```python while self.connect_status == 0: time.sleep(1) ``` `connect_status` is only ever set in `tunnel_up`. A hop leg that dies emits `conn_attempt_failed_hop` but never touches it, so the thread spins at 1 Hz forever, the main leg never starts, and no terminal signal for role "main" is ever emitted. Needs an abort path: hop failure ends the wait, cleans up (including the `firewall.allow_dest_ip(hop_ip, "-I")` rule inserted just before), and reports the attempt failed. A restart of aqomui-service is currently the only way to clear the stuck thread. **3. The cli is deaf to the suffixed signals.** `openvpn_log_monitor` matches only bare `connection_established` / `conn_attempt_failed`; the `_hop` variants (which the gui handles) pass it silently. Even with (2) fixed, `-v` needs the suffix handling to report the hop leg's fate instead of hanging. Fixing (1) makes doublehop work at all; (2) makes failure terminate instead of hang; (3) makes the cli report it. (2) and (3) are what turned a script error into a silent hang with zero feedback.
Author
Owner

Blast radius is worse than a stranded thread. Hit live 2026-08-27 during the PR #184 check: the failed hop leg's DCO device (tun_aqomui_h) survives both the OpenVPN death and a service restart — it sits state DOWN holding its kernel route (10.96.0.0/16 dev tun_aqomui_h linkdown). ProtonVPN assigns from that same 10.96.0.0/16, so every later single-hop connection comes up healthy (control channel rides the physical link) while its gateway resolves through the dead device — ip route get 10.96.0.1 → dev tun_aqomui_h. Connected-but-no-traffic, hours after the doublehop attempt, no error anywhere.

Recovery: ip link delete tun_aqomui_h (removes the route with the device).

Repair scope for this issue grows one item: on service start and on tunnel teardown, delete any leftover TUN_DEVICES interfaces — a dead DCO device is invisible in every log and poisons routing by subnet collision.

Blast radius is worse than a stranded thread. Hit live 2026-08-27 during the PR #184 check: the failed hop leg's DCO device (`tun_aqomui_h`) survives both the OpenVPN death and a service restart — it sits `state DOWN` holding its kernel route (`10.96.0.0/16 dev tun_aqomui_h linkdown`). ProtonVPN assigns from that same 10.96.0.0/16, so every later *single-hop* connection comes up healthy (control channel rides the physical link) while its gateway resolves through the dead device — `ip route get 10.96.0.1` → `dev tun_aqomui_h`. Connected-but-no-traffic, hours after the doublehop attempt, no error anywhere. Recovery: `ip link delete tun_aqomui_h` (removes the route with the device). Repair scope for this issue grows one item: on service start and on tunnel teardown, delete any leftover `TUN_DEVICES` interfaces — a dead DCO device is invisible in every log and poisons routing by subnet collision.
Sign in to join this conversation.
No milestone
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
mysticalsoap/aqomui#179
No description provided.