doublehop: hop.sh calls /sbin/route and a failed hop leg hangs the service #179
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Found live-testing #178:
-c,-t,-lwork, butaqomui-cli -v <hop> -c <server>hangs forever. The log shows the hop leg comes up (TLS handshake, ASSIGN_IP on tun_aqomui_h) and then dies running the up-script — three layers, all pre-dating #178 (the #132 fix itself worked: both projections built and the service launched the hop leg).1. hop.sh / hop_down.sh call
/sbin/route(net-tools), which no longer exists on Arch.OpenVPN treats the failed
--upas fatal and exits. This breaks doublehop for the gui and cli alike on any current system — nobody noticed because doublehop hasn't been live-run since the fork. The scripts already useip routefor reading; the fix is rewriting theroute add/delcalls as iproute2 (ip route add <ip>/32 via ...).hop_down.shalso passes a literalgw _gateway, which needs the same look while there. Both legs'--uphooks are affected (hop.sh -fon the first leg,hop.sh -son the second), so the second leg would have died the same way had the first survived.2. A failed hop leg hangs the service's TunnelThread forever.
tunnel.py run(), after starting the hop thread:
connect_statusis only ever set intunnel_up. A hop leg that dies emitsconn_attempt_failed_hopbut never touches it, so the thread spins at 1 Hz forever, the main leg never starts, and no terminal signal for role "main" is ever emitted. Needs an abort path: hop failure ends the wait, cleans up (including thefirewall.allow_dest_ip(hop_ip, "-I")rule inserted just before), and reports the attempt failed. A restart of aqomui-service is currently the only way to clear the stuck thread.3. The cli is deaf to the suffixed signals.
openvpn_log_monitormatches only bareconnection_established/conn_attempt_failed; the_hopvariants (which the gui handles) pass it silently. Even with (2) fixed,-vneeds the suffix handling to report the hop leg's fate instead of hanging.Fixing (1) makes doublehop work at all; (2) makes failure terminate instead of hang; (3) makes the cli report it. (2) and (3) are what turned a script error into a silent hang with zero feedback.
Blast radius is worse than a stranded thread. Hit live 2026-08-27 during the PR #184 check: the failed hop leg's DCO device (
tun_aqomui_h) survives both the OpenVPN death and a service restart — it sitsstate DOWNholding its kernel route (10.96.0.0/16 dev tun_aqomui_h linkdown). ProtonVPN assigns from that same 10.96.0.0/16, so every later single-hop connection comes up healthy (control channel rides the physical link) while its gateway resolves through the dead device —ip route get 10.96.0.1→dev tun_aqomui_h. Connected-but-no-traffic, hours after the doublehop attempt, no error anywhere.Recovery:
ip link delete tun_aqomui_h(removes the route with the device).Repair scope for this issue grows one item: on service start and on tunnel teardown, delete any leftover
TUN_DEVICESinterfaces — a dead DCO device is invisible in every log and poisons routing by subnet collision.