doublehop: switching from an active connection races the old tunnel's teardown #202

Closed
opened 2026-08-29 00:48:35 -04:00 by mysticalsoap · 0 comments
Owner

Seen live 2026-08-29 while validating #201: with a single-hop connection up, launching a doublehop attempt fails after ~60s with no-push-reply, while the same pair connects reliably from a fresh disconnect.

The race: establish_connection kills the current main tunnel and immediately launches the hop leg, whose pin via the physical link only lands at --up. Its handshake traffic therefore rides the dying tunnel's def1 routes; when they vanish mid-handshake, the PUSH reply is lost. OpenVPN would retry — but both clients kill the attempt on the first conn_attempt_failed_hop (the gui always has; the cli since #201), so the attempt dies. A single-hop switch survives the same race because nothing kills it before the retry succeeds.

Fix shapes, in rough preference order:

  1. Pin the hop endpoint via the physical gateway before launch (tunnel.py already does exactly this pre-launch pin for the bypass role) — folds naturally into the #176 reshape, where all hop routing moves into Python anyway.
  2. Have the service delay the new attempt until the old tunnel's teardown completes.
  3. Make clients tolerate the first hop-leg retry instead of killing on it.

Probably best absorbed into #176 rather than fixed standalone.

Seen live 2026-08-29 while validating #201: with a single-hop connection up, launching a doublehop attempt fails after ~60s with no-push-reply, while the same pair connects reliably from a fresh disconnect. The race: establish_connection kills the current main tunnel and immediately launches the hop leg, whose pin via the physical link only lands at --up. Its handshake traffic therefore rides the dying tunnel's def1 routes; when they vanish mid-handshake, the PUSH reply is lost. OpenVPN would retry — but both clients kill the attempt on the first conn_attempt_failed_hop (the gui always has; the cli since #201), so the attempt dies. A single-hop switch survives the same race because nothing kills it before the retry succeeds. Fix shapes, in rough preference order: 1. Pin the hop endpoint via the physical gateway before launch (tunnel.py already does exactly this pre-launch pin for the bypass role) — folds naturally into the #176 reshape, where all hop routing moves into Python anyway. 2. Have the service delay the new attempt until the old tunnel's teardown completes. 3. Make clients tolerate the first hop-leg retry instead of killing on it. Probably best absorbed into #176 rather than fixed standalone.
Sign in to join this conversation.
No milestone
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
mysticalsoap/aqomui#202
No description provided.