Skip to content

flaky: free_port() probe-then-rebind races, surfacing as AddrInUse launch failures #431

Description

@brayniac

The race

free_port() (in ringline/tests/echo.rs, duplicated across the other test
binaries) picks a port by binding a throwaway listener to :0, reading the
port, then dropping the listener so the server under test can bind it:

let listener = std::net::TcpListener::bind("127.0.0.1:0").unwrap();
let port = listener.local_addr().unwrap().port();
drop(listener);          // <-- window opens here

Between that drop and the server's own bind, anything on the machine can
take the port. The helper's CLAIMED set only dedupes within one test
binary, and cargo test --all runs several binaries in parallel, so the
kernel can hand the same ephemeral port to two of them.

It surfaces as:

launch failed: Io(Os { code: 98, kind: AddrInUse, message: "Address already in use" })

...attributed to whichever test lost the race, which is why it looks like a
different test each time.

Evidence

Hit while adding one more server-launching test in #432. 25x full-suite A/B on
a single exclusive anvil guest (debian-13-ci, kernel 6.12):

arm full-suite failures
base 0/25
base + one more free_port() launch 1/25

Low rate, but it scales with the number of concurrently-launched servers, so it
gets worse as the suite grows. It is also load-sensitive, which is consistent
with the GitHub-runner flakes in #378/#349 (different symptom, same "more tests
than cores" pressure).

The fix

Let the server own the bind and read the resolved address back — the runtime
already supports it, and shutdown_handle_reports_bound_addr already does it:

let (shutdown, handles) = RinglineBuilder::new(test_config())
    .bind("127.0.0.1:0".parse().unwrap())
    .launch::<H>()
    .expect("launch failed");
let addr = shutdown.bound_addr().expect("bound_addr after a TCP bind").to_string();

There is no window at all: the listening socket is never closed and reopened.

split_halves_echo_round_trip (#432) is converted to this pattern already.
The remaining ~131 free_port() call sites across echo.rs (104),
tls_echo.rs (14), stream_echo.rs (8), throughput.rs (3) and
signal_shutdown.rs (2) are a mechanical sweep.

Not all of them can convert blindly — a few need the address before launch
(e.g. handlers that stash a backend addr in a OnceLock), and those genuinely
need a reserved port. Those can keep free_port(), ideally holding the
listener until the server is up rather than dropping it early.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions