Skip to content

flaky: buffer_ring_exhaustion_recovers times out on GitHub runners (echo never arrives for 1-2 of 8 connections) #378

Description

@brayniac

buffer_ring_exhaustion_recovers (ringline/tests/echo.rs, io_uring only) fails intermittently on GitHub-hosted runners with one or two of its eight client connections reading nothing:

assertion `left == right` failed: conn 4 echo mismatch
  left: []
 right: [99, 111, 110, 110, 101, 99, 116, 105, 111, 110, 45, 52, ...]

The echo binary finishes in ~5.7 s, i.e. the client's 5 s read timeout expired (set_read_timeout surfaces as WouldBlock, which the read loop treats as end of data). The server did not close the connection; the echo simply never arrived within 5 s.

Occurrences

Not reproducible on our VMs

Debian 13 (kernel 6.12) anvil guests via systemslab, cargo test --release --test echo buffer_ring_exhaustion_recovers:

host vCPUs branch single-test runs full echo binary runs
delta 1 main @ 00c69bb 10/10 10/10
delta 1 #377 branch 10/10 10/10
hv01 4 main @ 00c69bb 20/20 5/5
hv01 4 #377 branch 20/20 5/5

Runner image was ubuntu-24.04 version 20260907.300.1 for both the passing (#376) and failing (#377) release runs, so the image is not the variable either.

Test shape

recv_buffer(4, 4096) and 8 concurrent connections, so the provided ring is exhausted by design and the multishot recvs terminate with ENOBUFS, park on recv_starved, and are re-armed by flush_replenish_and_rearm once bids are replenished. A connection that stays parked because nothing replenishes for it would show exactly this symptom. Candidates worth ruling out: a parked connection that is only re-armed "if replenished" in the same pass (a replenish that lands in a later iteration with no starved-connection wake), and the interaction with the runner's kernel scheduling under 4 vCPUs.

Suggested next steps

  1. Make the test's failure informative: on timeout, dump the server-side recv_starved, recv_park_count, BUFFER_RING_EMPTY/RECV_PARKED metrics and the provided-ring free count via a test hook, so the next CI failure says whether the connection was still parked.
  2. Run the test in a loop on a runner-like environment (a 4-vCPU guest with the Ubuntu 24.04 kernel; anvil supports swapping the guest kernel) to get a local reproduction.
  3. Until then, treat a solo failure of this test as a rerun candidate, like signal_wait_on_signal_shutdown.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions