Skip to content

flaky (io_uring): tls_segmented_recv_reassembles_and_eofs hangs — server stops echoing mid 100 KiB TLS message #423

Description

@brayniac

Summary

tls_segmented_recv_reassembles_and_eofs hangs a large fraction of runs on
io_uring. The server stops echoing part-way through a 100 KiB TLS message and
never resumes; the test then waits forever.

It is long-standing — reproduced on every sha tested, including before
#415 — and it is not a recent regression. It was invisible until now because
the test could not fail, only hang (see "Why this went unnoticed").

Repro

cargo test -p ringline --test tls_echo -- --test-threads=1

--test-threads=1 is the amplifier: ~40% there against ~6% in the default
parallel mode. Serialising increases the rate, so this is not
TEST_SERIALIZE contention.

Measured on an anvil guest (debian-13-ci, io_uring, 24 vCPU), 25 reps per sha:

sha hangs
0a8aabf #422 7/25
5afa2aa #421 15/25
b0e0605 #420 7/25
a91bde8 pre-#415 4/25

Rates vary run to run; presence does not.

What hangs

Single-threaded, the trailing test NAME ... line with no result names it
every time:

test tls_echo_with_external_client ... ok
test tls_info_accessors ... ok
test tls_outbound_connect_and_echo ... ok
test tls_segmented_recv_reassembles_and_eofs ...     <-- never completes

The test writes 100 KiB over TLS and reads the echo back. The read stalls
part-way: the client keeps getting WouldBlock, i.e. the server has stopped
sending. The EOF wait later in the same test is already bounded (300 x 10 ms),
so the hang is in the echo read, not the close path.

100 KiB is deliberately larger than rustls' internal plaintext buffer, so the
server's drain loop produces several owned segments across several TLS records
— that is what the test is for. Possibly related to the recorded pre-existing
">16 KiB TLS echo" hang, though that one was mio and this is io_uring.

Why this went unnoticed

TcpStream::set_read_timeout surfaces a timed-out read as
ErrorKind::WouldBlock, which is indistinguishable from "no data yet". The
test's WouldBlock arm retried unconditionally, making the 5-second timeout
inert — so a stall became an unbounded hang rather than a failure, and CI saw
a 1200-second job timeout with no failing test name.

The same pattern is suite-wide: 46 WouldBlock arms across 8 test files,
many under a set_read_timeout. tls_echo.rs alone has 9 set_read_timeout
calls and 10 such arms; echo.rs has 59 and 27.

  • Fixed for the two named tests in the PR that files this issue.
  • The remaining sites are worth converting to a shared deadline-aware helper,
    so no test in this suite can convert a stall into a hang again.

This issue tracks the underlying stall, not the test-harness defect.

Next steps

  • Capture thread stacks from a stuck run (gdb -p <tls_echo pid> -batch -ex "thread apply all bt"); gdb is not in the debian-13-ci guest image, so
    install it in the job or use a different image.
  • Determine whether the worker is parked on a CQE that never arrives, or
    whether rustls has buffered plaintext that the drain loop is not re-polling
    for.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions