Summary
tls_segmented_recv_reassembles_and_eofs hangs a large fraction of runs on
io_uring. The server stops echoing part-way through a 100 KiB TLS message and
never resumes; the test then waits forever.
It is long-standing — reproduced on every sha tested, including before
#415 — and it is not a recent regression. It was invisible until now because
the test could not fail, only hang (see "Why this went unnoticed").
Repro
cargo test -p ringline --test tls_echo -- --test-threads=1
--test-threads=1 is the amplifier: ~40% there against ~6% in the default
parallel mode. Serialising increases the rate, so this is not
TEST_SERIALIZE contention.
Measured on an anvil guest (debian-13-ci, io_uring, 24 vCPU), 25 reps per sha:
| sha |
|
hangs |
0a8aabf |
#422 |
7/25 |
5afa2aa |
#421 |
15/25 |
b0e0605 |
#420 |
7/25 |
a91bde8 |
pre-#415 |
4/25 |
Rates vary run to run; presence does not.
What hangs
Single-threaded, the trailing test NAME ... line with no result names it
every time:
test tls_echo_with_external_client ... ok
test tls_info_accessors ... ok
test tls_outbound_connect_and_echo ... ok
test tls_segmented_recv_reassembles_and_eofs ... <-- never completes
The test writes 100 KiB over TLS and reads the echo back. The read stalls
part-way: the client keeps getting WouldBlock, i.e. the server has stopped
sending. The EOF wait later in the same test is already bounded (300 x 10 ms),
so the hang is in the echo read, not the close path.
100 KiB is deliberately larger than rustls' internal plaintext buffer, so the
server's drain loop produces several owned segments across several TLS records
— that is what the test is for. Possibly related to the recorded pre-existing
">16 KiB TLS echo" hang, though that one was mio and this is io_uring.
Why this went unnoticed
TcpStream::set_read_timeout surfaces a timed-out read as
ErrorKind::WouldBlock, which is indistinguishable from "no data yet". The
test's WouldBlock arm retried unconditionally, making the 5-second timeout
inert — so a stall became an unbounded hang rather than a failure, and CI saw
a 1200-second job timeout with no failing test name.
The same pattern is suite-wide: 46 WouldBlock arms across 8 test files,
many under a set_read_timeout. tls_echo.rs alone has 9 set_read_timeout
calls and 10 such arms; echo.rs has 59 and 27.
- Fixed for the two named tests in the PR that files this issue.
- The remaining sites are worth converting to a shared deadline-aware helper,
so no test in this suite can convert a stall into a hang again.
This issue tracks the underlying stall, not the test-harness defect.
Next steps
- Capture thread stacks from a stuck run (
gdb -p <tls_echo pid> -batch -ex "thread apply all bt"); gdb is not in the debian-13-ci guest image, so
install it in the job or use a different image.
- Determine whether the worker is parked on a CQE that never arrives, or
whether rustls has buffered plaintext that the drain loop is not re-polling
for.
Summary
tls_segmented_recv_reassembles_and_eofshangs a large fraction of runs onio_uring. The server stops echoing part-way through a 100 KiB TLS message and
never resumes; the test then waits forever.
It is long-standing — reproduced on every sha tested, including before
#415 — and it is not a recent regression. It was invisible until now because
the test could not fail, only hang (see "Why this went unnoticed").
Repro
--test-threads=1is the amplifier: ~40% there against ~6% in the defaultparallel mode. Serialising increases the rate, so this is not
TEST_SERIALIZEcontention.Measured on an anvil guest (debian-13-ci, io_uring, 24 vCPU), 25 reps per sha:
0a8aabf5afa2aab0e0605a91bde8Rates vary run to run; presence does not.
What hangs
Single-threaded, the trailing
test NAME ...line with no result names itevery time:
The test writes 100 KiB over TLS and reads the echo back. The read stalls
part-way: the client keeps getting
WouldBlock, i.e. the server has stoppedsending. The EOF wait later in the same test is already bounded (300 x 10 ms),
so the hang is in the echo read, not the close path.
100 KiB is deliberately larger than rustls' internal plaintext buffer, so the
server's drain loop produces several owned segments across several TLS records
— that is what the test is for. Possibly related to the recorded pre-existing
">16 KiB TLS echo" hang, though that one was mio and this is io_uring.
Why this went unnoticed
TcpStream::set_read_timeoutsurfaces a timed-out read asErrorKind::WouldBlock, which is indistinguishable from "no data yet". Thetest's
WouldBlockarm retried unconditionally, making the 5-second timeoutinert — so a stall became an unbounded hang rather than a failure, and CI saw
a 1200-second job timeout with no failing test name.
The same pattern is suite-wide: 46
WouldBlockarms across 8 test files,many under a
set_read_timeout.tls_echo.rsalone has 9set_read_timeoutcalls and 10 such arms;
echo.rshas 59 and 27.so no test in this suite can convert a stall into a hang again.
This issue tracks the underlying stall, not the test-harness defect.
Next steps
gdb -p <tls_echo pid> -batch -ex "thread apply all bt"); gdb is not in thedebian-13-ciguest image, soinstall it in the job or use a different image.
whether rustls has buffered plaintext that the drain loop is not re-polling
for.