Skip to content

flaky (quic): large_payload_round_trip_via_endpoint_api — two reasoned fixes have missed #393

Description

@brayniac

Split out of #386, which was closed with this item explicitly unverified.

Symptom

ringline-quic/tests/peer.rs large_payload_round_trip_via_endpoint_api fails server should observe FIN under full-workspace parallel load (cargo test --all). It passes in isolation — 40/40 running the quic crate alone — so isolation cannot characterise it.

Two fixes attempted, both reasoned rather than measured, both missed

#389 replaced read_until_fin's fixed 64-round budget with a progress-based loop. It reproduced with that fix in place.

#392 diagnosed why: #389 defined progress as "application bytes arrived", but the stream FIN routinely needs several exchanges after the last data byte. Those rounds grow acc by nothing, so the budget drained in exactly the state the loop was waiting on. drain now returns whether any datagram crossed and that counts as progress.

That is hypothesis two, and it is not proven: 39/40 full-workspace runs passed with it, against no matched baseline at the same condition. It may have helped, done nothing, or changed the rate slightly.

What #392 does land

A diagnostic. When the loop gives up it now prints acc, expected, rounds and stall_budget, so the next failure says which bound fired:

  • stall_budget == 0 — both endpoints went quiet, and the FIN genuinely never came.
  • rounds at the cap — packets kept moving without the FIN ever arriving, which would mean my progress signal is now too permissive and the loop spins on timer-driven traffic.

These are different bugs wanting different fixes, and two attempts have failed to distinguish them from the bare assertion.

Why it stopped here

This is a pre-existing test-harness flake, not a product defect, and it was absorbing time from the release comparison campaign (#391). The next person to see it should read the diagnostic line rather than reason about mechanisms, as I did twice.

Reproduction notes

  • Needs cargo test --all; the quic crate alone will not do it.
  • Observed rate roughly 1 in 15-40 full-workspace runs on macOS, so a meaningful before/after needs >=100 runs per arm — an n=30 comparison on a sibling flake in flaky (mio): async_join3_mixed, signal_wait_on_signal_shutdown, and the quic peer FIN assert #386 produced a convincing result that reversed at n=120.
  • The only nondeterminism in the path is drain reading Instant::now() and driving QUIC's timers from it; the transport itself is an in-memory function call.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions