Skip to content

flaky (CI): Test (release) aborts with io_uring_setup ENOMEM on GitHub runners #426

Description

@brayniac

Symptom

A test binary aborts before running any test:

ringline: cannot create the test event loop on this host, aborting the test binary so this is reported once:
  ring setup: io_uring_setup(2): Cannot allocate memory (os error 12) (ENOMEM). Fix the underlying
  error, or build with the `force-mio` cargo feature to use the mio backend instead.
error: test failed, to rerun pass `-p ringline --lib`

Every test up to that point passes; the run dies at the point a new test event loop is created. Pre-existing on main and on PRs alike, most often on Test (release).

Ruled out

RLIMIT_MEMLOCK — measured and eliminated. The original version of this issue led with it ("characteristically RLIMIT_MEMLOCK exhaustion"), and that is wrong. The whole uring test suite passes with the soft limit at 8 KiB on kernel 6.12, so the ring is not charged against memlock there. This was recorded only in a comment in .github/workflows/ci.yml, not here, and as a result it has been re-proposed as the cause since. It is not the cause.

Host memory pressure — measured and eliminated. This was the surviving hypothesis, and the reason #474 added a resource-capture step. Those diagnostics have now fired. Run 36534300805, at the moment of the abort:

max locked memory  (kbytes, -l) 8192
Mem:  15988 total, 1041 used, 11511 free, 14947 available
Swap:  3071 total,     0 used

11.5 GiB free, 14.9 GiB available, zero swap used, memlock 8 MiB. Nothing was short of host memory, and no OOM kill appeared in dmesg.

Still open, and now instrumented (#505)

free -m reads host-wide /proc/meminfo, so it is structurally blind to the two remaining candidates. #505 captures both on failure:

  • A cgroup memory limit. io_uring charges ring memory to the memcg on modern kernels, so a cgroup limit produces precisely this signature — ENOMEM at io_uring_setup while the host reports gigabytes free. Now capturing memory.max / high / current / peak / events for the root and for the job's own cgroup, with a v1 fallback. memory.events and memory.peak are the load-bearing fields: they are retrospective and survive the death of the process that failed, which memory.current does not.
  • vm.max_map_count. io_uring_setup mmaps its rings, and the ceiling appears in neither ulimit -a nor free. ci: capture cgroup memory and the map-count ceiling on failure (#426) #505 records the limit only — the failing process is already gone when the step runs, so its own map count cannot be recovered post-hoc. If the cgroup numbers come back clean, the next step is capturing the map count in-process on the abort path.

/proc/pressure/memory (PSI) is also captured now, since it shows memory stall without an OOM kill.

Note: Test (release) is red for two independent reasons

Of three Test (release) failures on 2026-09-29, one was this ENOMEM, one was buffer_ring_exhaustion_recovers (#378 / #349), and one was both. Fixing this issue alone will not make that job trustworthy.

Why it matters beyond the noise

The abort is indistinguishable from a real failure at a glance — it exits 101 with error: test failed. A reviewer seeing a red Test (release) has to open the log to discover nothing actually failed, and the reflex becomes "rerun it", which is the same reflex that let #423 hide for months behind an opaque timeout.

The runtime's message here is good (it names the cause and the workaround, from #361). The gap is that CI treats an infrastructure failure as a test failure.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions