Describe the bug
In 1.10.1, calling obtain_shm_provider(true) (the new blocking argument) before the session's implicit SHM provider has finished initializing returns a Ready provider, but leaves the session's provider state permanently Initializing:
- every later non-blocking
obtain_shm_provider(false) returns Initializing for the life of the session
- a second blocking call returns
Error
Calling the blocking variant after the provider is already Ready is fine. Only a blocking call made while the state is still Initializing triggers this.
The implicit provider is documented as lazily initialized and available once initialization completes, so a state that never leaves Initializing does not look like the intended contract.
To reproduce
Self-contained, no router needed (zenoh-cpp over zenoh-c):
repro.cpp
// Minimal reproducer: obtain_shm_provider(true) before the provider is Ready leaves the
// session's implicit SHM provider permanently "Initializing".
#include <chrono>
#include <cstdio>
#include <string>
#include <thread>
#include <zenoh.hxx>
static const char* state(const std::variant<zenoh::SharedShmProvider, zenoh::ShmProviderNotReadyState>& r) {
if (std::holds_alternative<zenoh::SharedShmProvider>(r)) return "Ready";
switch (std::get<zenoh::ShmProviderNotReadyState>(r)) {
case zenoh::ShmProviderNotReadyState::SHM_PROVIDER_DISABLED: return "Disabled";
case zenoh::ShmProviderNotReadyState::SHM_PROVIDER_INITIALIZING: return "Initializing";
case zenoh::ShmProviderNotReadyState::SHM_PROVIDER_ERROR: return "Error";
}
return "?";
}
static zenoh::Session open_session() {
auto c = zenoh::Config::create_default();
c.insert_json5("scouting/multicast/enabled", "false");
c.insert_json5("transport/shared_memory/enabled", "true");
c.insert_json5("transport/shared_memory/transport_optimization/enabled", "true");
return zenoh::Session::open(std::move(c));
}
int main(int argc, char** argv) {
const std::string order = argc > 1 ? argv[1] : "blocking-first";
auto s = open_session();
std::printf("order: %s\n", order.c_str());
if (order == "blocking-first") {
auto first = s.obtain_shm_provider(true);
std::printf(" blocking obtain -> %s\n", state(first));
int ready = 0, init = 0, err = 0;
for (int i = 0; i < 200; ++i) {
std::string st = state(s.obtain_shm_provider(false));
ready += st == "Ready"; init += st == "Initializing"; err += st == "Error";
std::this_thread::sleep_for(std::chrono::milliseconds(10));
}
std::printf(" non-blocking x200 over 2 s -> Ready %d, Initializing %d, Error %d\n", ready, init, err);
std::printf(" blocking obtain again -> %s\n", state(s.obtain_shm_provider(true)));
} else {
int polls = 0; std::string st;
do { st = state(s.obtain_shm_provider(false)); ++polls; if (st != "Ready") std::this_thread::sleep_for(std::chrono::milliseconds(1)); }
while (st != "Ready" && polls < 3000);
std::printf(" non-blocking polling -> %s after %d polls\n", st.c_str(), polls);
int ready = 0;
for (int i = 0; i < 200; ++i) ready += std::string(state(s.obtain_shm_provider(false))) == "Ready";
std::printf(" non-blocking x200 -> Ready %d / 200\n", ready);
std::printf(" blocking obtain after Ready -> %s\n", state(s.obtain_shm_provider(true)));
}
return 0;
}
g++ -std=c++17 -DZENOHCXX_ZENOHC -I<prefix>/include repro.cpp -L<prefix>/lib -lzenohc -o repro
./repro blocking-first
./repro nonblocking-first
Output (identical over 3 runs):
order: blocking-first
blocking obtain -> Ready
non-blocking x200 over 2 s -> Ready 0, Initializing 200, Error 0
blocking obtain again -> Error
order: nonblocking-first
non-blocking polling -> Ready after 6 polls
non-blocking x200 -> Ready 200 / 200
blocking obtain after Ready -> Ready
The nonblocking-first run is the control: the same session reaches Ready after a few polls and stays Ready, and a blocking call made after that also returns Ready.
Likely cause
LazyShmProvider::try_get_provider creates a flume::bounded(1) channel, the init task sends the provider on it once, and the state moves to Ready only when a non-blocking caller's try_recv() succeeds:
The blocking path in zenoh-c receives that single message through the cloned receiver and returns it to the caller, but nothing moves the shared state to Ready:
After that the channel is empty and the sender is dropped, so every later try_recv() hits the Err(_) arm (interop.rs#L148) and returns Initializing again. A second blocking call gets None from the drained channel and reports Error.
Impact
Observed. An application that uses the implicit provider when it is Ready and otherwise sends a regular (non-SHM) payload, and that did one early blocking obtain right after opening the session, never saw the provider become Ready through its non-blocking path. Measured with a 6 MiB payload at 30 Hz for 12 s, 3 runs: 360 / 360 frames were received as non-SHM in every run (vs. 1 cold frame and the rest SHM without the early blocking call). No warning or error was logged. Publish p99 rose from about 0.9 ms to about 5.5 ms.
Expected, not directly tested.
- An application that treats "no Ready provider" as a failure would fail on every publish for the rest of the session.
- Transport-level implicit SHM for large regular payloads appears to be affected as well: in the run above the payloads were well above
message_size_threshold with transport_optimization enabled, yet the receiver saw no SHM samples, and the transport path uses the same try_get_provider. We did not isolate this case separately.
System info
- zenoh / zenoh-c / zenoh-cpp 1.10.1, rustc 1.97.1
- zenoh-c built with
shared-memory and unstable
- Linux x86_64, kernel 5.15.0
- 1.9.0 is not affected only because
obtain_shm_provider had no blocking argument there
Describe the bug
In 1.10.1, calling
obtain_shm_provider(true)(the newblockingargument) before the session's implicit SHM provider has finished initializing returns a Ready provider, but leaves the session's provider state permanentlyInitializing:obtain_shm_provider(false)returnsInitializingfor the life of the sessionErrorCalling the blocking variant after the provider is already Ready is fine. Only a blocking call made while the state is still
Initializingtriggers this.The implicit provider is documented as lazily initialized and available once initialization completes, so a state that never leaves
Initializingdoes not look like the intended contract.To reproduce
Self-contained, no router needed (zenoh-cpp over zenoh-c):
repro.cpp
Output (identical over 3 runs):
The
nonblocking-firstrun is the control: the same session reaches Ready after a few polls and stays Ready, and a blocking call made after that also returns Ready.Likely cause
LazyShmProvider::try_get_providercreates aflume::bounded(1)channel, the init task sends the provider on it once, and the state moves toReadyonly when a non-blocking caller'stry_recv()succeeds:The blocking path in zenoh-c receives that single message through the cloned receiver and returns it to the caller, but nothing moves the shared state to
Ready:After that the channel is empty and the sender is dropped, so every later
try_recv()hits theErr(_)arm (interop.rs#L148) and returnsInitializingagain. A second blocking call getsNonefrom the drained channel and reportsError.Impact
Observed. An application that uses the implicit provider when it is Ready and otherwise sends a regular (non-SHM) payload, and that did one early blocking obtain right after opening the session, never saw the provider become Ready through its non-blocking path. Measured with a 6 MiB payload at 30 Hz for 12 s, 3 runs: 360 / 360 frames were received as non-SHM in every run (vs. 1 cold frame and the rest SHM without the early blocking call). No warning or error was logged. Publish p99 rose from about 0.9 ms to about 5.5 ms.
Expected, not directly tested.
message_size_thresholdwithtransport_optimizationenabled, yet the receiver saw no SHM samples, and the transport path uses the sametry_get_provider. We did not isolate this case separately.System info
shared-memoryandunstableobtain_shm_providerhad noblockingargument there