Summary
Handle node-local schema absence gracefully in distributed stream queries. If the liaison knows the requested schema but a queried data node does not yet know it, that node must not fail the entire query. Return the results available from the other nodes.
This replaces the proposed schema-sync change in #14103. Schema convergence through incremental updates and periodic full reconciliation is intended behavior; this issue does not change its timing or consistency model.
Boundary and behavior
The production boundary is the data-node stream query response and its handling by the liaison's distributed stream query execution path.
- When a participating data node specifically lacks the requested stream schema, treat that node's contribution for that stream/group as empty and merge the remaining results normally.
- Preserve the normal schema-not-found error when the liaison itself cannot resolve the requested schema. Do not turn invalid queries into successful empty results.
- Preserve existing handling of unrelated failures, including timeouts, unavailable nodes, authorization failures, and query execution errors. Do not match arbitrary error-message text or suppress generic errors.
- For multi-group queries, missing schema for one group must not discard valid contributions for other requested groups on the same node. If all queried data nodes lack a schema that the liaison knows, the result for that schema is empty.
This intentionally permits results based on the currently available schema view during convergence. It does not promise a globally complete snapshot: node-local schema absence is not proof that the node has no persisted data.
Example and acceptance criteria
Use a controlled fixture, independent of native logging and reconciliation timers:
- The liaison and data nodes A/B know stream S. A contains an identifiable row A1; B contains B1. Queried node C lacks S.
- Query S through the public stream API on the liaison. It succeeds and returns exactly A1 and B1, respecting the requested ordering and limit, instead of failing because C lacks S.
- Verify the all-data-nodes-missing case, liaison-local schema absence, and a multi-group case where C can still contribute to a different requested group.
- Inject an unrelated node error and verify that the new schema-absence handling does not swallow it.
Add a focused regression test before implementation and demonstrate RED -> GREEN. Add the real distributed query scenario to the existing in-process integration suite; do not require external server binaries or a 150-second wait.
Scope and implementation pointers
Inspect the current main branch before implementation:
banyand/query/processor.go: data-node stream execution-context lookup currently wraps lookup failures in a generic common.Error.
banyand/stream/svc_standalone.go: ErrStreamNotExist identifies stream absence.
banyand/dquery/stream.go and the distributed execution/response aggregation path: distinguish liaison-local validation from remote node failures.
- Existing fixtures:
test/cases/stream and test/integration/distributed/query.
Preserve an unambiguous schema-absence distinction through the actual query transport, or normalize it at the node boundary without affecting liaison-local validation. Verify mixed-version behavior: old generic failures must not accidentally become ignorable. No persisted-format change is intended, and no unbounded retries or new buffering are needed.
Out of scope: revision watermarks, replay/full-reconciliation scheduling, storage durability, flush-on-close, and measure/trace query semantics. No dependency on the native-log implementation in apache/skywalking-banyandb#1368.
Readiness and verification
Bounded implementation leaf: one distributed stream-query behavior, with the data-node response consumed by the existing liaison query path. Relevant callers and fixtures have been inspected, but no new regression test has been authored or run for this issue; establish the failing test on current main before implementation or automation queueing.
Focused checks (after adding the tests):
go test ./banyand/query ./banyand/dquery
go test ./test/integration/distributed/query with the suite's normal prerequisites.
Summary
Handle node-local schema absence gracefully in distributed stream queries. If the liaison knows the requested schema but a queried data node does not yet know it, that node must not fail the entire query. Return the results available from the other nodes.
This replaces the proposed schema-sync change in #14103. Schema convergence through incremental updates and periodic full reconciliation is intended behavior; this issue does not change its timing or consistency model.
Boundary and behavior
The production boundary is the data-node stream query response and its handling by the liaison's distributed stream query execution path.
This intentionally permits results based on the currently available schema view during convergence. It does not promise a globally complete snapshot: node-local schema absence is not proof that the node has no persisted data.
Example and acceptance criteria
Use a controlled fixture, independent of native logging and reconciliation timers:
Add a focused regression test before implementation and demonstrate RED -> GREEN. Add the real distributed query scenario to the existing in-process integration suite; do not require external server binaries or a 150-second wait.
Scope and implementation pointers
Inspect the current main branch before implementation:
banyand/query/processor.go: data-node stream execution-context lookup currently wraps lookup failures in a genericcommon.Error.banyand/stream/svc_standalone.go:ErrStreamNotExistidentifies stream absence.banyand/dquery/stream.goand the distributed execution/response aggregation path: distinguish liaison-local validation from remote node failures.test/cases/streamandtest/integration/distributed/query.Preserve an unambiguous schema-absence distinction through the actual query transport, or normalize it at the node boundary without affecting liaison-local validation. Verify mixed-version behavior: old generic failures must not accidentally become ignorable. No persisted-format change is intended, and no unbounded retries or new buffering are needed.
Out of scope: revision watermarks, replay/full-reconciliation scheduling, storage durability, flush-on-close, and measure/trace query semantics. No dependency on the native-log implementation in apache/skywalking-banyandb#1368.
Readiness and verification
Bounded implementation leaf: one distributed stream-query behavior, with the data-node response consumed by the existing liaison query path. Relevant callers and fixtures have been inspected, but no new regression test has been authored or run for this issue; establish the failing test on current main before implementation or automation queueing.
Focused checks (after adding the tests):
go test ./banyand/query ./banyand/dquerygo test ./test/integration/distributed/querywith the suite's normal prerequisites.