Skip to content

Local mode: writing a sparse vector to a dense field leaves the collection unreadable #1461

Description

@HuaTNA

Reproduced on upstream dev at
589a87a10679680921571828d800a120425fa807, Python 3.13.13, on 2026-09-20.

Reproduction

from qdrant_client import QdrantClient, models

client = QdrantClient(":memory:")
client.create_collection(
    "example",
    vectors_config={"dense": models.VectorParams(size=2, distance=models.Distance.DOT)},
)
client.upsert("example", [models.PointStruct(id=1, vector={"dense": [1.0, 2.0]})])

try:
    client.upsert("example", [models.PointStruct(
        id=2,
        vector={"dense": models.SparseVector(indices=[0], values=[3.0])},
    )])
except Exception as exc:
    print(type(exc).__name__, str(exc))

print(client.count("example").count)
print(client.retrieve("example", [1, 2], with_vectors=True))
client.close()

Before the proposed fix, the upsert raises TypeError, but the count is now 2.
Retrieving the points then raises IndexError: the failed insertion has left
the point IDs and vector storage out of sync.

Expected: reject the incompatible vector before inserting the invalid point,
and leave the collection usable.

Cause and proposed fix

_validate_point and _validate_named_vectors check that a vector name exists,
but do not check whether it designates sparse storage before dispatching on the
input's type. A SparseVector can therefore pass validation for a dense or
multivector field. Conversely, a dense list can pass validation for a sparse
field. Failures occur later in the mutation path.

The local patch adds a shared sparse-versus-dense kind check in both validation
paths. It raises ValueError before collection state is changed. Batch updates
use these validation paths too. Dense-versus-multivector shape validation is
outside this patch's scope.

Validation

  • 18 new regression cases fail before the fix and pass after it.
  • Cases cover upsert, update_vectors and batch updates; sparse inputs for dense
    and multivector fields and dense inputs for sparse fields; memory and disk
    storage. They check original data, count, later queries and writes, and
    reopening persistent storage.
  • 32 tests pass together across the new test file, test_in_memory.py and
    test_local_persistence.py.
  • Ruff formatting, lint of the new test file, and git diff whitespace checks pass.
  • A temporary Qdrant v1.19.1 server rejects the nine corresponding single-invalid-
    operation cases with HTTP 400 and preserves the original point and count.

Server comparison limitation: a mixed request containing a valid point before
an invalid one can apply the valid point before returning an error. The local
regression tests preserve local mode's existing prevalidation behavior; they
do not assert that whole-request atomicity matches the server.

Duplicate check and disclosure

GitHub issues and PRs were searched on 2026-09-20 for sparse/dense validation,
wrong vector kinds, SparseVector TypeError, corruption, and
_validate_named_vectors. No exact duplicate was found in those searches.
Related PR #1445 concerns query errors; #1450 concerns non-finite vector values.

AI assistance was used to investigate, implement and test this local patch.
The reproduction and test results above were executed locally with AI assistance.

A focused fix and regression tests are prepared locally. Would this validation approach be welcome? I can submit the patch against dev after confirmation.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions