Skip to content

Prevent caller-controlled h5pyd endpoint overrides - #156

Merged
kerberizer merged 4 commits into
mainfrom
security/hsds-target-containment
Sep 6, 2026
Merged

kerberizer merged 4 commits into
mainfrom
security/hsds-target-containment

Conversation

@kerberiel

@kerberiel kerberiel commented Sep 5, 2026 •

Copy link
Copy Markdown
Member

Summary

Prevent request data from replacing h5pyd's configured HSDS network endpoint.

Motivation

The domain parameter should identify a resource on the configured HSDS service, not select another server. The pinned h5pyd version treats http://, https://, and http+unix:// values as endpoint overrides and may forward caller or ambient credentials there.

Solution

  • Reject those three endpoint-bearing prefixes case-insensitively before each existing direct h5pyd.File call.
  • Return 400 Bad Request from the HDF5 download route before h5pyd is invoked.
  • Keep the guard on the legacy and currently unused h5pyd call paths without removing or otherwise changing them.
  • Leave every other domain value unchanged.

Deliberate Scope

This PR only prevents caller-selected h5pyd network endpoints. It does not add a general HSDS path policy or change fragments, suffixes, Unicode, percent escapes, traversal-like strings, controls, relative paths, length limits, copying, errors, authorization, or producer behavior.

hdf5:// and //host/path remain unchanged because neither form selects another network endpoint in the pinned h5pyd implementation. The recognized endpoint forms must be re-audited before upgrading h5pyd.

Testing

  • poetry run pytest tests/test_hsds_endpoint_override.py: 9 passed
  • poetry run pytest: 54 passed, with the existing unrelated tests/test_api.py::test_info failure because the image-generated build version is unavailable locally
  • poetry check: completed with existing deprecation warnings only
  • git diff --check: passed

@kerberizer

Copy link
Copy Markdown
Member

fwiw @vedina

@vedina

vedina commented Sep 6, 2026 •

Copy link
Copy Markdown
Contributor

we should not reject segments like # , because it will break all h5viewers - we frequently want to focus on part of the h5 file, not entire one. So stripping # is no go, it will be functionality breaking.

If paths are validated (non ascii, long, etc), there should be offline validator implemented, so that nexus file writers know what to comply with.

Also check how applying validator on each call affect performance - e.g. h5pyd API requires 10-20 calls for properly showing domain (see e.g. latest in-host spectrasearch nexus overview router)

Besides, when we use HSDS via h5view or even nexusformat library, we are not calling h5pyd.File directly . There is also React code using HSDS API directly.

Therefore validation code should be in a shared library (pyambit) , not here in rcapi, otherwise our own writer may produce invalid domains. And the writer code should offer transformation given a local file path (i.e. original input file) to valid domain. This will be used in import / indexing pipeline.

[Opus review pending]

@vedina

vedina commented Sep 6, 2026

Copy link
Copy Markdown
Contributor

Reviewed at 77b4711. Ran the new suite locally against the committed config: 35 passed.

The core of this is right. Rejecting http://, //host/, .., \ and ?@: before anything reaches h5pyd.File closes the surface that actually matters — a caller being able to steer a server-side open at an arbitrary URL. That part I'd merge as-is.

My comments are all about the rules layered on top of that core, which reject a good deal more than the URL-shaped inputs they were aimed at.

On severity: I checked the rules in #1-#3 against real NeXus files and against the NXpaths the current writer produces, and each of them already has matches. There are files that store and index cleanly today which would return 400 once this lands, and study links that would stop resolving. Nothing about those inputs is hostile — they are ordinary names carrying a space, a percent sign or a non-ASCII character. The failure is also silent from the caller's side: a bare 400 Invalid HSDS domain with no indication of which rule fired, on a request that worked the day before. That is why I would treat #1-#3 as blocking rather than as polish.


1. The fragment is validated and then thrown away

validate_hsds_file_domain splits on #, validates both halves, and returns only the file part. So a request can be refused because of a component the function does not use and the caller never receives.

This is not hypothetical. The fragment is an NXpath, and NXpath segments are group names generated from endpoint and instrument metadata. Names with a leading or trailing space, or containing %, do occur in files written by the current NeXus writer — an endpoint expressed as a percentage is an ordinary way for an assay column to be labelled, and those labels come through from source spreadsheets unaltered. Links to those studies would be refused over a fragment that gets discarded either way.

Suggestion: split on #, validate the file part, return it, and drop the _validate_rooted_path(fragment) call.

2. isascii() is not doing security work

http://, //host/, .. and \ are all ASCII, so none of the rejections that matter depend on this check.

What it does do is reject legitimate file names. A non-breaking space (U+00A0) or an accented character reaching a file name — trivially easy via copy-paste from a spreadsheet — is enough for a file to upload fine and then be permanently undownloadable, returning a bare 400 with nothing to indicate which rule failed.

Suggestion: drop it, or normalize rather than reject. If it stays, the error detail should name the rule that failed so this is diagnosable from a log.

3. The % ban is wider than the thing it protects against

I read this one as guarding against double-decoding — /RRUF/%2e%2e/example.nxs is in the reject list, and if a value is decoded once by the framework and then again downstream, %2e%2e becomes ... That is worth keeping.

But "%" in value also rejects names like 50%_EtOH.nxs, and a percentage in a sample or endpoint name is ordinary. The guard can be kept while dropping the false rejections by rejecting only an actual escape sequence:

_PERCENT_ESCAPE = re.compile(r"%[0-9A-Fa-f]{2}")
...
if _PERCENT_ESCAPE.search(value):
    raise ValueError("invalid HSDS domain: percent-escape")

50%_EtOH.nxs passes, %2e%2e still does not.

Separately, # should stay restricted. A name containing # is genuinely ambiguous against # as the fragment separator and no split rule resolves that cleanly, so rejecting it is the right call — but it is worth a comment saying so, and it is an argument for stripping # in the writer (see #8) rather than only refusing it here.

Taken together with #1 and #2, the rejections these rules add fall almost entirely on valid data rather than on anything hostile.

4. The guard is applied at three call sites rather than one

The check is applied by reassigning a local at each of the three h5pyd.File sites. A fourth site added later inherits nothing, and neither a test nor a type would catch the omission.

A single open_hsds_file(domain, token=None, mode="r") in hsds_domain.py that validates and opens would keep the guard and the sink from drifting apart.

Related: hsds_dataset.py:107 and :118 still split on # inline instead of using the new canonicalizer, so the fragment convention is now encoded in three places.

5. Two guard clauses can never fire

  • if not segments: (hsds_domain.py:41) is unreachable — "".split("/") returns [''], never [].
  • value.startswith("//") is already covered by the empty-segment check in the loop below it.

Both are noise a reader still has to verify.

6. No test asserts a valid domain reaches h5pyd

All three integration-shaped tests assert rejection plus open_file.assert_not_called(). The parametrized accept list does cover the validator itself, but nothing checks that a valid domain gets through an endpoint with its fragment stripped.

That is the test that catches a rule which is defensible in isolation but wrong for the domains the endpoint actually receives — which is the failure mode of #1 through #3.

7. Scope and naming

The module validates syntax. Entitlement is settled downstream — the caller's token is forwarded to h5pyd.File, and HSDS enforces per-user visibility against it — so this is not the layer that decides who may read what, and it shouldn't try to be.

That makes the scope right. But test_hsds_containment.py reads as a broader promise than the code makes, and a later reader may take it for access control.

Suggestion: name it for what it does, e.g. rejecting URL-shaped and traversal-shaped domains.

8. Consider a matching rule on the producer side

This validates at the consumer. The producer — the NeXus writer and indexer — has no equivalent rule, so it is possible to write a file whose domain this API will refuse, and only discover it at download time.

pyambit is already a dependency here. A shared name sanitizer, link builder, and validator living there would let the writer guarantee exactly what this endpoint enforces, from one definition.

To be clear, this would not replace the check in this PR. The domain arrives from an HTTP client, so it has to be validated here regardless of who was supposed to have produced it. The two rules do different jobs; they just shouldn't be two different rules.


Summary: the URL and traversal rejections are worth merging and I would keep them as they are. #1-#3 I would fix before this lands rather than after — each already matches real stored data, and the resulting 400 gives the caller nothing to act on. The rest are cleanups that can follow.

--
Claude Opus 5

@kerberizer

kerberizer commented Sep 6, 2026 •

Copy link
Copy Markdown
Member

Haha, this time I omitted my own internal Opus review, cause it was getting late night -- and here we go. :D

@vedina

vedina commented Sep 6, 2026

Copy link
Copy Markdown
Contributor

Haha, this time I omitted my own internal Opus review, cause it was getting late night -- and here we go. :D

I pointed Opus at the bigger picture, several related repos that are producer / consumers ; so it is not necessary the same :)

@kerberiel

Copy link
Copy Markdown
Member Author

Thanks for the thorough review. I have addressed the compatibility blockers in 6d51260.

The validator now separates the file domain at the first # and validates only that portion. The part after # tells the browser-based viewer which object inside the file to show; it is not part of the file that h5pyd.File opens. Removing it from the h5pyd argument therefore does not alter the browser URL or the viewer's selected object, while validating that unused portion could reject legitimate NeXus object names.

The revision also:

  • accepts Unicode, non-breaking spaces, ordinary percent signs, punctuation, and segment-edge spaces produced by the existing writers;
  • rejects control characters instead of requiring ASCII/isprintable();
  • narrows the percent guard to %HH escape sequences;
  • preserves the legacy .chaold suffix explicitly;
  • removes the unreachable segment check;
  • adds a positive endpoint test proving that a valid composite reference reaches h5pyd as only its file-domain portion; and
  • renames the test module to describe domain validation rather than broader access control.

I kept the request-side validator in ramanchada-api because arbitrary HTTP input must be checked at this trust boundary. A shared producer-side serializer or sanitizer in pyambit would still be useful, but it should complement rather than replace this check and is better handled as a separate cross-repository change. I have likewise left the proposed h5pyd-opening wrapper for follow-up so this PR remains focused.

The validator runs once per h5pyd.File construction, not once per HSDS request, so its runtime cost is negligible. The revised focused suite passes 33 tests; the full local suite passes 77 tests with only the existing generated-version-file failure.

@vedina

vedina commented Sep 6, 2026 •

Copy link
Copy Markdown
Contributor

it would be useful to run the validator them against all domains in hsds so that we know if there is anything to reindex ..

@kerberizer

Copy link
Copy Markdown
Member

Claude Opus 5 here, reviewing on behalf of Luchesar.


Follow-up review of the aggregate diff 068d5d2..6d51260 (base 068d5d266fb875ca05a7e21ac9ae889367b7a9ea, head 6d512606cc7a88d737217376b16c6ed62663a564), not just the latest commit. I formed an independent assessment from the code first, then read @vedina's review at 77b4711 and the response. SHAs rechecked immediately before posting; unchanged. All 7 CI checks green at head.

Verdict: no blocking findings

The %HH, # and 500-vs-400 items below are real, but none of them is a correctness or security blocker, and I am not going to invent one to justify a second round. The containment property this PR claims actually holds, and I verified it against the pinned h5pyd rather than assuming it.

The short version of why: in h5pyd at the pinned rev 8d1c87e, the only construct by which a domain can change the destination is the protocol loop in h5pyd/_hl/files.py — a case-sensitive, literal startswith test for http://, https://, hdf5://, http+unix://. A value that begins with / and not // cannot match any of them. Past that loop the domain is handed to HttpConn and travels as the domain query-parameter value (params["domain"]), never as part of the URL path; the endpoint comes only from the endpoint kwarg or .hscfg/HS_ENDPOINT. So _validate_rooted_path is not a heuristic filter — it is the exact complement of the one construct that overrides the endpoint. That is the right shape for this guard.

All three direct h5pyd.File sinks are guarded and each receives only the validator's return value:

sink guard
convertor.py:131 :64-68, and the sink is reachable only under what == "h5" (a Literal, so no casing bypass)
hsds_dataset.py:182 :181, suffix=".chaold"
convertor_service.py:360 :358

There is no fourth. I also checked the indirect route: pynanomapper has its own h5pyd sinks in clients/h5service.py and clients/service_charisma.py, but this repo imports only datamodel_simple.StudyRaman, so none of them is reachable.

Worth stating the impact plainly, because the PR description undersells it: convertor.py:131 passes api_key=token, the caller's Keycloak bearer token. Before this change, an endpoint override was not merely SSRF — it was a token-exfiltration primitive: point the backend at a host you control and it hands you the token in a request header. That is now closed.

I ran ~30 adversarial variants against the head validator. All rejected, no bypass found: every URL scheme including http+unix:// and hdf5://, uppercase HTTP://, leading tab/newline/space before a scheme, //host/…, .., ., %2e%2e, %252e%252e, backslash, empty segments, trailing slash, CR/LF, DEL, U+2028, .NXS, missing suffix, over-length. The handful of odd inputs that are accepted are inert and I checked each one:

  • /x.nxs#http://evil.test/y.nxs → returns /x.nxs. A URL in the fragment is discarded, which is the whole point of treating it as opaque.
  • /a/b.nxs?x=1 → accepted. This is safe because the domain is a query-parameter value: I ran the pinned stack's requests (2.34.2) and it encodes to domain=%2Fa%2Fb.nxs%3Fx%3D1. ?, @ and : cannot split the query or introduce an authority.
  • U+2044 fraction slash, U+FF0F fullwidth solidus → accepted, but neither is a path separator for h5pyd and neither can form a scheme prefix. These are ordinary Unicode filename characters, which is exactly what review point DevOps enhancements #2 asked to preserve.

One deliberate divergence from the previous review, which I think is correct. @vedina wrote that the ?@: rejection was part of what she'd "merge as-is"; 6d51260 removed it. Given the query-parameter encoding above, that ban was buying nothing in containment terms while rejecting characters that are legal in POSIX filenames. Removing it was the right call, and I'd keep it removed.

Non-blocking findings

N1 — the %HH guard is narrower than before but still catches real names. _PERCENT_ESCAPE is %[0-9A-Fa-f]{2}, and a percent sign followed by two hex letters is not exotic in a materials corpus — a great many element symbols are pure hex. I confirmed these are rejected at head:

/BLOP/PET 50%Fe.nxs     REJECT
/BLOP/PET 50%Ca.nxs     REJECT
/BLOP/20%Ce doped.nxs   REJECT
/BLOP/50%_EtOH.nxs      ok      ← the case #3 was aimed at, now fixed
/OCEAN/café 50%.nxs     ok

Also worth writing down what the rule can and cannot buy, since it is easy to over-credit: Starlette decodes the query string before the validator runs, so a literal %HH arriving here means the client double-encoded. And requests re-encodes on the way out to HSDS (%2e%2e → %252e%252e), so HSDS would have to double-decode a query-parameter value for this to matter at all. It is defense-in-depth against a downstream behaviour, priced in some real names. Two defensible options: (a) keep it, add a one-line comment saying it guards against downstream double-decoding, and sweep the corpora once to confirm nothing stored matches; or (b) drop it and rely on the rooted-path rules, which are what does the security work. I'd take (a) — the sweep is cheap and settles it either way.

N2 — a # in a real file name is rejected, by design. I agree with the previous review that # should stay restricted; splitting at the first # is right for this application. Two properties worth recording, though. The good one: the failure is clean rather than silent — sample #3.nxs truncates to sample which then fails the suffix check, so it 400s instead of opening some other file. The sharp edge: a name shaped like a.nxs#b.nxs truncates to a.nxs and would open a different existing file. That ambiguity is inherent to the composite convention already used at hsds_dataset.py:107, not introduced here, and the real fix is producer-side rejection of # in names (review point #8).

N3 — /db/dataset returns 500, not 400, for an invalid legacy domain. read_cha raises a bare ValueError and get_dataset doesn't catch it, so the global handler produces {"detail":"Internal Server Error"}. I verified this locally (mocked, no network): ?domain=http://evil.test/x.chaold → 500, and h5pyd.File is never reached. Fail-closed, so this is correct security-wise — just inconsistent with the download endpoint's 400 and noisier in the logs than it needs to be.

N4 — the error still doesn't say which rule fired. Point #2 asked for this and it didn't land; all eight raise sites use the identical "invalid HSDS domain". The moment a real name starts 400ing, nobody can tell whether it was the suffix, the percent rule or a segment. A distinct suffix per raise (": percent-escape", ": not rooted", ": suffix") logged server-side, with the flat 400 kept on the wire, would make N1 and N2 diagnosable for roughly nothing.

N5 — the single-wrapper refactor (#4) is still open. All three sinks are guarded today and I confirmed there is no fourth, so the invariant holds — but it lives in three separate assignments, and read_cha's suffix=".chaold" is now a per-call-site contract rather than a property of the function. Reasonable to defer, as agreed; just don't lose it. Relatedly, :107 and :118 still split on # inline. Their semantics do match the new canonicalizer (split('#', 1)[0] ≡ partition("#")[0]) — I checked, so this is duplication rather than divergence.

N6 — /db/dataset accepts a bucket query parameter that is never used (hsds_dataset.py:47). Harmless today since it is never forwarded. Flagging it because bucket is an h5pyd.File kwarg, so if anyone wires it up later it becomes a second caller-controlled input into the same call, and this PR's guard would not cover it. Delete it or wire it deliberately.

N7 — cosmetic. The Cc check covers C0, C1 and DEL but not Cf (bidi overrides, BOM) or Zl/Zp. I confirmed /a/‹U+202E›evil.nxs is accepted mid-path. Inert for containment — a bidi override cannot form a scheme — so this is a log/UI display-spoofing nuisance at most. Mentioning it only so it is a known choice rather than an oversight; the previous review's push toward Unicode permissiveness argues for leaving it as is.

Status of the previous review's findings

# Finding Status at 6d51260
1 Fragment validated then thrown away Fixed. partition("#")[0]; only the file part is validated. The old len(parts) > 2 rejection is also gone, which is right — the first # is the separator, the rest is opaque.
2 isascii() not doing security work Fixed. isascii() and isprintable()/strip() replaced by a control-character check. Unicode, NBSP and segment-edge spaces confirmed accepted. Error-message half of the suggestion did not land → N4.
3 % ban too wide Fixed as suggested, using the exact regex proposed. Residual overbreadth remains → N1.
4 Three call sites, not one wrapper Deferred by agreement (was listed as a cleanup). → N5
5 Two unreachable guard clauses Partially. if not segments: removed. startswith("//") retained — it is redundant with the empty-segment check, but naming the network-path case explicitly is worth one redundant condition. Not worth another round.
6 No test proves a valid domain reaches h5pyd Fixed, and it is doing real work — see below.
7 Scope and naming Fixed. test_hsds_domain_validation.py, and the PR description's Scope section is accurate: it describes syntax and destination containment and explicitly disclaims authorization.
8 Producer-side rule in pyambit Deferred by agreement; complements rather than replaces this check.

Verification evidence

The tests do hit the real FastAPI decoding boundary. This was worth checking rather than assuming — if httpx had stripped the # client-side, test_download_passes_only_file_domain_to_h5pyd would prove much less than it appears to. It doesn't. The actual wire form is:

domain=%2FPROJECT%2Fcaf%C3%A9+50%25_EtOH.nxs%23%2F+endpoint+50%25+%2Fsignal%23raw

So the #, the literal %, the non-ASCII and the spaces all round-trip through Starlette's decoder, and the assertion that h5pyd received exactly /PROJECT/café 50%_EtOH.nxs genuinely proves both server-side fragment stripping and canonical forwarding. I also confirmed the validator sits on the correct side of the decoder — ?domain=%2FRRUF%2F..%2Fx.nxs and ?domain=http%3A%2F%2Fevil.test%2Fx.nxs arrive already decoded and are rejected. 33/33 pass locally at head.

Additional evidence I hadn't expected to get: CI runs the full suite for non-dependabot actors, which includes test_download_domain_h5 — it pulls a real domain from the live Solr index and downloads it through /db/download?what=h5. Green on 3.10/3.11/3.12 at head. That is a genuine end-to-end no-regression signal on real stored data, though pagesize: 1 means it is one sample, not a sweep (see Q1).

@vedina's statements about producer and consumer behaviour check out against source. Taking them as evidence and verifying rather than assuming:

  • Producer. nanodata import_pipeline/tasks/hsds_upload_lib.py:241 builds out_domain = "/{}/{}".format(domain, item.relative_to(folder).as_posix()), where domain is hsds_investigation (RRUFF, BLOP, SLOPP, …). Always rooted, never //, no .., no backslashes. The validator's model matches the producer's output shape exactly. The file-name component is unsanitized, though — e.g. pipeline_nexus/tasks/read_blop.py:326 uses "{}.nxs".format(substance.name) — which is precisely why N1/N2 turn on what those names actually contain.
  • "Stripping # will break all h5viewers." Confirmed as a concern and confirmed as addressed. SpectraSearch h5web.jsx takes the viewer's object path from window.location.hash, not from the query string, and browsers never transmit a fragment; NexusOverviewPage.jsx documents the same #/<entry>/<path> convention. Treating the fragment as opaque and dropping it before the h5pyd call only leaves the browser URL and the viewer's selected object untouched.
  • "We are not always calling h5pyd.File directly; there is React code using the HSDS API directly." Correct, and it is the right reason to scope this PR the way it is scoped: those are the browser talking to HSDS, not the server choosing a destination, so they are not server-side SSRF sinks and this guard neither covers nor needs to cover them.
  • Performance, "10-20 calls to show a domain." Confirmed not applicable to this validator. Those calls go browser → https://hsds.adma.ai directly (hsdsClient.js:39 with config.js:7), never through rcapi. In rcapi the validator runs once per h5pyd.File construction: 1.8 µs for a typical domain, 78 µs at the 2048-char cap. The length check runs before the O(n) scans, so it can't be amplified. Against one HSDS round trip this is unmeasurable.

Questions and assumptions

Q1. I assumed no stored .nxs domain contains % followed by two hex characters, or a #. I could not confirm that from here — the pipeline products aren't committed. @vedina, an rglob("*.nxs") over nexus_folder2import matched against %[0-9A-Fa-f]{2} and # would settle N1 and N2 in one pass, and is the thing I'd most like to see before merge — not as a blocker, but because it converts two judgement calls into facts.

Q2. I assumed the HSDS endpoint always comes from trusted runtime config. Confirmed for h5pyd's own resolution (endpoint kwarg → .hscfg/HS_ENDPOINT); I did not audit how that config is injected in the deployment.

Residual risks, intentionally outside this PR

Listing these so the boundary is explicit, not as asks:

  • This is syntax plus destination containment, not authorization. Entitlement is still decided downstream by HSDS against the forwarded token. The PR description says so; worth keeping that framing in the commit message too.
  • Producer-side sanitization / a shared validator in pyambit (handle broken jsons #8), and the single-wrapper refactor (ops: update the docker build triggers #4).
  • File-content links. A .nxs domain could contain an external link pointing elsewhere, and recursive_copy walks the opened file. Out of scope here and correctly so, but it is the next surface of the same shape.
  • No workload limit on the download path — recursive_copy into an in-memory BytesIO is unbounded in size and time.
  • h5grove is mounted at /h5grove over local files; I checked that fastapi_utils.settings.base_dir is pinned to NEXUS_DIR in main.py, so it is a separate, bounded surface. Noting it only because a reader may wonder whether this PR was meant to cover it. It wasn't.
  • Unrelated to this PR, but adjacent enough to be worth a one-line fix in SpectraSearch: h5web.jsx:52 builds db/download?what=h5&domain=${domain} without encodeURIComponent, unlike Chart.jsx:24. A name containing %, & or + is corrupted before the server ever sees it — e.g. 50%Fe.nxs arrives as mojibake and 404s rather than 400s. Pre-existing.

Bottom line

No blockers. The blocking findings from the previous review (#1, #2, #3) were addressed correctly and, as far as I can tell, without introducing a regression or a bypass: I looked specifically for one in the review-driven deltas, including the ?@: removal, and found none. The core containment property is sound and I verified it against the pinned dependency rather than taking it on trust. The worthwhile follow-ups, in the order I'd do them, are the corpus sweep (Q1), per-rule error detail (N4), the %HH decision that the sweep informs (N1), and then #4/#8.

--
Claude Opus 5

@vedina

vedina commented Sep 6, 2026 •

Copy link
Copy Markdown
Contributor

Let me use simple words, these reports are too long and awkwardly difficult to read by a dumb human.

what my concern is

  • our own writer may produce invalid domain , because the writers do not know the new rules

  • there are already many files that will NOT pass validation, thus breaking the current setup . Before deploying this change, we need to know what to reindex and how.

  • the gate is at rcapi - the HSDS API is available and not gated if directly accessed . There is JS code doing so, let alone direct API access

  • .chaold / .cha could be ignored, they will either migrated or removed. We do not plan supporting .cha . Besides, a nexus file may have other extensions.

  • one does not need a pipeline to check the existing domains !

Producer. nanodata import_pipeline/tasks/hsds_upload_lib.py:241 builds out_domain = "/{}/{}".format(domain, item.relative_to(folder).as_posix()), where domain is hsds_investigation (RRUFF, BLOP, SLOPP, …). Always rooted, never //, no .., no backslashes. The validator's model matches the producer's output shape exactly. The file-name component is unsanitized, though — e.g. pipeline_nexus/tasks/read_blop.py:326 uses "{}.nxs".format(substance.name) — which is precisely why N1/N2 turn on what those names actually contain.

Partly correct. Domain is not hsds_investigation, domain is the entire file path. Which is currently generated based on local paths , usually derived from user files. And they do have all varieties of nonascii symbols.

@kerberiel kerberiel changed the title Validate HSDS .nxs domains before h5pyd access Prevent caller-controlled h5pyd endpoint overrides Sep 6, 2026
@kerberiel
kerberiel requested a review from vedina September 6, 2026 09:45
@kerberiel

Copy link
Copy Markdown
Member Author

@vedina, this is now narrowed to the endpoint override only:

  • reject http://, https://, and http+unix:// before all three direct h5pyd.File calls;
  • pass every other value unchanged, including fragments, Unicode, percent signs, suffixes, hdf5://, and //host/path;
  • make no writer rules, path policy, stale-code removal, or other broader changes.

Because existing file names are no longer validated, this change should not require reindexing them. New head: d24b2c5. The focused suite passes 30 tests; the full local suite passes 75 with only the existing generated-version test failure.

Could you please re-review this smaller aggregate diff?

@kerberiel
kerberiel marked this pull request as draft September 6, 2026 10:05
@kerberiel

Copy link
Copy Markdown
Member Author

Small follow-up at 457a8b7: I trimmed the focused test matrix from 30 cases to 9. The removed cases were remnants of the earlier path-validation scope and made this endpoint-only change look broader than it is.

Production code is unchanged. The focused suite now covers only the three endpoint-replacing prefixes, the two explicitly non-overriding h5pyd forms, one ordinary pass-through case, and rejection before each direct h5pyd sink.

@kerberiel
kerberiel marked this pull request as ready for review September 6, 2026 10:07
@kerberizer
kerberizer merged commit 0903156 into main Sep 6, 2026
9 checks passed
@kerberizer
kerberizer deleted the security/hsds-target-containment branch September 6, 2026 12:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants