Skip to content

test: add nnunet pipeline tests - #395

Merged
JoshuaSiraj merged 11 commits into
mainfrom
JoshuaSiraj/nnunet_updates
Jun 30, 2026
Merged

JoshuaSiraj merged 11 commits into
mainfrom
JoshuaSiraj/nnunet_updates

Conversation

@JoshuaSiraj

@JoshuaSiraj JoshuaSiraj commented Jun 16, 2025 •

Copy link
Copy Markdown
Collaborator

Summary by CodeRabbit

  • New Features

    • Added train/test splitting for successfully processed cases via new --test-set-ratio and --random-seed CLI options, with deterministic selection and automatic dataset updates.
    • Updated output organization so saved files reflect the correct PatientID-based paths.
  • Bug Fixes

    • Made the pipeline crawl directory optional and safely handled absent values without forcing path conversion.
  • Tests

    • Added/extended integration coverage for CLI help, validation errors, end-to-end runs, and verification of expected imagesTs/labelsTs counts and fingerprint extraction.

@coderabbitai

coderabbitai Bot commented Jun 16, 2025 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Warning

Review limit reached

@JoshuaSiraj, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 46 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: c815bceb-41b6-44d3-bb8e-aca92734b2a4

📥 Commits

Reviewing files that changed from the base of the PR and between 4bb215c and 41a13d7.

⛔ Files ignored due to path filters (1)
  • pixi.lock is excluded by !**/*.lock and included by none
📒 Files selected for processing (2)
  • src/imgtools/io/nnunet_output.py
  • src/imgtools/nnunet_pipeline.py
📝 Walkthrough

Walkthrough

nnUNetPipeline now threads test_set_ratio and random_seed through the CLI, pipeline, and output layer, and nnUNetOutput can move a seeded subset of successful cases into the test split while patching the dataset index. New CLI integration tests cover help text, validation, and an end-to-end run.

Changes

nnUNet train/test split and CLI coverage

Layer / File(s) Summary
CLI options and command wiring
src/imgtools/cli/nnunet_pipeline.py
Adds --test-set-ratio and --random-seed, updates the Click command signature, and passes both values into nnUNetPipeline.
Pipeline parameter flow and split call
src/imgtools/nnunet_pipeline.py
Makes crawl_directory optional, forwards the split settings into nnUNetOutput, and calls split_dataset after processing successful samples.
Output split implementation
src/imgtools/io/nnunet_output.py
Adds split configuration fields, changes the filename template to use PatientID, and implements deterministic file moves plus CSV patching for the train/test split.
CLI integration tests
tests/integration/cli/test_nnunet_pipeline_cli.py
Adds fixtures and tests for CLI help, missing-option validation, and an end-to-end pipeline run that checks generated split counts and fingerprint extraction.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title is concise and accurately reflects the added nnUNet pipeline test coverage, though it omits the broader pipeline and CLI changes.
Docstring Coverage ✅ Passed Docstring coverage is 88.24% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch JoshuaSiraj/nnunet_updates

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@JoshuaSiraj

Copy link
Copy Markdown
Collaborator Author

@jjjermiah

Added testing with RADCURE, only issue is that nnunet is not available for py13.

@JoshuaSiraj

Copy link
Copy Markdown
Collaborator Author

The issue with env is nnunet install on py313(osx-64)

@jjjermiah jjjermiah left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

looks great to me, to address the py313 issue, I recommend having a check in the cli entry point that lets the user know

Comment thread tests/integration/cli/test_nnunet_pipeline_cli.py
Comment thread tests/integration/cli/test_nnunet_pipeline_cli.py Outdated
@JoshuaSiraj

Copy link
Copy Markdown
Collaborator Author

looks great to me, to address the py313 issue, I recommend having a check in the cli entry point that lets the user know

Thanks! How do I deal with the dependency issue? I need to add nnunetv2 to test as a dependency, instead should I add it only to py311 and py312?

@JoshuaSiraj
JoshuaSiraj marked this pull request as ready for review June 30, 2025 15:52

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

♻️ Duplicate comments (1)
tests/integration/cli/nnunet_pipeline_cli.py (1)

96-96: Consider implementing pytest-snapshot for consistent output validation.

As suggested in previous reviews, using pytest-snapshot would help ensure the generated outputs remain consistent across test runs and catch any unintended changes in the pipeline output format.

This would involve capturing and comparing the generated dataset structure, JSON configurations, and file contents against known good snapshots.

🧹 Nitpick comments (2)
tests/integration/cli/nnunet_pipeline_cli.py (2)

12-13: Fix class naming and docstring accuracy.

The class name should follow PEP 8 conventions, and the docstring mentions the wrong CLI command.

-class TestnnUNetCLI:
-    """Integration tests for the autopipeline CLI command using collections from the test data."""
+class TestNnUNetCLI:
+    """Integration tests for the nnunet_pipeline CLI command using collections from the test data."""

63-70: Consider moving YAML file creation to a fixture.

The temporary YAML file creation could be moved to a fixture for better test organization and reusability.

+    @pytest.fixture(scope="function")
+    def roi_yaml_file(self, tmp_path):
+        """Create a temporary ROI mapping YAML file."""
+        roi_dict = {
+            "BRAINSTEM": "Brainstem",
+            "SPINALCORD": "SpinalCord", 
+            "LARYNX": "Larynx",
+        }
+        roi_yaml_path = tmp_path / "roi_match.yaml"
+        with roi_yaml_path.open("w") as f:
+            yaml.dump(roi_dict, f)
+        return roi_yaml_path

Then use roi_yaml_file fixture in the test method instead of creating the file inline.

📜 Review details

Configuration used: .coderabbit.yaml
Review profile: CHILL
Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 4df690a and 69e697e.

⛔ Files ignored due to path filters (2)
  • pixi.lock is excluded by !**/*.lock and included by none
  • pixi.toml is excluded by none and included by none
📒 Files selected for processing (2)
  • src/imgtools/io/nnunet_output.py (1 hunks)
  • tests/integration/cli/nnunet_pipeline_cli.py (1 hunks)
🧰 Additional context used
📓 Path-based instructions (2)
`src/**/*.py`: Review the Python code for compliance with PEP 8 and PEP 257 (doc...

src/**/*.py: Review the Python code for compliance with PEP 8 and PEP 257 (docstring conventions). Ensure the following: - Variables and functions follow meaningful naming conventions. - Docstrings are present, accurate, and align with the implementation. - Code is efficient and avoids redundancy while adhering to DRY principles. - Consider suggestions to enhance readability and maintainability. - Highlight any potential performance issues, edge cases, or logical errors. - Ensure all imported libraries are used and necessary.

⚙️ Source: CodeRabbit Configuration File

List of files the instruction was applied to:

  • src/imgtools/io/nnunet_output.py
`tests/**/*`: Review the test code written with Pytest. Confirm: - Tests cover a...

tests/**/*: Review the test code written with Pytest. Confirm: - Tests cover all critical functionality and edge cases. - Test descriptions clearly describe their purpose. - Pytest best practices are followed, such as proper use of fixtures. - Ensure the tests are isolated and do not have external dependencies (e.g., network calls). - Verify meaningful assertions and avoidance of redundant tests. - Test code adheres to PEP 8 style guidelines.

⚙️ Source: CodeRabbit Configuration File

List of files the instruction was applied to:

  • tests/integration/cli/nnunet_pipeline_cli.py
🧠 Learnings (3)
📓 Common learnings
Learnt from: jjjermiah
PR: bhklab/med-imagetools#145
File: src/imgtools/utils/nnunet.py:0-0
Timestamp: 2024-11-29T21:18:38.153Z
Learning: Suggestions to modify the `save_json` function in `src/imgtools/utils/nnunet.py` to fix type annotations or add error handling are considered out of scope.
src/imgtools/io/nnunet_output.py (1)
Learnt from: jjjermiah
PR: bhklab/med-imagetools#145
File: src/imgtools/utils/nnunet.py:0-0
Timestamp: 2024-11-29T21:18:38.153Z
Learning: Suggestions to modify the `save_json` function in `src/imgtools/utils/nnunet.py` to fix type annotations or add error handling are considered out of scope.
tests/integration/cli/nnunet_pipeline_cli.py (3)
Learnt from: jjjermiah
PR: bhklab/med-imagetools#137
File: src/imgtools/dicom/sort/parser.py:132-137
Timestamp: 2024-11-21T21:03:45.548Z
Learning: Assertions are allowed for input validation throughout the project, including in `src/imgtools/dicom/sort/parser.py`.
Learnt from: jjjermiah
PR: bhklab/med-imagetools#137
File: src/imgtools/dicom/sort/utils.py:66-67
Timestamp: 2024-11-21T16:24:04.091Z
Learning: In the `src/imgtools/dicom/sort/utils.py` file and throughout the codebase, assertions are preferred for input validation instead of explicit type checks or raising exceptions.
Learnt from: jjjermiah
PR: bhklab/med-imagetools#137
File: src/imgtools/dicom/sort/utils.py:168-169
Timestamp: 2024-11-21T16:24:13.275Z
Learning: In `src/imgtools/dicom/sort/utils.py`, it's acceptable to use `assert` statements for input validation in the `read_tags` function.
🧬 Code Graph Analysis (1)
src/imgtools/io/nnunet_output.py (1)
src/imgtools/coretypes/base_masks.py (1)
  • roi_keys (161-163)
🔇 Additional comments (1)
src/imgtools/io/nnunet_output.py (1)

301-303: LGTM: Preserves background label correctly.

The change from reassigning to updating the labels dictionary in-place is a good fix. This ensures the background label (0) is preserved while adding ROI keys with proper incremental indexing, which aligns with nnUNet's expected label format.

Comment on lines +87 to +94
nnunet_result = subprocess.run([
"nnUNetv2_extract_fingerprint",
"-d", "1",
"--verify_dataset_integrity",
],
env=env,
stdout=subprocess.PIPE
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🛠️ Refactor suggestion

Improve subprocess call robustness.

The subprocess call lacks timeout and proper error handling, which could cause tests to hang or hide important error information.

         nnunet_result = subprocess.run([
             "nnUNetv2_extract_fingerprint",
             "-d", "1",
             "--verify_dataset_integrity",
             ],
             env=env,
-            stdout=subprocess.PIPE
+            stdout=subprocess.PIPE,
+            stderr=subprocess.PIPE,
+            timeout=300,  # 5 minute timeout
+            text=True
         )

Also consider logging the stderr output when the command fails for better debugging:

-        assert nnunet_result.returncode == 0, "nnUNetv2_extract_fingerprint failed"
+        assert nnunet_result.returncode == 0, f"nnUNetv2_extract_fingerprint failed: {nnunet_result.stderr}"
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
nnunet_result = subprocess.run([
"nnUNetv2_extract_fingerprint",
"-d", "1",
"--verify_dataset_integrity",
],
env=env,
stdout=subprocess.PIPE
)
nnunet_result = subprocess.run([
"nnUNetv2_extract_fingerprint",
"-d", "1",
"--verify_dataset_integrity",
],
env=env,
stdout=subprocess.PIPE,
stderr=subprocess.PIPE,
timeout=300, # 5 minute timeout
text=True
)
assert nnunet_result.returncode == 0, f"nnUNetv2_extract_fingerprint failed: {nnunet_result.stderr}"
🤖 Prompt for AI Agents
In tests/integration/cli/nnunet_pipeline_cli.py around lines 87 to 94, the
subprocess.run call lacks a timeout and error handling, risking hangs and
obscured errors. Add a timeout parameter to prevent indefinite blocking, use
check=True to raise an exception on failure, and catch
subprocess.CalledProcessError to log stderr output for better debugging.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
tests/integration/cli/test_nnunet_pipeline_cli.py (1)

35-52: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Tighten these error assertions.

Allowing or "Error:" in result.output means any Click failure can satisfy the test, even if the wrong validation fired. Assert the specific missing-option message in each case.

As per path instructions, tests/**/*: Verify meaningful assertions and avoidance of redundant tests.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/integration/cli/test_nnunet_pipeline_cli.py` around lines 35 - 52, The
invalid-args checks in test_invalid_args are too loose because they accept any
Click failure via the generic "Error:" fallback, which can hide the wrong
validation path. Tighten the assertions to match the exact missing-option
messages for each runner.invoke call in test_nnunet_pipeline_cli, using the
nnunet_pipeline CLI behavior as the source of truth. Keep the checks specific to
the expected missing --modalities and --roi-match-yaml/-ryaml failures so the
test only passes when the intended validation fires.

Source: Path instructions

🧹 Nitpick comments (1)
src/imgtools/nnunet_pipeline.py (1)

48-48: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Document the new crawl_directory parameter.

__init__ now exposes crawl_directory, but the constructor docstring still omits it, so the public API docs no longer match the implementation.

Proposed update
         mask_saving_strategy : MaskSavingStrategy
             Mask saving strategy
+        crawl_directory : str | Path | None, optional
+            Directory used for crawl metadata. If omitted, the default crawl
+            location is used.
         update_crawl : bool, optional
             Whether to force recrawling, by default False

As per path instructions, src/**/*.py: Review the Python code for compliance with PEP 8 and PEP 257 (docstring conventions).

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/imgtools/nnunet_pipeline.py` at line 48, Update the __init__ docstring in
nnunet_pipeline so it documents the new crawl_directory parameter alongside the
other constructor args. Add a concise parameter description that matches its
type and behavior, and keep the docstring aligned with PEP 257 conventions by
using the same style as the existing constructor documentation.

Source: Path instructions

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/integration/cli/test_nnunet_pipeline_cli.py`:
- Around line 53-55: Gate test_nnunet_pipeline on nnUNet availability before
invoking nnUNetv2_extract_fingerprint so it skips cleanly when the binary is
missing, and keep the subprocess-based integration path isolated. Also tighten
the existing macOS skip in test_nnunet_pipeline so it only excludes the actual
Python 3.13 nnUNet incompatibility rather than all Darwin runs; use the
test_nnunet_pipeline method and the nnUNetv2_extract_fingerprint call site as
the key places to update.

---

Outside diff comments:
In `@tests/integration/cli/test_nnunet_pipeline_cli.py`:
- Around line 35-52: The invalid-args checks in test_invalid_args are too loose
because they accept any Click failure via the generic "Error:" fallback, which
can hide the wrong validation path. Tighten the assertions to match the exact
missing-option messages for each runner.invoke call in test_nnunet_pipeline_cli,
using the nnunet_pipeline CLI behavior as the source of truth. Keep the checks
specific to the expected missing --modalities and --roi-match-yaml/-ryaml
failures so the test only passes when the intended validation fires.

---

Nitpick comments:
In `@src/imgtools/nnunet_pipeline.py`:
- Line 48: Update the __init__ docstring in nnunet_pipeline so it documents the
new crawl_directory parameter alongside the other constructor args. Add a
concise parameter description that matches its type and behavior, and keep the
docstring aligned with PEP 257 conventions by using the same style as the
existing constructor documentation.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: 640c32ef-8ff9-4339-9eb4-b2765b1dc241

📥 Commits

Reviewing files that changed from the base of the PR and between 69e697e and 2ff2c55.

⛔ Files ignored due to path filters (2)
  • pixi.lock is excluded by !**/*.lock and included by none
  • pixi.toml is excluded by none and included by none
📒 Files selected for processing (2)
  • src/imgtools/nnunet_pipeline.py
  • tests/integration/cli/test_nnunet_pipeline_cli.py

Comment on lines +53 to +55
@pytest.mark.skipif(platform == "darwin", reason="Test skipped on macOS, due to nnUNet py313 incompatibility")
@pytest.mark.parametrize("mask_saving_strategy", ["sparse_mask", "region_mask"])
def test_nnunet_pipeline(self, runner, temp_output_dir, DATA_DIR, mask_saving_strategy):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🩺 Stability & Availability | 🟠 Major | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf '\n== target file ==\n'
wc -l tests/integration/cli/test_nnunet_pipeline_cli.py
sed -n '1,220p' tests/integration/cli/test_nnunet_pipeline_cli.py

printf '\n== nnUNet references ==\n'
rg -n "nnUNetv2_extract_fingerprint|skipif|shutil.which|platform == \"darwin\"|py313|nnUNet" tests . -g '!**/.git/**'

Repository: bhklab/med-imagetools

Length of output: 12345


Gate this test on nnUNet availability, and narrow the macOS skip to the actual py3.13 incompatibility.
This test still unconditionally runs nnUNetv2_extract_fingerprint, so environments without that binary will fail instead of skipping cleanly. A shutil.which(...) check before the subprocess call would keep the integration test isolated.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/integration/cli/test_nnunet_pipeline_cli.py` around lines 53 - 55, Gate
test_nnunet_pipeline on nnUNet availability before invoking
nnUNetv2_extract_fingerprint so it skips cleanly when the binary is missing, and
keep the subprocess-based integration path isolated. Also tighten the existing
macOS skip in test_nnunet_pipeline so it only excludes the actual Python 3.13
nnUNet incompatibility rather than all Darwin runs; use the test_nnunet_pipeline
method and the nnUNetv2_extract_fingerprint call site as the key places to
update.

Source: Path instructions

@codecov

codecov Bot commented Jun 29, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 86.53846% with 7 lines in your changes missing coverage. Please review.
✅ Project coverage is 65.90%. Comparing base (b3ed192) to head (41a13d7).

Files with missing lines Patch % Lines
src/imgtools/io/nnunet_output.py 86.00% 7 Missing ⚠️
Additional details and impacted files
@@             Coverage Diff             @@
##             main     #395       +/-   ##
===========================================
+ Coverage   54.85%   65.90%   +11.05%     
===========================================
  Files          66       66               
  Lines        4370     4420       +50     
===========================================
+ Hits         2397     2913      +516     
+ Misses       1973     1507      -466     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

kellofoggs and others added 2 commits June 29, 2026 16:57
Hi sorry issue took a little while, but I wrote up a minimal way of
implementing a test_set_split for nnunet-pipeline.

The feature takes the successfully processed images and moves a
proportion of the images/labels to the appropriate dirs in the output
dir. I could not find any tests for nnunet-pipeline but this should
probably be added to them when they're ready. There is CLI support for
the feature as well as hints. Please let me know if there's anything II
should change style wise.


Fixes #443



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added a `test-set-ratio` option to the command-line workflow, letting
you split a portion of successfully processed data into a test set.
* The pipeline now automatically moves the selected files into the
appropriate test folders after processing.

* **Bug Fixes**
* Added validation for the new ratio value to keep it within a valid
range and prevent invalid dataset splits.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: kojukwu <kojukwu@gmail.com>
This update introduces a `test-set-ratio` and `random-seed` option to the nnUNet pipeline, allowing users to specify the proportion of successfully processed cases to be moved to the test set. The implementation includes validation for the ratio and updates to the dataset splitting logic. Additionally, the CLI has been updated to reflect these new options, ensuring a smoother user experience.

Tests have been added to verify the correct functionality of the new features.

Fixes #443

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (4)
src/imgtools/nnunet_pipeline.py (2)

49-59: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Keep the constructor docstring aligned with the new signature.

crawl_directory is now optional but still undocumented, and the new split parameters omit their defaults. This makes the public API harder to use correctly.

Proposed docstring update
+        crawl_directory : str | Path | None, optional
+            Directory for cached crawl data, by default None.
...
-        test_set_ratio : float
+        test_set_ratio : float, default=0.0
             Proportion of successful cases for the test set. Count is
             ceil(ratio * n_cases); 1.0 moves all cases to the test set.
-        random_seed : int
+        random_seed : int, default=42
             The random seed to use for the test set split.

As per path instructions, src/**/*.py: “Docstrings are present, accurate, and align with the implementation.”

Also applies to: 92-96

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/imgtools/nnunet_pipeline.py` around lines 49 - 59, The constructor
docstring for the nnUNet pipeline class is out of sync with the updated
signature: it must document the new optional crawl_directory parameter and
include the default values for the new split-related arguments. Update the
docstring where the initializer is described so it matches the public API
exposed by the constructor symbols, especially crawl_directory,
existing_file_mode, update_crawl, n_jobs, roi_ignore_case,
roi_allow_multi_key_matches, spacing, window, level, test_set_ratio, and
random_seed.

Source: Path instructions


159-159: 📐 Maintainability & Code Quality | 🔵 Trivial

Track this broad TODO outside the implementation.

The note is valid, but it is broad enough to become stale in code. Consider turning it into an issue or a more specific TODO. I can help draft the refactor ticket if useful.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/imgtools/nnunet_pipeline.py` at line 159, The TODO comment in the long
nnunet_pipeline function is too broad to keep in the implementation and will
likely go stale. Remove the inline “This function is long and has a lot of
concerns built into it” note and replace it with a more specific, actionable
TODO tied to the exact refactor needed, or move the broader refactor tracking to
an external issue/ticket. Use the surrounding pipeline function context in
src/imgtools/nnunet_pipeline.py to locate and update the comment.
src/imgtools/cli/nnunet_pipeline.py (1)

122-122: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Apply PEP 8 spacing in the new option and parameter.

Tiny cleanup, but it keeps the touched CLI surface consistent with the project’s Python style.

Proposed cleanup
-    type=click.FloatRange(0.0,1.0),
+    type=click.FloatRange(0.0, 1.0),
...
-    test_set_ratio:float,
+    test_set_ratio: float,

As per path instructions, src/**/*.py: “Review the Python code for compliance with PEP 8 and PEP 257.”

Also applies to: 158-158

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/imgtools/cli/nnunet_pipeline.py` at line 122, The new CLI option and its
corresponding parameter in nnunet_pipeline.py need PEP 8 spacing cleanup. Update
the affected click option definition and the matching function signature so the
decorator arguments and parameter list use standard spacing/style consistently
with the rest of the CLI, keeping the change localized to the relevant option
block and its associated callable.

Source: Path instructions

src/imgtools/io/nnunet_output.py (1)

19-19: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Make ProcessSampleResult a type-only import. It’s only used in split_dataset() annotations, so move it under TYPE_CHECKING (or rely on postponed annotations) to avoid importing imgtools.autopipeline at runtime.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/imgtools/io/nnunet_output.py` at line 19, `ProcessSampleResult` is only
referenced in `split_dataset()` type annotations, so it should not be imported
at runtime. Update the imports in `nnunet_output` to make `ProcessSampleResult`
type-only by moving it under a `TYPE_CHECKING` guard or relying on postponed
annotation behavior, and keep the runtime path free of the
`imgtools.autopipeline` import.

Source: Path instructions

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/imgtools/io/nnunet_output.py`:
- Line 201: The nnUNet output filename template in the path-building logic
currently embeds PatientID, which can expose PHI in generated artifacts. Update
the filename construction in the nnunet output helper to use a non-identifying
identifier such as SampleID or another generated ID instead of PatientID, and
make sure any related naming helpers or format strings that reference PatientID
are adjusted consistently.
- Around line 316-341: The split index patching logic in the helper that reads
self.writer.index_file should fail fast instead of silently returning when the
index file is missing after split_dataset() has moved files. Update the guard in
this path to raise an exception with clear context about the missing index and
the moved_sample_ids involved, so callers can detect the inconsistency rather
than continuing with stale metadata.

---

Nitpick comments:
In `@src/imgtools/cli/nnunet_pipeline.py`:
- Line 122: The new CLI option and its corresponding parameter in
nnunet_pipeline.py need PEP 8 spacing cleanup. Update the affected click option
definition and the matching function signature so the decorator arguments and
parameter list use standard spacing/style consistently with the rest of the CLI,
keeping the change localized to the relevant option block and its associated
callable.

In `@src/imgtools/io/nnunet_output.py`:
- Line 19: `ProcessSampleResult` is only referenced in `split_dataset()` type
annotations, so it should not be imported at runtime. Update the imports in
`nnunet_output` to make `ProcessSampleResult` type-only by moving it under a
`TYPE_CHECKING` guard or relying on postponed annotation behavior, and keep the
runtime path free of the `imgtools.autopipeline` import.

In `@src/imgtools/nnunet_pipeline.py`:
- Around line 49-59: The constructor docstring for the nnUNet pipeline class is
out of sync with the updated signature: it must document the new optional
crawl_directory parameter and include the default values for the new
split-related arguments. Update the docstring where the initializer is described
so it matches the public API exposed by the constructor symbols, especially
crawl_directory, existing_file_mode, update_crawl, n_jobs, roi_ignore_case,
roi_allow_multi_key_matches, spacing, window, level, test_set_ratio, and
random_seed.
- Line 159: The TODO comment in the long nnunet_pipeline function is too broad
to keep in the implementation and will likely go stale. Remove the inline “This
function is long and has a lot of concerns built into it” note and replace it
with a more specific, actionable TODO tied to the exact refactor needed, or move
the broader refactor tracking to an external issue/ticket. Use the surrounding
pipeline function context in src/imgtools/nnunet_pipeline.py to locate and
update the comment.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: f91f4f82-9b4e-405d-af3f-a5d49944b167

📥 Commits

Reviewing files that changed from the base of the PR and between 2ff2c55 and 4bb215c.

⛔ Files ignored due to path filters (1)
  • pixi.toml is excluded by none and included by none
📒 Files selected for processing (4)
  • src/imgtools/cli/nnunet_pipeline.py
  • src/imgtools/io/nnunet_output.py
  • src/imgtools/nnunet_pipeline.py
  • tests/integration/cli/test_nnunet_pipeline_cli.py
🚧 Files skipped from review as they are similar to previous changes (1)
  • tests/integration/cli/test_nnunet_pipeline_cli.py


self._file_name_format = (
"{DirType}{SplitType}/{Dataset}_{SampleID}.nii.gz"
"{DirType}{SplitType}/{PatientID}_{SampleID}.nii.gz"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟠 Major | ⚡ Quick win

Avoid embedding PatientID in generated filenames.

PatientID is patient-identifying metadata; putting it in every nnUNet filename can leak PHI through paths, logs, reports, and shared artifacts. Prefer the generated sample ID or another non-PHI identifier.

Proposed safer template
-            "{DirType}{SplitType}/{PatientID}_{SampleID}.nii.gz"
+            "{DirType}{SplitType}/{SampleID}.nii.gz"
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
"{DirType}{SplitType}/{PatientID}_{SampleID}.nii.gz"
"{DirType}{SplitType}/{SampleID}.nii.gz"
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/imgtools/io/nnunet_output.py` at line 201, The nnUNet output filename
template in the path-building logic currently embeds PatientID, which can expose
PHI in generated artifacts. Update the filename construction in the nnunet
output helper to use a non-identifying identifier such as SampleID or another
generated ID instead of PatientID, and make sure any related naming helpers or
format strings that reference PatientID are adjusted consistently.

Comment on lines +316 to +341
index_file = self.writer.index_file
if not moved_sample_ids or not index_file.exists():
return

dir_map = {"imagesTr": "imagesTs", "labelsTr": "labelsTs"}
df = pd.read_csv(index_file, dtype={"SampleID": str})

def belongs_to_moved_case(sample_id: str) -> bool:
sample_id = str(sample_id)
return any(
sample_id == mid or sample_id.startswith(f"{mid}_")
for mid in moved_sample_ids
)

def rewrite_split_path(filepath: str) -> str:
parts = Path(filepath).parts
if not parts or parts[0] not in dir_map:
return filepath
return str(Path(dir_map[parts[0]], *parts[1:]))

mask = df["SampleID"].map(belongs_to_moved_case)
df.loc[mask, "filepath"] = df.loc[mask, "filepath"].map(rewrite_split_path)
if "SplitType" in df.columns:
df.loc[mask, "SplitType"] = "Ts"

df.to_csv(index_file, index=False)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Fail fast when the split index cannot be patched.

split_dataset() moves files before calling this helper; if the index file is missing, Line 317 silently returns and leaves moved files without synchronized index metadata. Raise with context instead.

Proposed guard
         index_file = self.writer.index_file
-        if not moved_sample_ids or not index_file.exists():
+        if not moved_sample_ids:
             return
+        if not index_file.exists():
+            raise FileNotFoundError(
+                f"Cannot patch test split index because {index_file} does not exist."
+            )
+
+        required_columns = {"SampleID", "filepath"}
         dir_map = {"imagesTr": "imagesTs", "labelsTr": "labelsTs"}
         df = pd.read_csv(index_file, dtype={"SampleID": str})
+        missing_columns = required_columns - set(df.columns)
+        if missing_columns:
+            raise ValueError(
+                f"Cannot patch test split index; missing columns: {sorted(missing_columns)}"
+            )
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
index_file = self.writer.index_file
if not moved_sample_ids or not index_file.exists():
return
dir_map = {"imagesTr": "imagesTs", "labelsTr": "labelsTs"}
df = pd.read_csv(index_file, dtype={"SampleID": str})
def belongs_to_moved_case(sample_id: str) -> bool:
sample_id = str(sample_id)
return any(
sample_id == mid or sample_id.startswith(f"{mid}_")
for mid in moved_sample_ids
)
def rewrite_split_path(filepath: str) -> str:
parts = Path(filepath).parts
if not parts or parts[0] not in dir_map:
return filepath
return str(Path(dir_map[parts[0]], *parts[1:]))
mask = df["SampleID"].map(belongs_to_moved_case)
df.loc[mask, "filepath"] = df.loc[mask, "filepath"].map(rewrite_split_path)
if "SplitType" in df.columns:
df.loc[mask, "SplitType"] = "Ts"
df.to_csv(index_file, index=False)
index_file = self.writer.index_file
if not moved_sample_ids:
return
if not index_file.exists():
raise FileNotFoundError(
f"Cannot patch test split index because {index_file} does not exist."
)
required_columns = {"SampleID", "filepath"}
dir_map = {"imagesTr": "imagesTs", "labelsTr": "labelsTs"}
df = pd.read_csv(index_file, dtype={"SampleID": str})
missing_columns = required_columns - set(df.columns)
if missing_columns:
raise ValueError(
f"Cannot patch test split index; missing columns: {sorted(missing_columns)}"
)
def belongs_to_moved_case(sample_id: str) -> bool:
sample_id = str(sample_id)
return any(
sample_id == mid or sample_id.startswith(f"{mid}_")
for mid in moved_sample_ids
)
def rewrite_split_path(filepath: str) -> str:
parts = Path(filepath).parts
if not parts or parts[0] not in dir_map:
return filepath
return str(Path(dir_map[parts[0]], *parts[1:]))
mask = df["SampleID"].map(belongs_to_moved_case)
df.loc[mask, "filepath"] = df.loc[mask, "filepath"].map(rewrite_split_path)
if "SplitType" in df.columns:
df.loc[mask, "SplitType"] = "Ts"
df.to_csv(index_file, index=False)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/imgtools/io/nnunet_output.py` around lines 316 - 341, The split index
patching logic in the helper that reads self.writer.index_file should fail fast
instead of silently returning when the index file is missing after
split_dataset() has moved files. Update the guard in this path to raise an
exception with clear context about the missing index and the moved_sample_ids
involved, so callers can detect the inconsistency rather than continuing with
stale metadata.

@JoshuaSiraj
JoshuaSiraj requested a review from skim2257 June 30, 2026 16:11
@JoshuaSiraj
JoshuaSiraj merged commit 130c271 into main Jun 30, 2026
111 of 127 checks passed
@JoshuaSiraj
JoshuaSiraj deleted the JoshuaSiraj/nnunet_updates branch June 30, 2026 18:14
@JoshuaSiraj JoshuaSiraj linked an issue Jun 30, 2026 that may be closed by this pull request
This was referenced Jun 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

nnUNetPipeline naming convention

3 participants