test: add nnunet pipeline tests - #395
Conversation
|
Warning Review limit reached
Next review available in: 46 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Pro Run ID: ⛔ Files ignored due to path filters (1)
📒 Files selected for processing (2)
📝 WalkthroughWalkthrough
ChangesnnUNet train/test split and CLI coverage
Estimated code review effort🎯 4 (Complex) | ⏱️ ~45 minutes 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Added testing with RADCURE, only issue is that nnunet is not available for py13. |
|
The issue with env is nnunet install on py313(osx-64) |
jjjermiah
left a comment
There was a problem hiding this comment.
looks great to me, to address the py313 issue, I recommend having a check in the cli entry point that lets the user know
Thanks! How do I deal with the dependency issue? I need to add nnunetv2 to test as a dependency, instead should I add it only to py311 and py312? |
There was a problem hiding this comment.
Actionable comments posted: 1
♻️ Duplicate comments (1)
tests/integration/cli/nnunet_pipeline_cli.py (1)
96-96: Consider implementing pytest-snapshot for consistent output validation.As suggested in previous reviews, using pytest-snapshot would help ensure the generated outputs remain consistent across test runs and catch any unintended changes in the pipeline output format.
This would involve capturing and comparing the generated dataset structure, JSON configurations, and file contents against known good snapshots.
🧹 Nitpick comments (2)
tests/integration/cli/nnunet_pipeline_cli.py (2)
12-13: Fix class naming and docstring accuracy.The class name should follow PEP 8 conventions, and the docstring mentions the wrong CLI command.
-class TestnnUNetCLI: - """Integration tests for the autopipeline CLI command using collections from the test data.""" +class TestNnUNetCLI: + """Integration tests for the nnunet_pipeline CLI command using collections from the test data."""
63-70: Consider moving YAML file creation to a fixture.The temporary YAML file creation could be moved to a fixture for better test organization and reusability.
+ @pytest.fixture(scope="function") + def roi_yaml_file(self, tmp_path): + """Create a temporary ROI mapping YAML file.""" + roi_dict = { + "BRAINSTEM": "Brainstem", + "SPINALCORD": "SpinalCord", + "LARYNX": "Larynx", + } + roi_yaml_path = tmp_path / "roi_match.yaml" + with roi_yaml_path.open("w") as f: + yaml.dump(roi_dict, f) + return roi_yaml_pathThen use
roi_yaml_filefixture in the test method instead of creating the file inline.
📜 Review details
Configuration used: .coderabbit.yaml
Review profile: CHILL
Plan: Pro
⛔ Files ignored due to path filters (2)
pixi.lockis excluded by!**/*.lockand included by nonepixi.tomlis excluded by none and included by none
📒 Files selected for processing (2)
src/imgtools/io/nnunet_output.py(1 hunks)tests/integration/cli/nnunet_pipeline_cli.py(1 hunks)
🧰 Additional context used
📓 Path-based instructions (2)
`src/**/*.py`: Review the Python code for compliance with PEP 8 and PEP 257 (doc...
src/**/*.py: Review the Python code for compliance with PEP 8 and PEP 257 (docstring conventions). Ensure the following: - Variables and functions follow meaningful naming conventions. - Docstrings are present, accurate, and align with the implementation. - Code is efficient and avoids redundancy while adhering to DRY principles. - Consider suggestions to enhance readability and maintainability. - Highlight any potential performance issues, edge cases, or logical errors. - Ensure all imported libraries are used and necessary.
⚙️ Source: CodeRabbit Configuration File
List of files the instruction was applied to:
src/imgtools/io/nnunet_output.py
`tests/**/*`: Review the test code written with Pytest. Confirm: - Tests cover a...
tests/**/*: Review the test code written with Pytest. Confirm: - Tests cover all critical functionality and edge cases. - Test descriptions clearly describe their purpose. - Pytest best practices are followed, such as proper use of fixtures. - Ensure the tests are isolated and do not have external dependencies (e.g., network calls). - Verify meaningful assertions and avoidance of redundant tests. - Test code adheres to PEP 8 style guidelines.
⚙️ Source: CodeRabbit Configuration File
List of files the instruction was applied to:
tests/integration/cli/nnunet_pipeline_cli.py
🧠 Learnings (3)
📓 Common learnings
Learnt from: jjjermiah
PR: bhklab/med-imagetools#145
File: src/imgtools/utils/nnunet.py:0-0
Timestamp: 2024-11-29T21:18:38.153Z
Learning: Suggestions to modify the `save_json` function in `src/imgtools/utils/nnunet.py` to fix type annotations or add error handling are considered out of scope.
src/imgtools/io/nnunet_output.py (1)
Learnt from: jjjermiah
PR: bhklab/med-imagetools#145
File: src/imgtools/utils/nnunet.py:0-0
Timestamp: 2024-11-29T21:18:38.153Z
Learning: Suggestions to modify the `save_json` function in `src/imgtools/utils/nnunet.py` to fix type annotations or add error handling are considered out of scope.
tests/integration/cli/nnunet_pipeline_cli.py (3)
Learnt from: jjjermiah
PR: bhklab/med-imagetools#137
File: src/imgtools/dicom/sort/parser.py:132-137
Timestamp: 2024-11-21T21:03:45.548Z
Learning: Assertions are allowed for input validation throughout the project, including in `src/imgtools/dicom/sort/parser.py`.
Learnt from: jjjermiah
PR: bhklab/med-imagetools#137
File: src/imgtools/dicom/sort/utils.py:66-67
Timestamp: 2024-11-21T16:24:04.091Z
Learning: In the `src/imgtools/dicom/sort/utils.py` file and throughout the codebase, assertions are preferred for input validation instead of explicit type checks or raising exceptions.
Learnt from: jjjermiah
PR: bhklab/med-imagetools#137
File: src/imgtools/dicom/sort/utils.py:168-169
Timestamp: 2024-11-21T16:24:13.275Z
Learning: In `src/imgtools/dicom/sort/utils.py`, it's acceptable to use `assert` statements for input validation in the `read_tags` function.
🧬 Code Graph Analysis (1)
src/imgtools/io/nnunet_output.py (1)
src/imgtools/coretypes/base_masks.py (1)
roi_keys(161-163)
🔇 Additional comments (1)
src/imgtools/io/nnunet_output.py (1)
301-303: LGTM: Preserves background label correctly.The change from reassigning to updating the labels dictionary in-place is a good fix. This ensures the background label (0) is preserved while adding ROI keys with proper incremental indexing, which aligns with nnUNet's expected label format.
| nnunet_result = subprocess.run([ | ||
| "nnUNetv2_extract_fingerprint", | ||
| "-d", "1", | ||
| "--verify_dataset_integrity", | ||
| ], | ||
| env=env, | ||
| stdout=subprocess.PIPE | ||
| ) |
There was a problem hiding this comment.
🛠️ Refactor suggestion
Improve subprocess call robustness.
The subprocess call lacks timeout and proper error handling, which could cause tests to hang or hide important error information.
nnunet_result = subprocess.run([
"nnUNetv2_extract_fingerprint",
"-d", "1",
"--verify_dataset_integrity",
],
env=env,
- stdout=subprocess.PIPE
+ stdout=subprocess.PIPE,
+ stderr=subprocess.PIPE,
+ timeout=300, # 5 minute timeout
+ text=True
)Also consider logging the stderr output when the command fails for better debugging:
- assert nnunet_result.returncode == 0, "nnUNetv2_extract_fingerprint failed"
+ assert nnunet_result.returncode == 0, f"nnUNetv2_extract_fingerprint failed: {nnunet_result.stderr}"📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| nnunet_result = subprocess.run([ | |
| "nnUNetv2_extract_fingerprint", | |
| "-d", "1", | |
| "--verify_dataset_integrity", | |
| ], | |
| env=env, | |
| stdout=subprocess.PIPE | |
| ) | |
| nnunet_result = subprocess.run([ | |
| "nnUNetv2_extract_fingerprint", | |
| "-d", "1", | |
| "--verify_dataset_integrity", | |
| ], | |
| env=env, | |
| stdout=subprocess.PIPE, | |
| stderr=subprocess.PIPE, | |
| timeout=300, # 5 minute timeout | |
| text=True | |
| ) | |
| assert nnunet_result.returncode == 0, f"nnUNetv2_extract_fingerprint failed: {nnunet_result.stderr}" |
🤖 Prompt for AI Agents
In tests/integration/cli/nnunet_pipeline_cli.py around lines 87 to 94, the
subprocess.run call lacks a timeout and error handling, risking hangs and
obscured errors. Add a timeout parameter to prevent indefinite blocking, use
check=True to raise an exception on failure, and catch
subprocess.CalledProcessError to log stderr output for better debugging.
…ameter and remove obsolete CLI tests
There was a problem hiding this comment.
Actionable comments posted: 1
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
tests/integration/cli/test_nnunet_pipeline_cli.py (1)
35-52: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick winTighten these error assertions.
Allowing
or "Error:" in result.outputmeans any Click failure can satisfy the test, even if the wrong validation fired. Assert the specific missing-option message in each case.As per path instructions,
tests/**/*: Verify meaningful assertions and avoidance of redundant tests.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@tests/integration/cli/test_nnunet_pipeline_cli.py` around lines 35 - 52, The invalid-args checks in test_invalid_args are too loose because they accept any Click failure via the generic "Error:" fallback, which can hide the wrong validation path. Tighten the assertions to match the exact missing-option messages for each runner.invoke call in test_nnunet_pipeline_cli, using the nnunet_pipeline CLI behavior as the source of truth. Keep the checks specific to the expected missing --modalities and --roi-match-yaml/-ryaml failures so the test only passes when the intended validation fires.Source: Path instructions
🧹 Nitpick comments (1)
src/imgtools/nnunet_pipeline.py (1)
48-48: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winDocument the new
crawl_directoryparameter.
__init__now exposescrawl_directory, but the constructor docstring still omits it, so the public API docs no longer match the implementation.Proposed update
mask_saving_strategy : MaskSavingStrategy Mask saving strategy + crawl_directory : str | Path | None, optional + Directory used for crawl metadata. If omitted, the default crawl + location is used. update_crawl : bool, optional Whether to force recrawling, by default FalseAs per path instructions,
src/**/*.py: Review the Python code for compliance with PEP 8 and PEP 257 (docstring conventions).🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/imgtools/nnunet_pipeline.py` at line 48, Update the __init__ docstring in nnunet_pipeline so it documents the new crawl_directory parameter alongside the other constructor args. Add a concise parameter description that matches its type and behavior, and keep the docstring aligned with PEP 257 conventions by using the same style as the existing constructor documentation.Source: Path instructions
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/integration/cli/test_nnunet_pipeline_cli.py`:
- Around line 53-55: Gate test_nnunet_pipeline on nnUNet availability before
invoking nnUNetv2_extract_fingerprint so it skips cleanly when the binary is
missing, and keep the subprocess-based integration path isolated. Also tighten
the existing macOS skip in test_nnunet_pipeline so it only excludes the actual
Python 3.13 nnUNet incompatibility rather than all Darwin runs; use the
test_nnunet_pipeline method and the nnUNetv2_extract_fingerprint call site as
the key places to update.
---
Outside diff comments:
In `@tests/integration/cli/test_nnunet_pipeline_cli.py`:
- Around line 35-52: The invalid-args checks in test_invalid_args are too loose
because they accept any Click failure via the generic "Error:" fallback, which
can hide the wrong validation path. Tighten the assertions to match the exact
missing-option messages for each runner.invoke call in test_nnunet_pipeline_cli,
using the nnunet_pipeline CLI behavior as the source of truth. Keep the checks
specific to the expected missing --modalities and --roi-match-yaml/-ryaml
failures so the test only passes when the intended validation fires.
---
Nitpick comments:
In `@src/imgtools/nnunet_pipeline.py`:
- Line 48: Update the __init__ docstring in nnunet_pipeline so it documents the
new crawl_directory parameter alongside the other constructor args. Add a
concise parameter description that matches its type and behavior, and keep the
docstring aligned with PEP 257 conventions by using the same style as the
existing constructor documentation.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Pro
Run ID: 640c32ef-8ff9-4339-9eb4-b2765b1dc241
⛔ Files ignored due to path filters (2)
pixi.lockis excluded by!**/*.lockand included by nonepixi.tomlis excluded by none and included by none
📒 Files selected for processing (2)
src/imgtools/nnunet_pipeline.pytests/integration/cli/test_nnunet_pipeline_cli.py
| @pytest.mark.skipif(platform == "darwin", reason="Test skipped on macOS, due to nnUNet py313 incompatibility") | ||
| @pytest.mark.parametrize("mask_saving_strategy", ["sparse_mask", "region_mask"]) | ||
| def test_nnunet_pipeline(self, runner, temp_output_dir, DATA_DIR, mask_saving_strategy): |
There was a problem hiding this comment.
🩺 Stability & Availability | 🟠 Major | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
set -euo pipefail
printf '\n== target file ==\n'
wc -l tests/integration/cli/test_nnunet_pipeline_cli.py
sed -n '1,220p' tests/integration/cli/test_nnunet_pipeline_cli.py
printf '\n== nnUNet references ==\n'
rg -n "nnUNetv2_extract_fingerprint|skipif|shutil.which|platform == \"darwin\"|py313|nnUNet" tests . -g '!**/.git/**'Repository: bhklab/med-imagetools
Length of output: 12345
Gate this test on nnUNet availability, and narrow the macOS skip to the actual py3.13 incompatibility.
This test still unconditionally runs nnUNetv2_extract_fingerprint, so environments without that binary will fail instead of skipping cleanly. A shutil.which(...) check before the subprocess call would keep the integration test isolated.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/integration/cli/test_nnunet_pipeline_cli.py` around lines 53 - 55, Gate
test_nnunet_pipeline on nnUNet availability before invoking
nnUNetv2_extract_fingerprint so it skips cleanly when the binary is missing, and
keep the subprocess-based integration path isolated. Also tighten the existing
macOS skip in test_nnunet_pipeline so it only excludes the actual Python 3.13
nnUNet incompatibility rather than all Darwin runs; use the test_nnunet_pipeline
method and the nnUNetv2_extract_fingerprint call site as the key places to
update.
Source: Path instructions
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #395 +/- ##
===========================================
+ Coverage 54.85% 65.90% +11.05%
===========================================
Files 66 66
Lines 4370 4420 +50
===========================================
+ Hits 2397 2913 +516
+ Misses 1973 1507 -466 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Hi sorry issue took a little while, but I wrote up a minimal way of implementing a test_set_split for nnunet-pipeline. The feature takes the successfully processed images and moves a proportion of the images/labels to the appropriate dirs in the output dir. I could not find any tests for nnunet-pipeline but this should probably be added to them when they're ready. There is CLI support for the feature as well as hints. Please let me know if there's anything II should change style wise. Fixes #443 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a `test-set-ratio` option to the command-line workflow, letting you split a portion of successfully processed data into a test set. * The pipeline now automatically moves the selected files into the appropriate test folders after processing. * **Bug Fixes** * Added validation for the new ratio value to keep it within a valid range and prevent invalid dataset splits. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: kojukwu <kojukwu@gmail.com>
This update introduces a `test-set-ratio` and `random-seed` option to the nnUNet pipeline, allowing users to specify the proportion of successfully processed cases to be moved to the test set. The implementation includes validation for the ratio and updates to the dataset splitting logic. Additionally, the CLI has been updated to reflect these new options, ensuring a smoother user experience. Tests have been added to verify the correct functionality of the new features. Fixes #443
There was a problem hiding this comment.
Actionable comments posted: 2
🧹 Nitpick comments (4)
src/imgtools/nnunet_pipeline.py (2)
49-59: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winKeep the constructor docstring aligned with the new signature.
crawl_directoryis now optional but still undocumented, and the new split parameters omit their defaults. This makes the public API harder to use correctly.Proposed docstring update
+ crawl_directory : str | Path | None, optional + Directory for cached crawl data, by default None. ... - test_set_ratio : float + test_set_ratio : float, default=0.0 Proportion of successful cases for the test set. Count is ceil(ratio * n_cases); 1.0 moves all cases to the test set. - random_seed : int + random_seed : int, default=42 The random seed to use for the test set split.As per path instructions,
src/**/*.py: “Docstrings are present, accurate, and align with the implementation.”Also applies to: 92-96
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/imgtools/nnunet_pipeline.py` around lines 49 - 59, The constructor docstring for the nnUNet pipeline class is out of sync with the updated signature: it must document the new optional crawl_directory parameter and include the default values for the new split-related arguments. Update the docstring where the initializer is described so it matches the public API exposed by the constructor symbols, especially crawl_directory, existing_file_mode, update_crawl, n_jobs, roi_ignore_case, roi_allow_multi_key_matches, spacing, window, level, test_set_ratio, and random_seed.Source: Path instructions
159-159: 📐 Maintainability & Code Quality | 🔵 TrivialTrack this broad TODO outside the implementation.
The note is valid, but it is broad enough to become stale in code. Consider turning it into an issue or a more specific TODO. I can help draft the refactor ticket if useful.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/imgtools/nnunet_pipeline.py` at line 159, The TODO comment in the long nnunet_pipeline function is too broad to keep in the implementation and will likely go stale. Remove the inline “This function is long and has a lot of concerns built into it” note and replace it with a more specific, actionable TODO tied to the exact refactor needed, or move the broader refactor tracking to an external issue/ticket. Use the surrounding pipeline function context in src/imgtools/nnunet_pipeline.py to locate and update the comment.src/imgtools/cli/nnunet_pipeline.py (1)
122-122: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winApply PEP 8 spacing in the new option and parameter.
Tiny cleanup, but it keeps the touched CLI surface consistent with the project’s Python style.
Proposed cleanup
- type=click.FloatRange(0.0,1.0), + type=click.FloatRange(0.0, 1.0), ... - test_set_ratio:float, + test_set_ratio: float,As per path instructions,
src/**/*.py: “Review the Python code for compliance with PEP 8 and PEP 257.”Also applies to: 158-158
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/imgtools/cli/nnunet_pipeline.py` at line 122, The new CLI option and its corresponding parameter in nnunet_pipeline.py need PEP 8 spacing cleanup. Update the affected click option definition and the matching function signature so the decorator arguments and parameter list use standard spacing/style consistently with the rest of the CLI, keeping the change localized to the relevant option block and its associated callable.Source: Path instructions
src/imgtools/io/nnunet_output.py (1)
19-19: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winMake
ProcessSampleResulta type-only import. It’s only used insplit_dataset()annotations, so move it underTYPE_CHECKING(or rely on postponed annotations) to avoid importingimgtools.autopipelineat runtime.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/imgtools/io/nnunet_output.py` at line 19, `ProcessSampleResult` is only referenced in `split_dataset()` type annotations, so it should not be imported at runtime. Update the imports in `nnunet_output` to make `ProcessSampleResult` type-only by moving it under a `TYPE_CHECKING` guard or relying on postponed annotation behavior, and keep the runtime path free of the `imgtools.autopipeline` import.Source: Path instructions
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@src/imgtools/io/nnunet_output.py`:
- Line 201: The nnUNet output filename template in the path-building logic
currently embeds PatientID, which can expose PHI in generated artifacts. Update
the filename construction in the nnunet output helper to use a non-identifying
identifier such as SampleID or another generated ID instead of PatientID, and
make sure any related naming helpers or format strings that reference PatientID
are adjusted consistently.
- Around line 316-341: The split index patching logic in the helper that reads
self.writer.index_file should fail fast instead of silently returning when the
index file is missing after split_dataset() has moved files. Update the guard in
this path to raise an exception with clear context about the missing index and
the moved_sample_ids involved, so callers can detect the inconsistency rather
than continuing with stale metadata.
---
Nitpick comments:
In `@src/imgtools/cli/nnunet_pipeline.py`:
- Line 122: The new CLI option and its corresponding parameter in
nnunet_pipeline.py need PEP 8 spacing cleanup. Update the affected click option
definition and the matching function signature so the decorator arguments and
parameter list use standard spacing/style consistently with the rest of the CLI,
keeping the change localized to the relevant option block and its associated
callable.
In `@src/imgtools/io/nnunet_output.py`:
- Line 19: `ProcessSampleResult` is only referenced in `split_dataset()` type
annotations, so it should not be imported at runtime. Update the imports in
`nnunet_output` to make `ProcessSampleResult` type-only by moving it under a
`TYPE_CHECKING` guard or relying on postponed annotation behavior, and keep the
runtime path free of the `imgtools.autopipeline` import.
In `@src/imgtools/nnunet_pipeline.py`:
- Around line 49-59: The constructor docstring for the nnUNet pipeline class is
out of sync with the updated signature: it must document the new optional
crawl_directory parameter and include the default values for the new
split-related arguments. Update the docstring where the initializer is described
so it matches the public API exposed by the constructor symbols, especially
crawl_directory, existing_file_mode, update_crawl, n_jobs, roi_ignore_case,
roi_allow_multi_key_matches, spacing, window, level, test_set_ratio, and
random_seed.
- Line 159: The TODO comment in the long nnunet_pipeline function is too broad
to keep in the implementation and will likely go stale. Remove the inline “This
function is long and has a lot of concerns built into it” note and replace it
with a more specific, actionable TODO tied to the exact refactor needed, or move
the broader refactor tracking to an external issue/ticket. Use the surrounding
pipeline function context in src/imgtools/nnunet_pipeline.py to locate and
update the comment.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Pro
Run ID: f91f4f82-9b4e-405d-af3f-a5d49944b167
⛔ Files ignored due to path filters (1)
pixi.tomlis excluded by none and included by none
📒 Files selected for processing (4)
src/imgtools/cli/nnunet_pipeline.pysrc/imgtools/io/nnunet_output.pysrc/imgtools/nnunet_pipeline.pytests/integration/cli/test_nnunet_pipeline_cli.py
🚧 Files skipped from review as they are similar to previous changes (1)
- tests/integration/cli/test_nnunet_pipeline_cli.py
|
|
||
| self._file_name_format = ( | ||
| "{DirType}{SplitType}/{Dataset}_{SampleID}.nii.gz" | ||
| "{DirType}{SplitType}/{PatientID}_{SampleID}.nii.gz" |
There was a problem hiding this comment.
🔒 Security & Privacy | 🟠 Major | ⚡ Quick win
Avoid embedding PatientID in generated filenames.
PatientID is patient-identifying metadata; putting it in every nnUNet filename can leak PHI through paths, logs, reports, and shared artifacts. Prefer the generated sample ID or another non-PHI identifier.
Proposed safer template
- "{DirType}{SplitType}/{PatientID}_{SampleID}.nii.gz"
+ "{DirType}{SplitType}/{SampleID}.nii.gz"📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| "{DirType}{SplitType}/{PatientID}_{SampleID}.nii.gz" | |
| "{DirType}{SplitType}/{SampleID}.nii.gz" |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/imgtools/io/nnunet_output.py` at line 201, The nnUNet output filename
template in the path-building logic currently embeds PatientID, which can expose
PHI in generated artifacts. Update the filename construction in the nnunet
output helper to use a non-identifying identifier such as SampleID or another
generated ID instead of PatientID, and make sure any related naming helpers or
format strings that reference PatientID are adjusted consistently.
| index_file = self.writer.index_file | ||
| if not moved_sample_ids or not index_file.exists(): | ||
| return | ||
|
|
||
| dir_map = {"imagesTr": "imagesTs", "labelsTr": "labelsTs"} | ||
| df = pd.read_csv(index_file, dtype={"SampleID": str}) | ||
|
|
||
| def belongs_to_moved_case(sample_id: str) -> bool: | ||
| sample_id = str(sample_id) | ||
| return any( | ||
| sample_id == mid or sample_id.startswith(f"{mid}_") | ||
| for mid in moved_sample_ids | ||
| ) | ||
|
|
||
| def rewrite_split_path(filepath: str) -> str: | ||
| parts = Path(filepath).parts | ||
| if not parts or parts[0] not in dir_map: | ||
| return filepath | ||
| return str(Path(dir_map[parts[0]], *parts[1:])) | ||
|
|
||
| mask = df["SampleID"].map(belongs_to_moved_case) | ||
| df.loc[mask, "filepath"] = df.loc[mask, "filepath"].map(rewrite_split_path) | ||
| if "SplitType" in df.columns: | ||
| df.loc[mask, "SplitType"] = "Ts" | ||
|
|
||
| df.to_csv(index_file, index=False) |
There was a problem hiding this comment.
🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win
Fail fast when the split index cannot be patched.
split_dataset() moves files before calling this helper; if the index file is missing, Line 317 silently returns and leaves moved files without synchronized index metadata. Raise with context instead.
Proposed guard
index_file = self.writer.index_file
- if not moved_sample_ids or not index_file.exists():
+ if not moved_sample_ids:
return
+ if not index_file.exists():
+ raise FileNotFoundError(
+ f"Cannot patch test split index because {index_file} does not exist."
+ )
+
+ required_columns = {"SampleID", "filepath"}
dir_map = {"imagesTr": "imagesTs", "labelsTr": "labelsTs"}
df = pd.read_csv(index_file, dtype={"SampleID": str})
+ missing_columns = required_columns - set(df.columns)
+ if missing_columns:
+ raise ValueError(
+ f"Cannot patch test split index; missing columns: {sorted(missing_columns)}"
+ )📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| index_file = self.writer.index_file | |
| if not moved_sample_ids or not index_file.exists(): | |
| return | |
| dir_map = {"imagesTr": "imagesTs", "labelsTr": "labelsTs"} | |
| df = pd.read_csv(index_file, dtype={"SampleID": str}) | |
| def belongs_to_moved_case(sample_id: str) -> bool: | |
| sample_id = str(sample_id) | |
| return any( | |
| sample_id == mid or sample_id.startswith(f"{mid}_") | |
| for mid in moved_sample_ids | |
| ) | |
| def rewrite_split_path(filepath: str) -> str: | |
| parts = Path(filepath).parts | |
| if not parts or parts[0] not in dir_map: | |
| return filepath | |
| return str(Path(dir_map[parts[0]], *parts[1:])) | |
| mask = df["SampleID"].map(belongs_to_moved_case) | |
| df.loc[mask, "filepath"] = df.loc[mask, "filepath"].map(rewrite_split_path) | |
| if "SplitType" in df.columns: | |
| df.loc[mask, "SplitType"] = "Ts" | |
| df.to_csv(index_file, index=False) | |
| index_file = self.writer.index_file | |
| if not moved_sample_ids: | |
| return | |
| if not index_file.exists(): | |
| raise FileNotFoundError( | |
| f"Cannot patch test split index because {index_file} does not exist." | |
| ) | |
| required_columns = {"SampleID", "filepath"} | |
| dir_map = {"imagesTr": "imagesTs", "labelsTr": "labelsTs"} | |
| df = pd.read_csv(index_file, dtype={"SampleID": str}) | |
| missing_columns = required_columns - set(df.columns) | |
| if missing_columns: | |
| raise ValueError( | |
| f"Cannot patch test split index; missing columns: {sorted(missing_columns)}" | |
| ) | |
| def belongs_to_moved_case(sample_id: str) -> bool: | |
| sample_id = str(sample_id) | |
| return any( | |
| sample_id == mid or sample_id.startswith(f"{mid}_") | |
| for mid in moved_sample_ids | |
| ) | |
| def rewrite_split_path(filepath: str) -> str: | |
| parts = Path(filepath).parts | |
| if not parts or parts[0] not in dir_map: | |
| return filepath | |
| return str(Path(dir_map[parts[0]], *parts[1:])) | |
| mask = df["SampleID"].map(belongs_to_moved_case) | |
| df.loc[mask, "filepath"] = df.loc[mask, "filepath"].map(rewrite_split_path) | |
| if "SplitType" in df.columns: | |
| df.loc[mask, "SplitType"] = "Ts" | |
| df.to_csv(index_file, index=False) |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/imgtools/io/nnunet_output.py` around lines 316 - 341, The split index
patching logic in the helper that reads self.writer.index_file should fail fast
instead of silently returning when the index file is missing after
split_dataset() has moved files. Update the guard in this path to raise an
exception with clear context about the missing index and the moved_sample_ids
involved, so callers can detect the inconsistency rather than continuing with
stale metadata.
Summary by CodeRabbit
New Features
--test-set-ratioand--random-seedCLI options, with deterministic selection and automatic dataset updates.PatientID-based paths.Bug Fixes
Tests
imagesTs/labelsTscounts and fingerprint extraction.