Skip to content

fix(training): read UTF-8 BOM CSV metadata - #1335

Open
YaoxinHuang wants to merge 2 commits into
ace-step:mainfrom
YaoxinHuang:fix/read-bom-training-metadata
Open

YaoxinHuang wants to merge 2 commits into
ace-step:mainfrom
YaoxinHuang:fix/read-bom-training-metadata

Conversation

@YaoxinHuang

@YaoxinHuang YaoxinHuang commented Sep 23, 2026 •

Copy link
Copy Markdown

A UTF-8 BOM on the first File header makes dataset scanning silently ignore the entire CSV, losing its BPM, key and caption annotations. Decode with utf-8-sig so both plain UTF-8 and BOM-prefixed exports populate the existing metadata fields.

Added regressions through DatasetBuilder.scan_directory using real WAV/CSV files: comma, semicolon and tab separators, Unicode filenames/captions, a BOM inside caption text, and the missing-File-column control. The BOM cases fail before this change; all 3 new test methods and the combined 17-test dataset/path suite pass. New-test Ruff, compilation and diff checks pass; the CSV module retains 8 pre-existing Ruff diagnostics.

Scope is CSV decoding; non-target hardware/runtime paths are unchanged. No model training, GPU tests or full repository suite were run. Prepared and tested with Codex.

Summary by CodeRabbit

  • Bug Fixes
    • CSV metadata with a UTF-8 byte order mark is now read correctly, so metadata can be applied from files with or without the mark. Existing caption text and files without a File header remain handled appropriately.

This file contains regression tests for UTF-8 CSV metadata used by training dataset scans, including tests for metadata application, BOM preservation, and handling of missing file headers.
@coderabbitai

coderabbitai Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 12e6bb0c-a007-4483-87ea-255b23a4a5a3

📥 Commits

Reviewing files that changed from the base of the PR and between ca1e85f and fc813c8.

📒 Files selected for processing (2)
  • acestep/training/dataset_builder_modules/csv_metadata.py
  • acestep/training/dataset_builder_modules/csv_metadata_test.py

Included review availability: Your plan provides up to 10 included reviews per hour; 9 remain after this review.


📝 Walkthrough

Walkthrough

CSV metadata loading now opens files with utf-8-sig encoding. Regression tests cover UTF-8 and UTF-8-with-BOM files, multiple delimiters, embedded BOMs in captions, and CSVs without a File header.

Changes

CSV metadata loading

Layer / File(s) Summary
BOM-aware CSV loading
acestep/training/dataset_builder_modules/csv_metadata.py, acestep/training/dataset_builder_modules/csv_metadata_test.py
load_csv_metadata uses utf-8-sig encoding. Tests cover metadata scans with and without a leading BOM, comma, semicolon, and tab delimiters, preservation of an embedded caption BOM, and ignoring CSVs without a File header.

Priority: ⬇️ Low

Estimated code review effort: 2 (Simple) | ~10 minutes

Change: Bug fix

Merge Risk: ⚪ Minimal · up to fc813

The CSV metadata change is mergeable after normal checks; no actionable risk is established.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: fixing CSV metadata decoding for UTF-8 files with a byte order mark.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 6 functions across 2 files.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

A rabbit reads a CSV with care,
A leading BOM is cleared from there.
Commas, tabs, and semicolons align,
While caption BOMs stay in their line.
No File header? The scan lets it pass,
Then hops along through the meadow grass.

Comment @coderabbitai help to get the list of available commands.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant