fix(training): read UTF-8 BOM CSV metadata - #1335
YaoxinHuang wants to merge 2 commits into
Conversation
This file contains regression tests for UTF-8 CSV metadata used by training dataset scans, including tests for metadata application, BOM preservation, and handling of missing file headers.
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (2)
Included review availability: Your plan provides up to 10 included reviews per hour; 9 remain after this review. 📝 WalkthroughWalkthroughCSV metadata loading now opens files with ChangesCSV metadata loading
Priority: ⬇️ Low Estimated code review effort: 2 (Simple) | ~10 minutes Change: Bug fix Merge Risk: ⚪ Minimal · up to The CSV metadata change is mergeable after normal checks; no actionable risk is established. 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. A rabbit reads a CSV with care, Comment |
A UTF-8 BOM on the first
Fileheader makes dataset scanning silently ignore the entire CSV, losing its BPM, key and caption annotations. Decode withutf-8-sigso both plain UTF-8 and BOM-prefixed exports populate the existing metadata fields.Added regressions through
DatasetBuilder.scan_directoryusing real WAV/CSV files: comma, semicolon and tab separators, Unicode filenames/captions, a BOM inside caption text, and the missing-File-column control. The BOM cases fail before this change; all 3 new test methods and the combined 17-test dataset/path suite pass. New-test Ruff, compilation and diff checks pass; the CSV module retains 8 pre-existing Ruff diagnostics.Scope is CSV decoding; non-target hardware/runtime paths are unchanged. No model training, GPU tests or full repository suite were run. Prepared and tested with Codex.
Summary by CodeRabbit
Fileheader remain handled appropriately.