Add bounded Parquet output and an MLCommons Croissant sidecar after canonical training-row validation. JSONL and the Sediment manifest remain canonical.
Preserve nested messages, omitted values, source evidence, and recipe semantics. Decision: define how operators supply the dataset's actual license and rights metadata; the software license must not be assumed to license captured data.
Acceptance criteria:
Relevant files:
packages/export/sediment_export/schema_contracts.py
packages/export/sediment_export/jsonl.py
packages/export/sediment_export/bounded_training.py
docs/exports/training-exports.md
Add bounded Parquet output and an MLCommons Croissant sidecar after canonical training-row validation. JSONL and the Sediment manifest remain canonical.
Preserve nested messages, omitted values, source evidence, and recipe semantics. Decision: define how operators supply the dataset's actual license and rights metadata; the software license must not be assumed to license captured data.
Acceptance criteria:
Relevant files:
packages/export/sediment_export/schema_contracts.pypackages/export/sediment_export/jsonl.pypackages/export/sediment_export/bounded_training.pydocs/exports/training-exports.md