Canonical dataset contract¶
Maturity: verified
One canonical bundle has this layout:
<bundle-root>/
manifest.json
metadata.parquet
signals/
train/part-00000-of-000NN.array_record
val/part-00000-of-000NN.array_record
test/part-00000-of-000NN.array_record
Signals¶
Each ArrayRecord record is one contiguous uncompressed little-endian
float32 I/Q array shaped (2, T). Writers use
group_size:1,uncompressed; ArrayRecord’s index provides direct access by the
record index local to that split.
Metadata¶
metadata.parquet has one row per signal record. Common columns are
sample_id, source_record_id, label, and snr_db. Rows are grouped in
train, validation, and test order. Dataset-specific columns become sample
metadata attributes.
The bundle reads the table once and slices it using the manifest’s split counts. Within each slice, row position directly addresses the corresponding split-owned signal record.
Manifest¶
The manifest identifies the dataset snapshot, source provenance, signal dtype and shape, physical signal facts, split policy and counts, and checksums for every generated artifact.
Reader¶
CanonicalRFBundle shares the manifest and metadata across split datasets.
Production pretraining exposes split-local scalar records through
CanonicalGrainSource; Grain owns sampling, concurrent reads, buffering, and
batching. CanonicalRFDataset remains the map-style PyTorch reader for the
completed post-training path and standalone experiments.