Canonical dataset contract

Maturity: verified

One canonical bundle has this layout:

<bundle-root>/
  manifest.json
  metadata.parquet
  signals/
    train/part-00000-of-000NN.array_record
    val/part-00000-of-000NN.array_record
    test/part-00000-of-000NN.array_record

Signals

Each ArrayRecord record is one contiguous uncompressed little-endian float32 I/Q array shaped (2, T). Writers use group_size:1,uncompressed; ArrayRecord’s index provides direct access by the record index local to that split.

Metadata

metadata.parquet has one row per signal record. Common columns are sample_id, source_record_id, label, and snr_db. Rows are grouped in train, validation, and test order. Dataset-specific columns become sample metadata attributes.

The bundle reads the table once and slices it using the manifest’s split counts. Within each slice, row position directly addresses the corresponding split-owned signal record.

Manifest

The manifest identifies the dataset snapshot, source provenance, signal dtype and shape, physical signal facts, split policy and counts, and checksums for every generated artifact.

Reader

CanonicalRFBundle shares the manifest and metadata across split datasets. Production pretraining exposes split-local scalar records through CanonicalGrainSource; Grain owns sampling, concurrent reads, buffering, and batching. CanonicalRFDataset remains the map-style PyTorch reader for the completed post-training path and standalone experiments.

See Also