Canonical datasets¶
Maturity: verified
Training reads generated canonical bundles, never dataset-specific raw formats.
Bundle flow¶
raw source
-> typed dataset adapter
-> common ArrayRecord and metadata writer
-> manifest + ArrayRecord shards + row-aligned metadata
-> one shared manifest and metadata bundle
-> split-owned train and validation Grain sources
-> Grain sampling, batching, prefetch, and iterator state
-> typed CPU batch
ArrayRecord supplies indexed random access to individual uncompressed signal records. Metadata stores each split assignment and split-local record index. Shuffled batches do not scan shards or load neighboring samples. Train and validation share manifest and metadata identity while each opens only its own ArrayRecord shard set.
Dataset readiness¶
present means raw source material exists. A dataset becomes training-ready
only after its ArrayRecord bundle is generated and qualified through the same
production Grain pipeline used by training.
The pretraining application resolves the registry’s canonical location to a
path visible to the process. Scheme-less paths work directly; a gs:// path
requires an executor-supplied mounted bucket namespace.