NEWS · RESEARCH · #318
Croissant 1.0: a schema.org-based metadata format for ML-ready datasets, with support from major dataset hosts and frameworks
Croissant is a new ML-oriented metadata format (v1.0) built on schema.org that standardizes description and organization of datasets without changing underlying file formats. The release includes a spec, example datasets, an open-source Python validator/consumer/generator, a visual editor, and a Responsible AI (RAI) vocabulary extension, and is being adopted by Kaggle, Hugging Face, OpenML, indexed by Dataset Search, and made loadable in major frameworks via TensorFlow Datasets.
KEY POINTS
- Croissant is a new ML-oriented metadata format (v1.0) built on schema.org that standardizes description and organization of datasets without changing underlying file formats.
- The release includes a spec, example datasets, an open-source Python validator/consumer/generator, a visual editor, and a Responsible AI (RAI) vocabulary extension, and is being adopted by Kaggle, Hugging Face, OpenML, indexed by Dataset Search, and made loadable in major frameworks via TensorFlow Datasets.
- A shared ML-focused metadata format can reduce the "data development burden" by improving dataset discoverability, tooling, reuse, and tracking of Responsible AI attributes across repositories and frameworks.
WHY IT MATTERS
A shared ML-focused metadata format can reduce the "data development burden" by improving dataset discoverability, tooling, reuse, and tracking of Responsible AI attributes across repositories and frameworks.