Tech Meridian ← LIVE FEED
RU

NEWS · RESEARCH · #318

Croissant 1.0: a schema.org-based metadata format for ML-ready datasets, with support from major dataset hosts and frameworks

Croissant is a new ML-oriented metadata format (v1.0) built on schema.org that standardizes description and organization of datasets without changing underlying file formats. The release includes a spec, example datasets, an open-source Python validator/consumer/generator, a visual editor, and a Responsible AI (RAI) vocabulary extension, and is being adopted by Kaggle, Hugging Face, OpenML, indexed by Dataset Search, and made loadable in major frameworks via TensorFlow Datasets.

KEY POINTS

  1. Croissant is a new ML-oriented metadata format (v1.0) built on schema.org that standardizes description and organization of datasets without changing underlying file formats.
  2. The release includes a spec, example datasets, an open-source Python validator/consumer/generator, a visual editor, and a Responsible AI (RAI) vocabulary extension, and is being adopted by Kaggle, Hugging Face, OpenML, indexed by Dataset Search, and made loadable in major frameworks via TensorFlow Datasets.
  3. A shared ML-focused metadata format can reduce the "data development burden" by improving dataset discoverability, tooling, reuse, and tracking of Responsible AI attributes across repositories and frameworks.

WHY IT MATTERS

A shared ML-focused metadata format can reduce the "data development burden" by improving dataset discoverability, tooling, reuse, and tracking of Responsible AI attributes across repositories and frameworks.

SOURCES & TIMELINE

1