RESEARCH · RESEARCH · #602
QVAC Genesis III: 191.43B-token synthetic STEM corpus for efficient LM pre-training
An arXiv preprint introduces QVAC Genesis III, a 191.43B-token, multi-domain synthetic STEM corpus covering 19 domains and multiple difficulty/educational styles; it is generated via a dual teacher‑distillation strategy that uses a weak edge-scale student model to produce corrective explanations and contrastive option-level reasoning, and includes an LLM-as-a-parser evaluation protocol. Controlled from‑scratch experiments with 1.7B-parameter models show consistent improvements over the open synthetic corpus Cosmopedia-v2 and the public Cosmo-1B model on ARC, GPQA Diamond and MMLU STEM benchmarks (up to +28.57% on ARC-E and +21.35% on ARC-C), with Valid Answer Rates up to 99.45%.
KEY POINTS
- An arXiv preprint introduces QVAC Genesis III, a 191.43B-token, multi-domain synthetic STEM corpus covering 19 domains and multiple difficulty/educational styles; it is generated via a dual teacher‑distillation strategy that uses a weak edge-scale student model to produce corrective explanations and contrastive option-level reasoning, and includes an LLM-as-a-parser evaluation protocol.
- Controlled from‑scratch experiments with 1.7B-parameter models show consistent improvements over the open synthetic corpus Cosmopedia-v2 and the public Cosmo-1B model on ARC, GPQA Diamond and MMLU STEM benchmarks (up to +28.57% on ARC-E and +21.35% on ARC-C), with Valid Answer Rates up to 99.45%.
- A large, STEM-focused synthetic pretraining corpus designed for sample-efficient learning on small/edge models could materially improve the quality of educational and on-device LMs and broaden open alternatives to private proprietary corpora.
WHY IT MATTERS
A large, STEM-focused synthetic pretraining corpus designed for sample-efficient learning on small/edge models could materially improve the quality of educational and on-device LMs and broaden open alternatives to private proprietary corpora.