IDEA Prune: An integrated enlarge-and-prune pipeline for generative language model pretraining
The paper advocates incorporating enlarged-model pretraining into structured pruning pipelines and treats the enlarge-and-prune process as a single integrated system. It studies whether pretraining a larger model is worthwhile even if the larger model is never deployed and how to optimize the pipeline for token efficiency.