Paper 'Smart Content Ingestion' presents a production-ready content-extraction system for generative AI workloads
This arXiv paper (arXiv:2610.07091v1) describes a production-ready content-extraction pipeline for enterprise generative-AI workloads that treats content extraction as an explicit, measurable lifecycle stage. The system combines selective OCR routing, a scarcity-first curation engine, a reference-based extraction scorer, a deterministic structure-aware parent-child chunker, and a read-only retrieval evaluator; reported results on a 180-document corpus include a top extractor score of 97.4/100 (CER 0.13%, table similarity 0.995) and chunker retrieval metrics hit@1 68.6%, hit@10 92.8%, MRR 0.77 over 25,050 generated questions.