RESEARCH · RESEARCH · #1340
CARAT (arXiv:2609.38340v1) benchmark probes whether materials LLMs reason or recite
The new arXiv preprint (arXiv:2609.38340v1) introduces CARAT, a benchmark that fixes question and gold answer across eight matched views and applies controls (answer masking, evidence injection, paired inference, matched fine-tuning) to test whether materials LLMs derive structural relations from input or merely repeat text. Key reported results: a grounded view outperforms formula inputs by 17.3 points; GraphSpace beats a plain periodic graph by 19.3 points (driven by a 1.96-point gain when the plain rendering already contains needed fields and a 46.7-point gain when it omits them); models often quote links without using them (95.6% paired agreement) but matched supervision can push that to 99.8%, while deleting the link can drop performance to 23.4%, below a 27.0% shortcut baseline.
KEY POINTS
- The new arXiv preprint (arXiv:2609.38340v1) introduces CARAT, a benchmark that fixes question and gold answer across eight matched views and applies controls (answer masking, evidence injection, paired inference, matched fine-tuning) to test whether materials LLMs derive structural relations from input or merely repeat text.
- Key reported results: a grounded view outperforms formula inputs by 17.3 points; GraphSpace beats a plain periodic graph by 19.3 points (driven by a 1.96-point gain when the plain rendering already contains needed fields and a 46.7-point gain when it omits them); models often quote links without using them (95.6% paired agreement) but matched supervision can push that to 99.8%, while deleting the link can drop performance to 23.4%, below a 27.0% shortcut baseline.
- This matters because CARAT exposes how domain LLMs can appear accurate by reciting input or copied links rather than grounding answers, stressing the need for controlled benchmarks and training/regression tests in scientific applications.
WHY IT MATTERS
This matters because CARAT exposes how domain LLMs can appear accurate by reciting input or copied links rather than grounding answers, stressing the need for controlled benchmarks and training/regression tests in scientific applications.