Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

TOPIC · ENTITY #11755

arXiv:2609.38340v1

Related event timeline, sources and context from the news index.

EVENT TIMELINE

1

RESEARCH · 1 SOURCE · arXiv cs.AI

CARAT (arXiv:2609.38340v1) benchmark probes whether materials LLMs reason or recite

The new arXiv preprint (arXiv:2609.38340v1) introduces CARAT, a benchmark that fixes question and gold answer across eight matched views and applies controls (answer masking, evidence injection, paired inference, matched fine-tuning) to test whether materials LLMs derive structural relations from input or merely repeat text. Key reported results: a grounded view outperforms formula inputs by 17.3 points; GraphSpace beats a plain periodic graph by 19.3 points (driven by a 1.96-point gain when the plain rendering already contains needed fields and a 46.7-point gain when it omits them); models often quote links without using them (95.6% paired agreement) but matched supervision can push that to 99.8%, while deleting the link can drop performance to 23.4%, below a 27.0% shortcut baseline.

6.0