Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1340

CARAT (arXiv:2609.38340v1) benchmark probes whether materials LLMs reason or recite

The new arXiv preprint (arXiv:2609.38340v1) introduces CARAT, a benchmark that fixes question and gold answer across eight matched views and applies controls (answer masking, evidence injection, paired inference, matched fine-tuning) to test whether materials LLMs derive structural relations from input or merely repeat text. Key reported results: a grounded view outperforms formula inputs by 17.3 points; GraphSpace beats a plain periodic graph by 19.3 points (driven by a 1.96-point gain when the plain rendering already contains needed fields and a 46.7-point gain when it omits them); models often quote links without using them (95.6% paired agreement) but matched supervision can push that to 99.8%, while deleting the link can drop performance to 23.4%, below a 27.0% shortcut baseline.

KEY POINTS

  1. The new arXiv preprint (arXiv:2609.38340v1) introduces CARAT, a benchmark that fixes question and gold answer across eight matched views and applies controls (answer masking, evidence injection, paired inference, matched fine-tuning) to test whether materials LLMs derive structural relations from input or merely repeat text.
  2. Key reported results: a grounded view outperforms formula inputs by 17.3 points; GraphSpace beats a plain periodic graph by 19.3 points (driven by a 1.96-point gain when the plain rendering already contains needed fields and a 46.7-point gain when it omits them); models often quote links without using them (95.6% paired agreement) but matched supervision can push that to 99.8%, while deleting the link can drop performance to 23.4%, below a 27.0% shortcut baseline.
  3. This matters because CARAT exposes how domain LLMs can appear accurate by reciting input or copied links rather than grounding answers, stressing the need for controlled benchmarks and training/regression tests in scientific applications.

WHY IT MATTERS

This matters because CARAT exposes how domain LLMs can appear accurate by reciting input or copied links rather than grounding answers, stressing the need for controlled benchmarks and training/regression tests in scientific applications.

SOURCES & TIMELINE

1