Do LLMs Understand Context? A Knowledge Graph–Based Evaluation Framework (S3KG)
A new arXiv paper (arXiv:2609.30484v1) proposes a knowledge-graph-based evaluation framework for LLM contextual understanding in question answering centered on S3KG (Semantic Structural Similarity for KGs), a hybrid metric that combines structural and semantic signals. The work also introduces a triplet-level diagnostic analysis to categorize reasoning errors; across nine benchmarks S3KG reports up to +7.6 F1 points over the strongest baseline and AUROC up to 0.973.