RESEARCH · RESEARCH · #1410
Build2SPARQL: large-scale text-to-SPARQL dataset for building knowledge graphs
The paper introduces Build2SPARQL, a benchmark produced by a KG-grounded pipeline that auto-generates 6,136 executable SPARQL queries and 30,680 natural-language questions over 201 building KGs (180 Brick, 21 ASHRAE 223P). Queries are produced and validated by graph-traversal code while LLMs only generate the natural-language paraphrases; human validation of 300 questions reports 98.8% semantic fidelity and naturalness and 84.0% operational plausibility, and retrieval-augmented evaluation boosted exact-match accuracy on three open-weight models from 0.2–20% (zero-shot) to 56–65% (three-shot retrieved).
KEY POINTS
- The paper introduces Build2SPARQL, a benchmark produced by a KG-grounded pipeline that auto-generates 6,136 executable SPARQL queries and 30,680 natural-language questions over 201 building KGs (180 Brick, 21 ASHRAE 223P).
- Queries are produced and validated by graph-traversal code while LLMs only generate the natural-language paraphrases; human validation of 300 questions reports 98.8% semantic fidelity and naturalness and 84.0% operational plausibility, and retrieval-augmented evaluation boosted exact-match accuracy on three open-weight models from 0.2–20% (zero-shot) to 56–65% (three-shot retrieved).
- Provides a large, KG-grounded benchmark that fixes query correctness independently of model generation and materially advances evaluation and development of text-to-SPARQL systems for building-domain knowledge graphs.
WHY IT MATTERS
Provides a large, KG-grounded benchmark that fixes query correctness independently of model generation and materially advances evaluation and development of text-to-SPARQL systems for building-domain knowledge graphs.