IntLawNER: token-level NER dataset and benchmark for international law (arXiv)
The paper introduces IntLawNER, a token-level NER dataset and benchmark for codified international law, containing 2,987 gold-annotated sentences and 8,094 entity spans from ICJ decisions, UN Security Council resolutions, and ECtHR judgments annotated with seven institution-specific entity types. The authors describe a hybrid pipeline (candidate retrieval, LLM vetting, human review) that reduced 468k source sentences to the final set, analyse silver-to-gold annotation mismatches, and benchmark models—finding zero-shot span-based GLiNER performs poorly on institution-function labels (0.243 micro-F1) while few-shot prompting substantially improves LLMs, with Claude Opus 4.6 reaching 0.873 micro-F1."