Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

MODEL · ENTITY #6542

Claude Opus 4.6

Related event timeline, sources and context from the news index.

EVENT TIMELINE

2

RESEARCH · 1 SOURCE · arXiv cs.AI

IntLawNER: token-level NER dataset and benchmark for international law (arXiv)

The paper introduces IntLawNER, a token-level NER dataset and benchmark for codified international law, containing 2,987 gold-annotated sentences and 8,094 entity spans from ICJ decisions, UN Security Council resolutions, and ECtHR judgments annotated with seven institution-specific entity types. The authors describe a hybrid pipeline (candidate retrieval, LLM vetting, human review) that reduced 468k source sentences to the final set, analyse silver-to-gold annotation mismatches, and benchmark models—finding zero-shot span-based GLiNER performs poorly on institution-function labels (0.243 micro-F1) while few-shot prompting substantially improves LLMs, with Claude Opus 4.6 reaching 0.873 micro-F1."

6.0

RESEARCH · 1 SOURCE · Hugging Face

UK AISI publishes verified benchmark results via EvalEval's Evaluation Cards

The UK AI Security Institute (AISI) is using the EvalEval Coalition's Every Eval Ever schema and Evaluation Cards platform to publicly share verified evaluation runs tied to its Terminal-Bench 2.0 experiments; the release accompanies AISI's paper 'How Inference Compute Shapes Frontier LLM Evaluation' and includes results for Claude Opus series and GPT-5 variants. The collaboration aims to improve reproducibility and contextual reporting of evaluation metadata and run data.

7.0