JusticeAxis: multimodal legal-judgment benchmark and JusticeAgent harness (arXiv:2610.00353v1)
The paper introduces JusticeAxis, a multimodal benchmark of 256 real-world criminal cases from 18 countries that includes audio, image, and text evidence plus three lawyer-written judgments per case, and proposes JusticeAgent, a modular harness where fact-establishing agents feed a judge agent that applies statutes under skills distilled from execution trajectories and admitted via Bayesian credible bounds. Experiments on arXiv:2610.00353v1 report that model failure modes shift with scale (open-weight backbones drift to unsupported grounds; frontier models revert to statutory defaults) and that JusticeAgent can act as a plugin to bring a frozen open-weight backbone to commercial-level performance; project resources are available at the linked GitHub repository.