Tech Meridian ← ENTITY INDEX
RU

MODEL · ENTITY #4284

Llama 3.2 3B

Related event timeline, sources and context from the news index.

EVENT TIMELINE

1

RESEARCH · 1 SOURCE · arXiv cs.AI

Benchmarking LLM Safety for Vehicle Voice Command Authorization

The paper (arXiv:2609.19630v1) introduces a 202-scenario benchmark and seven-class taxonomy to evaluate pre-action authorization decisions for vehicle voice assistants. It evaluates two local open-weight models and three API-based LLMs, reporting decision alignment from 40.1% (Llama 3.2 3B) to 89.1% (Gemini 3.1 Pro Preview), finds API models at 83.2–89.1% with no significant differences, observes two to three False Executes among 161 non-execution scenarios, and concludes that structured LLM decisions alone are insufficient—deployments require an independent enforcement layer to verify permissions and vehicle-state constraints.

7.0