Tech Meridian ← ENTITY INDEX
RU

TOPIC · ENTITY #1407

benchmarks

Related event timeline, sources and context from the news index.

EVENT TIMELINE

2

RESEARCH · 1 SOURCE · arXiv cs.AI

little m: an AI agent to formulate industrial process optimization models

Researchers released 'little m', an AI agent that combines a domain-specific knowledge repository with LLM-driven interaction to translate messy, multimodal industrial specifications (text and process diagrams) into mathematical optimization models. They also introduced IPC-Bench, a 50-scenario multimodal benchmark for industrial process control; evaluations (automated structural checks and double-blind human review) show little m generates substantially more semantically correct formulations than state-of-the-art LLMs, though the paper does not evaluate solver feasibility, physical validity, or closed-loop performance.

6.0

RESEARCH · 1 SOURCE · Apple Machine Learning Research

Agent Seer: Synthesizing realistic agent evaluation scenarios from tool specifications

Agent Seer is a method that synthesizes realistic evaluation scenarios for AI agents by using tool specifications—function names, natural-language descriptions, and typed parameter schemas—rather than relying on hand-crafted scenarios or live tool execution. The approach aims to scale scenario generation across tool ecosystems and avoid static benchmarks that cannot keep up with evolving APIs.

7.0