Elastic presents a reusable, production-grade evaluation framework for agentic AI
At QCon AI, Susan Chang (Principal Data Scientist at Elastic) described how Elastic moved from siloed, ad-hoc agent evaluations to a unified production-grade evaluation framework that combines LLM-as-judge and deterministic rules, connects Python data-science evals with TypeScript production code, and uses deep tracing to detect regressions across RAG and cybersecurity workloads while preserving domain context.