Agent Evaluation Metric (AEM) for multi-turn conversations
An AWS Machine Learning post introduces the Agent Evaluation Metric (AEM), a decomposable, turn-level metric for evaluating multi-turn conversational agents. The post demonstrates AEM's first dimension—correctness—showing how it can pinpoint the specific turn that caused a failure and distinguish that root cause from later turns that inherited the error.