Microsoft Research Podcast interviews Jennifer Neville on AI evaluation and unexpected failures
In a Microsoft Research Podcast episode, partner research manager Jennifer Neville discusses the role of evaluation in pushing AI performance, surprising model failures that appear beyond standard benchmarks, practical guidance for working with current systems, and why close attention to training data matters when results defy expectations.