BioEVAL: global, multi-institutional benchmark evaluating LLMs and multimodal models on bioengineering
BioEVAL (BioEngineering Validation of AI and LLMs) is a multi-institutional, PhD-level benchmark released on arXiv (2609.30489v1) that spans 11 bioengineering subfields and comprises 608 items: 380 MCQs (359 retained after audit), 218 literature-synthesis tasks, and 10 multimodal experimental image problems. The authors evaluated cloud-scale and locally deployable foundation/multimodal models (e.g., ChatGPT, Gemini, Grok), reporting up to 90% MCQ accuracy, 0.72 literature-synthesis similarity, and 80% accuracy on a small multimodal sample, and maintain BioEVAL as an extensible benchmark with standardized contribution and evaluation protocols.