RESEARCH · RESEARCH · #1007
PFArena: benchmark of PLMs, LLMs, and agents for protein modification (arXiv:2609.28921v1)
PFArena is a new benchmark (arXiv:2609.28921v1) that provides four controlled task interfaces for protein modification, covering single-mutant generation and multi-mutant ranking under varying levels of mutation fitness data. The authors evaluate six PLMs, six LLMs, and five LLM-based agents using complementary metrics, find that PLMs excel at open-ended single-mutant generation while LLMs and agents perform better at multi-mutant ranking when target-specific fitness data exist, and report that all model families struggle as search-space size and mutation depth increase; the code and benchmark suite are released to support reproducible work.
KEY POINTS
- PFArena is a new benchmark (arXiv:2609.28921v1) that provides four controlled task interfaces for protein modification, covering single-mutant generation and multi-mutant ranking under varying levels of mutation fitness data.
- The authors evaluate six PLMs, six LLMs, and five LLM-based agents using complementary metrics, find that PLMs excel at open-ended single-mutant generation while LLMs and agents perform better at multi-mutant ranking when target-specific fitness data exist, and report that all model families struggle as search-space size and mutation depth increase; the code and benchmark suite are released to support reproducible work.
- Provides a standardized, reproducible benchmark that clarifies where PLMs, LLMs, and agents succeed or fail in realistic protein-modification decision settings, guiding future model development and experimental design.
WHY IT MATTERS
Provides a standardized, reproducible benchmark that clarifies where PLMs, LLMs, and agents succeed or fail in realistic protein-modification decision settings, guiding future model development and experimental design.
SOURCES & TIMELINE
1arXiv:2609.28921v1 Announce Type: new Abstract: Protein modification requires navigating an immense sequence space, yet wet-lab validation remains low-throughput and costly. Although computational paradigms including protein language models (PLMs), large language models (LLMs), and LLM-based agents have shown promise in protein modification, their relative efficacy across realistic experimental decision-making settin…
arXiv:2609.28850v1 Announce Type: new Abstract: Reproducing a machine learning paper involves most research steps, from installing software and debugging to running experiments, work that AI agents increasingly do. We introduce RECLAIM, a benchmark of 100 NeurIPS 2025 papers that can be rebuilt yearly from new conferences. For each paper we fix in advance the result to reproduce, what counts as a successful reproduct…