RESEARCH · RESEARCH · #1099
ORCA benchmark evaluates LLMs on data-science code translation
ORCA is a new benchmark for Data Science Code Translation (DSCT) introduced on arXiv (2609.30749v1). It contains ORCA-MAIN (1,600 grounding-level tasks across Data Querying, Data Manipulation, and Deep Learning) and ORCA-PROJECT (200 full-project translations), with annotated reference translations and test cases; evaluations show limited LLM performance (Claude-Opus-4.6: 56.92% on ORCA-MAIN, 33.67% on ORCA-PROJECT) and report that an intent-augmented two-stage method yields modest absolute success-rate gains (~4.80% and ~5.33%).
KEY POINTS
- ORCA is a new benchmark for Data Science Code Translation (DSCT) introduced on arXiv (2609.30749v1).
- It contains ORCA-MAIN (1,600 grounding-level tasks across Data Querying, Data Manipulation, and Deep Learning) and ORCA-PROJECT (200 full-project translations), with annotated reference translations and test cases; evaluations show limited LLM performance (Claude-Opus-4.6: 56.92% on ORCA-MAIN, 33.67% on ORCA-PROJECT) and report that an intent-augmented two-stage method yields modest absolute success-rate gains (~4.80% and ~5.33%).
- ORCA provides a targeted, functionally validated benchmark and baseline results that quantify shortcomings of current LLMs on translating real data-science code and suggests a simple intent-augmentation that improves performance.
WHY IT MATTERS
ORCA provides a targeted, functionally validated benchmark and baseline results that quantify shortcomings of current LLMs on translating real data-science code and suggests a simple intent-augmentation that improves performance.