EpiBench: Multi-turn Research Workflows for Multimodal Agents (April 2026)
FreeBenchmarks multimodal agents on episodic scientific research workflows — literature search, figure extraction, cross-paper synthesis; built on smolagents with persistent memory and tool use
About EpiBench: Multi-turn Research Workflows for Multimodal Agents (April 2026)
EpiBench is an episodic multi-turn multimodal benchmark designed to evaluate AI agents on short scientific research workflows. It requires agents to proactively search literature, extract information from figures and tables, integrate evidence across multiple papers, and sustain use of accumulated evidence over multiple turns to answer objective questions that demand cross-paper comparisons and multi-figure integration. The benchmark introduces a process-level evaluation framework for fine-grained testing and diagnosis of research agents. Experiments show that even the leading model achieves only 29.23% accuracy on the hard split, highlighting significant room for improvement in multi-turn, multi-evidence research workflows.
Key Features
Pros & Cons
- Introduces a novel benchmark for multi-turn, multi-evidence research workflows not covered by existing evaluations
- Provides process-level evaluation for detailed diagnostic insights
- Open-access benchmark available on arXiv with downloadable code and data
- Reveals significant performance gaps in current models (leading model only 29.23% on hard split)
- Currently only covers short research workflows; longer or more complex workflows not yet addressed
- Accuracy on the hard split is very low (29.23%), indicating the benchmark is extremely challenging
- Limited to scientific research domain; not generalizable to other multi-turn tasks