PRMBENCH: A Fine-grained and Challenging Benchmark for Process-Level Reward Models
FreeA fine-grained benchmark for process-level reward model evaluation
FreeFree tier
About PRMBENCH: A Fine-grained and Challenging Benchmark for Process-Level Reward Models
PRMBench is a benchmark designed to evaluate the fine-grained error detection capabilities of Process-Level Reward Models (PRMs). It comprises 6,216 carefully designed problems and 83,456 step-level labels, assessing models across multiple dimensions including simplicity, soundness, and sensitivity. The benchmark reveals significant weaknesses in current PRMs, highlighting challenges in process-level evaluation and guiding future research. Accepted at ACL 2025 Main.
Key Features
6,216 carefully designed problems with 83,456 step-level labels
Evaluates PRMs across multiple dimensions: simplicity, soundness, and sensitivity
Designed for fine-grained error detection in complex reasoning and decision-making tasks
Tests 15 models including open-source PRMs and closed-source LLMs prompted as critics
Accepted at ACL 2025 Main
Pros & Cons
Pros
- Comprehensive evaluation with a large number of fine-grained labels
- Systematic assessment across multiple dimensions (simplicity, soundness, sensitivity)
- Reveals significant weaknesses in current PRMs, guiding future improvements
- Open-source and freely available for research
Cons
- Limited to process-level reward models; may not apply to other types of reward models
- Benchmark results may not directly translate to real-world deployment performance
Best For
Evaluating process-level reward models (PRMs) for complex reasoning tasksAssessing fine-grained error detection in step-by-step reasoning processesBenchmarking model performance on implicit error types in real-world scenariosGuiding research and development of more robust process-level reward models
FAQ
What is PRMBench?
PRMBench is a benchmark designed to evaluate the fine-grained error detection capabilities of Process-Level Reward Models (PRMs). It consists of 6,216 problems and 83,456 step-level labels.
What dimensions does PRMBench evaluate?
PRMBench evaluates PRMs across multiple dimensions including simplicity, soundness, and sensitivity.
What models were tested in PRMBench?
The benchmark tested 15 models, spanning both open-source PRMs and closed-source large language models prompted as critic models.