PPRBench: A Process-level Benchmark for LLMs’ Physical Reasoning
Published in under review, 2026
We introduce PPRBench, a contamination-free benchmark designed to evaluate the process-level physical reasoning capabilities of large language models (LLMs). Unlike existing benchmarks that focus solely on final-answer correctness, PPRBench assesses the quality of each reasoning step.
We also develop PhysGrader, a rubric-guided LLM-as-a-judge framework that provides fine-grained, step-by-step evaluation of physics reasoning.
Key contributions:
- A novel contamination-free benchmark covering diverse physics reasoning tasks
- PhysGrader: a scalable, rubric-guided process-level evaluator
- Exploration of GRPO with process rewards generated by PhysGrader
- Comprehensive evaluation of 16 frontier LLMs, revealing systematic reasoning fidelity gaps
Recommended citation: Denghong Luan, Wentao Shi, Wenjie Wang, Xiangnan He. "PPRBench: A Process-level Benchmark for LLMs' Physical Reasoning." under review.
