PPRBench: A Process-level Benchmark for LLMs’ Physical Reasoning

Published in under review, 2026

We introduce PPRBench, a contamination-free benchmark designed to evaluate the process-level physical reasoning capabilities of large language models (LLMs). Unlike existing benchmarks that focus solely on final-answer correctness, PPRBench assesses the quality of each reasoning step.

We also develop PhysGrader, a rubric-guided LLM-as-a-judge framework that provides fine-grained, step-by-step evaluation of physics reasoning.

Key contributions:

  • A novel contamination-free benchmark covering diverse physics reasoning tasks
  • PhysGrader: a scalable, rubric-guided process-level evaluator
  • Exploration of GRPO with process rewards generated by PhysGrader
  • Comprehensive evaluation of 16 frontier LLMs, revealing systematic reasoning fidelity gaps

Recommended citation: Denghong Luan, Wentao Shi, Wenjie Wang, Xiangnan He. "PPRBench: A Process-level Benchmark for LLMs' Physical Reasoning." under review.