← Back to all stories

Grading the Journey: How Process Reward Models (PRMs) Replaced Flawed Outcome-Based Evaluation

Consider a high school math teacher grading an exam. If a student writes five lines of completely nonsensical, mathematically impossible arithmetic, but accidentally makes two canceling errors and writes the correct final number '42' at the bottom of the page, a strict teacher does not award 100%. They mark the faulty algebra with red ink. For years, AI training made that exact error by relying exclusively on Outcome-Based Reward Models (ORMs).

The Dangerous Flaw of Outcome Rewards (ORMs)

In Outcome-Based Reinforcement Learning, the model generates an entire 20-step derivation, and the training system checks only whether the final answer matches ground truth ($+1$ for correct, $0$ for wrong).

This creates severe pathologies during training:

  • Rewarding Lucky False Logic: The model is rewarded for hallucinations and flawed steps if the final number accidentally lands on the right answer.
  • Punishing Brilliant Derivations: If a model executes 19 steps of flawless Olympiad-level mathematical reasoning but makes a tiny arithmetic slip in the final addition, an ORM assigns a score of zero, discouraging deep reasoning.
  • Sparse Credit Assignment: The model receives no granular feedback on which specific step was the turning point between success and failure.
[Outcome Reward Model (ORM) vs. Process Reward Model (PRM)]

Outcome Reward Model (ORM: Grades only the finish line):
[Step 1: OK] ──► [Step 2: Flawed math!] ──► [Step 3: Lucky cancellation] ──► "Answer: 42"
└── ORM Feedback: Reward = +1.0 (Mistakenly reinforces faulty reasoning!)

Process Reward Model (PRM: Grades every step):
[Step 1: Expand brackets]      ──► PRM Step Score: +1.0 (Valid ✓)
[Step 2: Divide by zero error]  ──► PRM Step Score: -1.0 (ERROR LOCATED HERE ✗)
[Step 3: Conclude answer 42]   ──► PRM Step Score:  0.0 (Pruned branch!)

The Process Reward Model (PRM) Revolution

Pioneered by OpenAI's PRM800K research, Process Reward Models evaluate the logical correctness of every single intermediate step in the chain of thought:

  1. Step-Level Value Scoring: At each step delimiter (e.g. \n\n), the PRM computes a probability score that the current mathematical state can lead to a valid solution.
  2. Best-of-N Guided Beam Search: During inference, the system generates multiple candidate continuations at each step, using the PRM to prune branches that receive low step scores and expanding only high-scoring logical branches.
  3. Active Error Detection: When the PRM score drops sharply between Step $K$ and Step $K+1$, the model immediately backtracks and explores alternative derivations.

Engineering Takeaway

True cognitive reasoning requires verifying every step of the journey, not just celebrating the destination. Use Process Reward Models and step-level verifiers to guide complex problem-solving pipelines.

Reference Paper / Context: Let's Verify Step by Step: Process Reward Models in Mathematical Reasoning (Lightman et al.) — Read source ↗
👨‍💻
About the Author

I am Vikram Samal, an AI systems architect exploring how intelligent systems reason, adapt, and act—and how to make them reliable at scale. I connect emerging AI capabilities with the architectural decisions that shape performance, trust, and practical value. Through this blog, I share insights into the ideas and engineering choices shaping AI’s next chapter. As a proud father of two, I believe curiosity, human judgment, and continuous learning are essential in a world being transformed by AI.

Read full bio & connect on LinkedIn →
Previous
← Tracing the Swarm: How OpenTelemetry and Semantic Spans Conquered Multi-Agent Observability
Next
The Unified Store: Why Dedicated Vector Databases Lost to Columnar and Relational Engines →