DeepSignal
© 2026 DeepSignal · About
  • All
  • Featured
  • Latest
  • Guides
  • Daily
  • Weekly
  • Saved
  • Subscribe
  • Sources
  • About
  • Feedback
Sign in
  • Featured
  • Latest
  • Guides
  • Daily
  • Weekly

    AI Glossary

    What is GRPO?

    Overview

    GRPO, or Group Relative Policy Optimization, is a reinforcement-learning method that trains a model by comparing rewards across a group of sampled responses instead of relying on a separate value model. It matters because it can make reasoning-model post-training more memory-efficient while still encouraging responses that score better on verifiable tasks.

    Why it matters

    GRPO has become a practical post-training technique for improving reasoning behavior in language and multimodal models.

    Where it appears in AI research

    • Reasoning model technical reports
    • Reinforcement learning post-training
    • Verifiable reward experiments
    • Open-weight model training recipes

    Related terms

    Direct Preference OptimizationAgent EvaluationLarge Language Models (LLMs)