Procedural Memory Distillation: Online Reflection for Self-Improving Language Models
Quick Answer
This paper shows that Procedural Memory Distillation (PMD) enhances reinforcement learning by converting cross-episode signals into reusable memory, improving Qwen3-8B and OLMo3-Instruct-7B models by 3.8-5.5% on SCIKNOWEVAL and 7.9-13.6% on LIVECODEBENCH.
Quick Take
The co-evolution of policy and memory allows for more effective self-supervision, demonstrating significant performance gains when both components are active.
Key Points
- PMD organizes memory into raw trajectories, strategies, and higher-level patterns.
- Empirical results show PMD outperforms SDPO by significant margins.
- Freezing either memory or policy reduces PMD's performance by over 10%.
- The model learns from its own experiences for better procedural knowledge.
- PMD facilitates online reflection for continuous self-improvement.
DeepSignal Analysis
What happened
The Procedural Memory Distillation (PMD) method enhances reinforcement learning by converting cross-episode signals into reusable memory. This approach has shown performance improvements in the Qwen3-8B and OLMo3-Instruct-7B models, achieving gains of 3.8-5.5% on SCIKNOWEVAL and 7.9-13.6% on LIVECODEBENCH. The method emphasizes the co-evolution of policy and memory, which is crucial for effective self-supervision.
Key evidence
- PMD improves the Qwen3-8B and OLMo3-Instruct-7B models by 3.8-5.5% on SCIKNOWEVAL and 7.9-13.6% on LIVECODEBENCH.
- The method organizes memory into three levels: raw trajectories, self-reflected strategies, and higher-level behavioral patterns.
- Freezing either the memory or the policy results in a performance drop of more than 10% across SCIKNOWEVAL domains.
Why it matters
The introduction of PMD represents a significant advancement in reinforcement learning by enabling models to retain and utilize procedural knowledge across episodes. This could lead to more robust AI systems capable of learning from past experiences, thereby improving their decision-making processes. The co-evolution aspect suggests that the interaction between policy and memory is vital for optimizing performance, which may influence future research directions in AI.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Reinforcement learning with verifiable rewards (RLVR), along with recent selfdistillation variants such as SDPO, evaluates each rollout against a verifier and updates the policy from that episode-level signal. However, the richer procedural information in the rollout is rarely retained or reused. Across episodes and epochs, the model repeatedly encounters related problems under a changing policy, producing cross-episode signals that episode-local updates cannot capture: which strategies consistently pass verification, which failure modes persist, which patterns recur. We propose Procedural Memory Distillation (PMD), which converts these crossepisode signals into reusable procedural memory and distills it into the policy's weights during training. This memory functions as a training scaffold, absorbed into the policy itself, yielding a memory-free model at inference. PMD organizes the memory at three levels of abstraction: raw trajectories, self-reflected strategies and lessons, and higher-level behavioral patterns that recur across problems, all extracted online from the model's own trajectories. A memory-conditioned self-teacher draws on the accumulated experience to supervise the student on its own rollouts, enabling student to progressively internalize procedural knowledge within its parameters. The central design principle is co-evolution: the policy generates rollouts that update the memory, and memory shapes the supervision that updates the policy. Empirically, across Qwen3-8B and OLMo3-Instruct-7B, PMD improves over SDPO by 3.8-5.5% on SCIKNOWEVAL and 7.9-13.6% on LIVECODEBENCH. Co-evolution powers these gains: freezing either the memory or the policy trails PMD by more than 10% across SCIKNOWEVAL domains.
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.01480 [cs.AI] |
| (or arXiv:2607.01480v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.01480 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Semih Yavuz [view email]
[v1]
Wed, 1 Jul 2026 21:20:57 UTC (696 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.