Codifying the Judge: Scalable Evaluation via Program Distillation
Quick Answer
The PAJAMA system introduces program distillation as a scalable alternative to LLMs for automated evaluation, matching the performance of a 13B-size LLM while significantly reducing API costs.
Quick Take
It synthesizes programmatic judges that provide transparency and efficiency, achieving better accuracy and throughput across five datasets. Additionally, a reward model distilled from these programs outperforms traditional -based models at a fraction of the cost.
Key Points
- PAJAMA synthesizes programs that act as judges, improving evaluation transparency.
- Programmatic judges match the performance of a 13B-size LLM across five datasets.
- Using program outputs as routing signals enhances both accuracy and throughput.
- The distilled reward model outperforms LLM-based models at 100x lower API costs.
- Program distillation eliminates per-sample API costs, enhancing scalability.
DeepSignal Analysis
What happened
The PAJAMA system introduces program distillation as a cost-effective alternative to large language models (LLMs) for automated evaluation. It matches the performance of a 13B-size LLM while reducing API costs significantly. The system synthesizes programmatic judges that enhance transparency and efficiency in decision-making.
Key evidence
- PAJAMA synthesizes programs that act as judges, aggregating their decisions into a final verdict while allowing for inspection and editing.
- The programmatic judges achieved performance comparable to a 13B-size LLM across five datasets and four model families.
- A reward model derived from the programmatic judges outperformed a traditional LLM-based model at a significantly lower API cost, demonstrating improved efficiency.
Why it matters
The introduction of program distillation could reshape automated evaluation by addressing the high costs and latency associated with LLMs. This approach not only enhances transparency but also allows for more efficient decision-making processes. By matching the performance of large models at a fraction of the cost, PAJAMA may enable broader adoption of automated evaluation systems in various applications.
Paper Resources
📖 Reader Mode
~2 min readAbstract:LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, significant latency, and opaque decisions -- limitations that undermine its scalability and reliability. We address these with a simple, efficient alternative: program distillation. Instead of prompting an LLM at the evaluation time, we distill its decision logic into a committee of programs that score candidates directly. These programmatic judges offer transparency, are easily inspected or edited, and eliminate per-sample API costs. Building on this notion, we introduce PAJAMA, a system that synthesizes programs as judges, aggregates their decisions into a joint verdict, and incorporates a fallback mechanism to selectively escalate low-confidence cases to an LLM. Across five datasets and four model families, we show that programmatic judges can match the performance of a 13B-size LLM judge. When using program outputs as routing signals, PAJAMA improves both accuracy and throughput and advances the Pareto frontier. Beyond evaluation, programmatic judges produce cheap and effective reward signals: on RewardBench, a reward model distilled from programs' verdicts outperforms one trained on a proprietary LLM's labels at two orders of magnitude lower API cost.
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.22561 [cs.AI] |
| (or arXiv:2607.22561v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.22561 arXiv-issued DOI via DataCite |
Submission history
From: Tzu-Heng Huang [view email]
[v1]
Fri, 29 May 2026 03:13:24 UTC (9,073 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.