TAPR: Enhancing LLM Performance with a Task-Aware Prompt Rewriter
Quick Answer
This paper shows that The Task-Aware Prompt Rewriter (TAPR) enhances LLM performance by reformulating prompts for tasks like question answering and summarization.
Quick Take
Trained with reinforcement learning, TAPR shows consistent improvements over base models, achieving higher accuracy on benchmarks such as Natural Questions and GSM8K. The code is available for further exploration.
Key Points
- TAPR reformulates user prompts into task-optimized versions for better performance.
- Utilizes for training with LLM-as-judge evaluations.
- Demonstrated consistent gains in prompt rewriting across various tasks.
- Fine-tuning Phi-4-mini-instruct leads to clearer, more instructive prompts.
- Achieved higher accuracy on benchmarks like Natural Questions and GSM8K.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Large Language Models (LLMs) often require carefully crafted prompts to unlock their full potential, which can be a barrier for non-expert users. This work addresses the challenge by introducing a Task-Aware Prompt Rewriter (TAPR), a model that reformulates user prompts into task-optimized prompts with the explicit goal of improving downstream LLM performance. We train TAPR using reinforcement learning with Group Relative Policy Optimization (GRPO), where rewards are derived from LLM-as-judge evaluations of both the reformulated prompt and the corresponding task output. Experimental results on diverse tasks, such as question answering, summarization, and arithmetic reasoning, show that our method yields consistent gains over base models in prompt rewriting ability. Fine-tuning Phi-4-mini-instruct (as the base model for TAPR) produces prompts that contain clearer and more instructive language, leading to higher accuracy on established benchmarks such as Natural Questions and GSM8K. Our code is available at: this https URL
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.28657 [cs.AI] |
| (or arXiv:2607.28657v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.28657 arXiv-issued DOI via DataCite |
Submission history
From: Hosein Azarbonyad [view email]
[v1]
Fri, 17 Jul 2026 15:37:07 UTC (4,618 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.