Context Recycling for Long-Horizon LLM Inference
Quick Answer
ContextForge enhances long-horizon reasoning in large language models (LLMs) by recycling context through structured query generation and external memory retrieval.
Quick Take
In a 15-turn conversational benchmark, it shows improved consistency and reduced token usage compared to baseline models, maintaining response accuracy. This approach allows to extend their capabilities without larger context windows or retraining.
Key Points
- ContextForge reduces token overhead while preserving answer quality in LLMs.
- The system enables efficient reuse of prior computations across conversational turns.
- In tests, ContextForge improved consistency over a 15-turn healthcare query benchmark.
- No need for larger context windows or model retraining with ContextForge.
- Code and evaluation artifacts are available on GitHub.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Large language models (LLMs) exhibit strong capabilities in short-context reasoning but degrade in performance over long conversational horizons due to context window limitations and inefficient token usage. We introduce ContextForge, a system for context recycling that maintains task-relevant information across turns by combining structured query generation, external memory retrieval, and controlled synthesis. The system enables efficient reuse of prior computation without relying on full context replay, reducing token overhead while preserving answer quality. We evaluate ContextForge using a 15-turn conversational benchmark that tests multi-turn reasoning, back-references, and domain shifts across structured healthcare queries. Compared to a baseline agent using identical underlying models, ContextForge demonstrates improved consistency and reduced token consumption, while maintaining comparable response accuracy. These results suggest that context recycling provides a practical approach for extending LLM capabilities in long-horizon tasks without requiring larger context windows or model retraining. Code and evaluation artifacts are available at this https URL.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2606.26105 [cs.CL] |
| (or arXiv:2606.26105v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2606.26105 arXiv-issued DOI via DataCite |
Submission history
From: Derek Thomas [view email]
[v1]
Fri, 1 May 2026 04:45:08 UTC (19 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.