TabRank: Chain-of-Thought Distillation for Table Re-Rankers
Quick Answer
TabRank introduces a novel framework for training reasoning rerankers in tabular retrieval, achieving significant performance improvements across multiple datasets.
Quick Take
It enhances Acc@10 by 30.5% on HybridQA and 52.9% on TabFact, demonstrating effective generalization in multi-table scenarios. The framework leverages a dataset of 6728 reasoning traces for optimal training.
Key Points
- TabRank improves table retrieval performance by training reasoning rerankers.
- Achieved a 30.5% increase in Acc@10 on HybridQA dataset.
- Demonstrated a 52.9% improvement on TabFact subset.
- Utilizes a dataset of 6728 reasoning traces for training.
- Generalizes effectively to multi-table reasoning scenarios.
Paper Resources
📖 Reader Mode
~2 min readAbstract:The ability to retrieve relevant tables for answering questions is a key task for structured information retrieval. Multi-stage retrieval systems rely heavily on rerankers to refine candidate lists produced by efficient first-stage retrievers. As a result, neural rerankers and LLM-based reranking methods have become increasingly important due to their superior capacity for semantic understanding and reasoning compared to conventional sparse or dense retrieval models. Recently, Large Reasoning Models (LRMs) equipped with explicit chain-of-thought (CoT) reasoning have shown strong improvements in ranking quality in unstructured passage retrieval. In this work, we present TabRank, a framework for training reasoning rerankers for Tabular Retrieval. We first present a comprehensive dataset of 6728 reasoning traces for tabular reranking on the Natural Questions Tables dataset. We then explore two variants of training a compact reasoning model on these reasoning traces: explicit CoT distillation and conditioning the student reranker on the teacher's reasoning trace within the prompt. We stress-test TabRank on several out-of-distribution generalization settings on diverse domains and multi-table scenarios. Our approach significantly improves performance across a variety of table retrieval datasets, increasing Acc@10 by 30.5% on HybridQA, 15.2% on SQA, 52.9% on TabFact, and 13.1% on TATQA subsets of the Multi-Table QA Benchmark compared to the base model. Notably, TabRank generalizes effectively to multi-table reasoning. Our code, data and models are available at this https URL
| Comments: | 8 pages, 3 figures |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR) |
| Cite as: | arXiv:2607.25182 [cs.CL] |
| (or arXiv:2607.25182v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.25182 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Adarsh Singh [view email]
[v1]
Tue, 28 Jul 2026 01:23:04 UTC (881 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.