Lossy Compressive Text Autoencoders
Quick Answer
The proposed lossy compressive text autoencoder achieves a compression rate of 2.24 bits per byte, matching lossless algorithms while maintaining strong reconstruction and performance on downstream tasks like question-answering and semantic similarity.
Quick Take
The architecture utilizes residual downscaling and upscaling of hidden representations, evaluated across various quantization methods and datasets.
Key Points
- Autoencoder architecture employs residual downscaling and upscaling of hidden representations.
- Achieves 2.24 bits per byte compression on web text data.
- Evaluated using BLEU for surface-level and -based judge for semantic-level similarity.
- Demonstrates competitive performance on downstream benchmarks for question-answering.
- Explores various quantization methods and training objectives.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Our work explores learning a compressed latent representation of text, at the intersection of data compression and representation learning. We propose an autoencoder architecture that performs residual downscaling and upscaling of hidden representations along the time axis, with a residual low-dimension discrete bottleneck. We analyze our approach for different quantization methods, training objectives, and datasets. For different levels of compression, we evaluate the similarity between the original and reconstructed text both at the surface-level (BLEU) and at the semantic-level (LLM-based judge). Additionally, we evaluate our models on downstream question-answering and semantic text similarity benchmarks. Our approach results in compressed representations which are on par with lossless text compression algorithms at 2.24 bits per byte on web text data, while having good reconstruction and downstream task performance.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.10738 [cs.CL] |
| (or arXiv:2610.10738v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10738 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: David Grangier [view email]
[v1]
Wed, 7 Oct 2026 18:08:40 UTC (14,137 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →The "10th Juror": Open-Set Standpoint Screening for Bureaucratic Bias Detection
MARS-Gov introduces a framework for detecting bureaucratic bias in Dutch government documents, achieving a new state-of-the-art F1 score of 0.880. This model outperforms existing zero-shot detectors by 20.2 points and reduces unnecessary interventions to just 2.5%. The framework's dynamic '10th juror' adapts to emerging biases, enhancing legal language processing.