A $1,500 foundation model that rivals larger LLMs - Venturebeat
Quick Answer
Researchers at Sapient developed HRM-Text, a 1B-parameter foundation model trained from scratch for $1,500, achieving competitive performance against larger LLMs on industry benchmarks.
Quick Take
This model utilizes a Hierarchical Recurrent Model architecture, focusing on instruction-response pairs instead of traditional autoregressive methods, significantly reducing training costs and data requirements.
Key Points
- HRM-Text trained for $1,500, a fraction of typical costs.
- Utilizes a Hierarchical Recurrent Model for enhanced sample efficiency.
- Achieved competitive performance against larger open models.
- Focuses on instruction-response pairs rather than raw text.
- Reduces the need for extensive internet-scale data.
Article Excerpt
From source RSS / original summary# Researchers say they trained a foundation model from scratch for about $1,500. Training a foundation from scratch costs millions and requires internet-scale data — which is why most enterprises don't bother. To overcome this brute-force scaling dogma, researchers at Sapient developed HRM-Text, which replaces standard Transformers with a highly sample-efficient Hierarchical Recurrent Model (HRM), an architecture they first introduced last year.
Instead of brute-force autoregressive prediction on raw text, HRM-Text trains exclusively on instruction-response pairs. The researchers were able to train a 1B-parameter HRM-Text from scratch at a fraction of the cost and tokens of normal LLMs. Their model achieved performance competitive with much larger open models on key industry benchmarks
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from WebSearch (Tavily)
See more →全球AI芯片峰会,9月上海见!
The 2026 Global AI Chip Summit will take place in Shanghai on September 22-23, focusing on the evolving AI chip landscape, including the shift from training to inference, the rise of diverse chip technologies, and the restructuring of industry competition. Notable speakers include experts from leading universities and companies, discussing advancements in AI chip architecture and applications.