Parcae: Doing more with fewer parameters using stable looped models
Quick Answer
Parcae, a new stable looped architecture by Together AI, achieves up to 6.3% lower validation perplexity than previous models while using only 770M parameters, matching the performance of a 1.3B parameter transformer.
Quick Take
This innovation allows for scaling model quality without increasing memory footprint, addressing the challenges of training looped models effectively.
Key Points
- Parcae achieves 6.3% lower validation perplexity than previous looped models.
- The model uses 770M parameters, matching a 1.3B parameter transformer.
- First scaling laws for looping established, requiring increased looping and data.
- Looped models traditionally suffer from training instability and residual state explosion.
- Parcae stabilizes training by constraining input injection parameters.
Source Excerpt
Parcae is a stable looped language model that matches the quality of a Transformer twice its size — a 770M model reaching 1. 3B-level performance. We introduce the first scaling laws for looping and show that increasing recurrence, not just data, is a compute-efficient path to bet
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from Together AI
See more →
Open, convenient and predictable: Introducing Provisioned Throughput
Together AI introduces Provisioned Throughput, offering guaranteed inference capacity for MiniMax M3 and GLM-5.2 at $0.05 per PTU per minute, achieving costs up to 90% lower than Claude Opus 4.8. This new model provides predictable pricing and a 99% uptime SLA, catering to companies transitioning to open weight models for production workloads.

