Plan, divide, and conquer: How weak models excel at long context tasks
Quick Answer
The research from Together AI reveals that smaller models using a 'Divide & Conquer' framework can outperform GPT-4o in long context tasks, demonstrating significant performance gains while being cheaper and faster.
Quick Take
This approach addresses model confusion and aggregation noise, making it effective for tasks like QA and summarization, though it has limitations for high-synergy tasks.
Key Points
- Smaller models like Llama-3-70B outperform GPT-4o in long context tasks using Divide & Conquer.
- The framework reduces model confusion and aggregation noise for better performance.
- Testing only 5 random samples can optimize chunk size efficiently.
- Divide & Conquer is cheaper and faster, leveraging parallel processing of smaller models.
- Best suited for moderate cross-chunk dependency tasks like QA and summarization.
Source Excerpt
As context windows grow, performance degrades in unexpected ways. We show how a "Divide & Conquer" framework — breaking long documents into parallel chunks with a planner, workers, and manager — lets smaller models like Llama-3-70B and Qwen-72B outperform GPT-4o single-shot.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from Together AI
See more →
Open, convenient and predictable: Introducing Provisioned Throughput
Together AI introduces Provisioned Throughput, offering guaranteed inference capacity for MiniMax M3 and GLM-5.2 at $0.05 per PTU per minute, achieving costs up to 90% lower than Claude Opus 4.8. This new model provides predictable pricing and a 99% uptime SLA, catering to companies transitioning to open weight models for production workloads.

