
Meta AI uses a second AI agent as a memory coach to keep long tasks on track
Quick Answer
Meta AI introduces a dual-agent system where a 'memory agent' enhances an 'action agent' by selectively recalling relevant information, improving performance on benchmarks like Terminal-Bench 2.0 from 38% to 46%.
Quick Take
This innovative approach reduces errors and latency, demonstrating that targeted memory interventions outperform constant recall strategies.
Key Points
- The memory agent updates a structured memory bank to aid the action agent.
- Selective reminders improved task completion rates by up to 10 percentage points.
- Performance gains were observed even with stronger action agents like Opus 4.6.
- The system outperformed traditional memory retrieval methods like Mem0.
- Training smaller models like Qwen3.5-27B improved memory management through fine-tuning.
DeepSignal Analysis
What happened
Meta AI has developed a dual-agent system consisting of a 'memory agent' and an 'action agent' to enhance task performance. This system improved scores on benchmarks like Terminal-Bench 2.0, increasing success rates from 38% to 46%. The memory agent selectively decides when to provide reminders, which helps reduce errors and latency during task execution.
Key evidence
- The memory agent updates a structured memory bank and decides whether to remind the action agent based on recent task steps.
- In tests, the system achieved a 46% success rate on Terminal-Bench 2.0, compared to 38% for the baseline without the memory agent.
- The memory agent's selective intervention outperformed a version that provided constant reminders, indicating that targeted memory use is more effective.
Why it matters
This approach addresses the challenge of 'behavioral state decay' in AI agents, where relevant information can become obscured over time. By improving how agents recall and utilize past information, Meta AI's system could enhance the efficiency and accuracy of AI in complex tasks. This could have implications for various applications, from customer service to autonomous systems.
Source Excerpt
Meta AI wants to stop AI agents from forgetting errors they've already diagnosed and repeating failed steps during complex tasks. A separate memory agent maintains a structured memory bank and decides when to remind the main agent and when to stay silent. The system improved scores by up to 8. 3 percentage points across two benchmarks.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

