AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adaptive Recovery
Quick Answer
AnovaX is a local voice assistant that operates entirely on a user's computer, utilizing a multi-agent architecture and LLM planning for task execution.
Quick Take
It features a safety layer, adaptive recovery, and a companion Flask server for remote control via mobile devices, demonstrating that a lightweight assistant can effectively manage desktop tasks without relying on cloud services.
Key Points
- AnovaX runs entirely on the user's computer, avoiding cloud dependency.
- Utilizes a with specialized agents for various tasks.
- Incorporates an adaptive recovery loop to handle execution failures.
- Features a Flask server for real-time remote control via mobile devices.
- Demonstrates effective desktop task management with a few thousand lines of code.
DeepSignal Analysis
What happened
AnovaX is a local voice assistant designed to operate entirely on a user's computer, utilizing a multi-agent architecture and LLM planning for task execution. It incorporates a safety layer and adaptive recovery mechanisms, allowing it to manage desktop tasks without cloud dependency.
Key evidence
- AnovaX operates as a local-first assistant, meaning it runs entirely on the user's computer without sending data to the cloud.
- The architecture includes a multi-agent orchestrator that translates plans into typed child agents, each with specific timeout and retry policies.
- A Flask server allows remote control of AnovaX via mobile devices, enabling real-time monitoring of agent activities and screen streaming.
Why it matters
The development of AnovaX highlights a shift towards local processing in voice assistants, which can enhance user privacy and reduce latency. By demonstrating that a lightweight assistant can effectively manage desktop tasks, it challenges the prevailing reliance on cloud-based solutions, potentially influencing future designs in the industry.
What to watch
Paper Resources
📖 Reader Mode
~2 min readAbstract:Desktop voice assistants are still dominated by cloud pipelines that ship raw audio off the machine and expose a fixed set of skills. We describe AnovaX, a small local-first assistant that runs entirely on the user's computer and treats the desktop itself as its action surface. A single Python process wires together a wake-word gate, a speech pipeline, an LLM planner (Gemini) that emits a JSON plan of tool calls, a whitelist-and-denylist safety layer, a multi-agent orchestrator that translates each plan into typed child agents on a bounded thread pool, and an adaptive recovery loop that takes over whenever a core step fails. Every tool corresponds to a specialized agent class (AppAgent, TypingAgent, BrowserAgent and six others) with its own timeout, retry policy, and shared-resource locks. A recursive MetaAgent lets the planner delegate a sub-goal back to itself, capped at two levels of nesting. The recovery loop uses a compact ReAct-style prompt and hides Gemini's latency behind speculative execution of read-only tools. A companion Flask server exposes a phone-friendly remote over the local WiFi, mirrors every agent lifecycle event to the phone in real time, and streams the laptop's screen back over MJPEG so the user can watch remote commands land as they run. The point of the project is less to compete with Siri or Alexa than to show that a legible, few-thousand-line assistant is enough to open apps, type into them, run searches, coordinate concurrent actions, recover from single-step failures, and be driven entirely from a phone in another room -- without the LLM ever touching the keyboard.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.15367 [cs.AI] |
| (or arXiv:2607.15367v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.15367 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Raunak B Sinha [view email]
[v1]
Thu, 16 Jul 2026 18:09:36 UTC (17 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.