Model-Native Computing Architecture: Envisioning Future System Architecture Through the Lens of Computer Architecture
Quick Answer
The paper proposes the Intelligent Computing Architecture Model (ICAM), a six-layer framework for model-native computing, addressing issues like cache reuse and agent scheduling in large language models (LLMs) such as Codex and Claude Code.
Quick Take
It introduces design laws to optimize performance and highlights the need for a unified model in systems, while also outlining a research roadmap for future developments.
Key Points
- ICAM resolves the debate on LLMs as CPUs or operating systems with a dual-plane view.
- Introduces three design laws for optimizing cache reuse, context management, and agent collaboration.
- Validates design laws against existing system-level data and agentic software practices.
- Highlights the need for a unified model in the emerging model-native stack.
- Outlines a research roadmap for advancing model-native computing.
Paper Resources
Article Content
From source RSS / original summaryarXiv:2606. 00288v1 Announce Type: new Abstract: are undergoing a transition from model technology to system technology. As developers use Codex, Claude Code, AutoGPT, and related agents to write code, manage projects, and execute multi-step tasks, recurring engineering problems such as cache reuse, context management, agent scheduling, and permission control increasingly resemble classical computer systems problems. This paper develops that analogy as a visionary survey.
We map concepts from computer architecture to the emerging model-native stack and review work on LLM-as-OS, memory management, agent frameworks, tool protocols, coordination, cognitive architectures, and safety governance. We argue that these strands address different layers of the same system but lack a unified model. To fill this gap, we propose the Intelligent Computing Architecture Model (ICAM), a six-layer framework for model-native computing with explicit interface contracts and design axioms.
ICAM resolves the apparent tension over whether an LLM is more like a CPU or an operating system through a dual-plane view: a probabilistic execution plane concerned with what can be computed, and a deterministic control plane concerned with what should be computed.
We further introduce three design laws: the Semantic Locality Law for KV-cache reuse and inference speedup, the Context Budget Law for effective working sets under finite windows and attention decay, and the Agent Speedup Law for diminishing returns in multi-agent collaboration. We validate these laws against published system-level data and relate them to recent evidence on agentic software practices.
We conclude by identifying where the analogy breaks down and outlining a research roadmap for model-native computing. This is a conceptual and survey contribution; it does not report new experiments.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.