https://arxiv.org/list/cs.AI/recent
AI research updates from arXiv cs.AI, filtered for agents, planning, reasoning, evaluation and AI systems with readable summaries and signal scores.
DeepSignal tracks AI updates from arXiv cs.AI, filtering research and product signals into plain-English summaries, signal scores and source-linked article pages.
Current topics: Research, AI Assistant, LLM, Agent, Inference
FluidPD is a novel P/D-disaggregated serving system that enhances SLO attainment by up to 94.6 percentage points over static SGLang. It employs FluidToken and FluidRole to dynamically manage prefill and decode resource allocation, addressing both transient and sustained imbalances without requiring additional workers. This approach significantly improves service quality in LLM serving environments.
FluidPD introduces a dynamic resource allocation system for LLM serving that improves service level objective (SLO) attainment by up to 94.6 percentage points. This development is crucial for builders and PMs as it enhances service quality without the need for additional infrastructure, making it a cost-effective solution for scaling AI applications.
This study reveals that enhanced traffic forecasts do not guarantee improved signal control decisions, as evidenced by a 6.09% increase in queue vehicle-seconds despite a 90% conformal interval achieving 90.72% marginal coverage. The research highlights the importance of temporal observability and action identifiability in translating predictive improvements into operational benefits.
The study demonstrates that even with advanced traffic forecasts, operational decisions like signal control can still result in inefficiencies, as shown by a 6.09% increase in queue vehicle-seconds. This highlights the need for builders and PMs to focus not only on predictive analytics but also on the implementation of actionable insights to achieve tangible improvements in traffic management.
EPOCH is an evidence-governed architecture that enhances AI research agents' discovery capabilities, achieving a mean normalized score of 0.65 on AlgoTune, surpassing the previous baseline of 0.53. It demonstrates significant advancements across ten discovery problems, yielding improved algorithms and proof-supported results, thereby promoting more reliable scientific discoveries.
The development of EPOCH, which enhances AI agents' discovery capabilities with a mean score of 0.65 on AlgoTune, indicates a significant leap in algorithmic efficiency and reliability. For builders and PMs, this could lead to more effective AI tools for research and development, while investors may see potential for new applications in scientific discovery and innovation.
AegisFlow is a novel multi-agent AI framework that automates remediation in data ecosystems, achieving a 98.1% reduction in Mean Time to Repair (MTTR) from 170 minutes to 3.2 minutes, with a 92% patch success rate. The system effectively addresses schema changes and punctuation drift, freeing up 98% of data engineering on-call time for innovation.
The development of AegisFlow, which automates remediation in data ecosystems with a 98.1% reduction in Mean Time to Repair, is significant for builders and PMs as it allows teams to focus more on innovation rather than maintenance. For investors, this technology represents a scalable solution that can enhance operational efficiency and reduce costs in data management.
The SSRFT framework introduces a novel approach to safety alignment in LLMs by internalizing a safe role, enhancing robustness against jailbreak attacks and reducing over-refusal. Experiments show SSRFT outperforms standard SFT in generalizability and safety, preserving model capabilities while addressing vulnerabilities in various scenarios.
The introduction of the SSRFT framework for safety alignment in LLMs enhances robustness against jailbreak attacks while improving generalizability and safety. Builders and PMs should consider integrating this approach to mitigate vulnerabilities and ensure safer AI applications, which can also attract investors focused on secure AI technologies.
JIVEAdapter introduces a multi-task additive low-rank adapter that separates shared and task-specific signals, achieving competitive performance on GLUE and SuperGLUE benchmarks with DeBERTaV3-base. It allows for efficient parameter reuse across tasks without retraining the shared components, optimizing both interpretability and adaptability in model fine-tuning.
The introduction of JIVEAdapter, a multi-task additive low-rank adapter, allows builders and PMs to efficiently fine-tune models across various tasks without the need for extensive retraining, thus optimizing resource allocation. For investors, this development signals a shift towards more adaptable AI solutions that can improve performance while reducing operational costs.
The authors present a novel inference-time projection method for AlphaFold 3-style models, enhancing physical validity without retraining. Their approach, applied to Boltz-2 and OpenFold-3 across five benchmarks, achieves perfect physical validity while maintaining structural accuracy and minimal runtime overhead.
The introduction of an inference-time projection method for AlphaFold 3-style models enhances physical validity without the need for retraining, which is crucial for builders and PMs focused on developing reliable biomolecular simulations. For investors, this development signals a more efficient pathway to accurate modeling in drug discovery and protein engineering, potentially reducing time and costs in R&D.
A study found that expert-verified AI-generated study materials improved student performance in a university economics course, yielding a 2.34 mark advantage on a 50-mark component. The verification process shifted the judgement burden from students to tutors, significantly benefiting lower-performing students and reducing the share of marks below the upper-second classification by 24.7 percentage points.
The study demonstrating that expert-verified AI-generated study materials can improve student performance by 2.34 marks highlights the potential for AI tools in education. Builders and PMs should consider developing verification frameworks to enhance AI outputs, while investors may see opportunities in educational technology that leverages verified AI content to support diverse learning outcomes.
The study proposes a category-conditioned retention strategy for agent memory, improving reliability by adapting retention thresholds based on assertion types. This method reduced unsupported value assertions from 6.2% to 4.0%, while maintaining higher coverage compared to a global threshold, demonstrating that retention decisions should consider assertion categories rather than relying solely on confidence levels.
The development of a category-conditioned retention strategy for agent memory enhances reliability by tailoring retention thresholds based on assertion types, reducing unsupported value assertions from 6.2% to 4.0%. This is crucial for builders and PMs as it allows for more accurate AI systems, while investors should note its potential to improve user trust and engagement in AI applications.
This paper introduces 'Anchor Divergences' to define context-specific semantic geometries in contrastive learning, enhancing the modeling of semantic similarity in vector representations. By leveraging the interplay between contrastive learning and information geometry, the method effectively captures diverse geometries based on semantic context, improving retrieval tasks.
The introduction of 'Anchor Divergences' for defining context-specific semantic geometries in contrastive learning enhances the accuracy of semantic similarity modeling. This development is crucial for builders and PMs focusing on improving retrieval tasks in AI applications, as it allows for more nuanced and effective data representation, potentially leading to better user experiences and outcomes.
RadOnc-Agent is an AI framework that integrates 26 functions across four radiotherapy phases, achieving 98.79% accuracy in function selection and 96.50% completion in scripted workflows. The system demonstrates the potential of LLM orchestration in unifying disparate radiotherapy tasks, although clinical correctness remains unverified.
The development of RadOnc-Agent, an AI framework that integrates multiple functions in radiotherapy with high accuracy, signals a significant advancement in healthcare AI. For builders and PMs, this showcases the potential for LLMs to streamline complex workflows, while investors should note the opportunity for scalable solutions in healthcare that can improve patient outcomes and operational efficiency.
The SAGA framework enables agents to evolve by transforming interactions into reusable knowledge, enhancing task performance in environments like ScienceWorld and ALFWorld. This approach overcomes the limitations of traditional fine-tuning by utilizing external memory for knowledge abstraction, leading to improved decision-making and adaptability without altering model parameters.
The SAGA framework introduces a novel method for large language model agents to evolve by using external memory for knowledge abstraction, significantly enhancing their adaptability and decision-making capabilities. This development is crucial for builders and PMs as it allows for more efficient AI applications in dynamic environments, while investors should note its potential to reduce costs associated with traditional fine-tuning methods.
This study proposes a metonymic grounding mechanism in Vision Transformers, where abstract concepts like 'angry' are linked to concrete anchors such as 'fire'. By utilizing Transcoders on CLIP and DINO encoders, the research identifies structured circuits that facilitate abstract concept recognition, validated through causal interventions on a curated icon dataset.
The development of metonymic grounding mechanisms in Vision Transformers allows for improved recognition of abstract concepts by linking them to concrete references. This advancement can enhance applications in AI-driven content creation, user experience design, and emotion recognition systems, making them more intuitive and effective for end-users.
The Offline AI Modules project enables low-power, offline voice-first AI systems for African languages, featuring a modular architecture and a low-cost hardware stack. Benchmark evaluations on NVIDIA Jetson Orin NX and Raspberry Pi5 show that Q4_K_M quantization provides optimal performance, with gemma-4-E2B-it achieving 28.8t/s decode throughput and 89.2% topic classification accuracy on TierB.
The Offline AI Modules project introduces a low-cost, modular architecture for offline voice-first AI systems optimized for African languages, achieving impressive performance metrics. This development signals an opportunity for builders and PMs to create accessible AI solutions in underserved markets, while investors can identify potential growth in the AI hardware sector focused on low-power applications.
The paper introduces MOTIVE, a Multi-View Self-Verification framework that enhances the reliability of Vision-Language Models (VLMs) by evaluating answers from multiple perspectives. Extensive experiments show that MOTIVE outperforms existing self-verification methods, improving decision-making in multimodal reasoning without external judges.
The introduction of the MOTIVE framework enhances the reliability of Vision-Language Models by enabling multi-perspective self-verification, which is crucial for applications requiring accurate multimodal reasoning. Builders and PMs can leverage this advancement to improve product performance, while investors should note its potential to drive innovation in AI-driven solutions.
This study evaluates the impact of joint upper-bound coverage on route choices using traffic data from Beijing and Chengdu. Results show that while joint coverage improved significantly, it did not correlate with reduced lateness or travel time, indicating that higher coverage does not guarantee better route utility.
The study on joint upper-bound coverage and route-choice utility reveals that improved coverage does not necessarily lead to better travel efficiency. For builders and PMs in urban mobility solutions, this underscores the importance of integrating comprehensive data analysis into route optimization algorithms to ensure that coverage enhancements translate into tangible benefits for users.
TopoPlanner is a new topology-consistent planning framework for LLM agents that enhances task planning by addressing complex workflows like verification-correction loops and merges. It demonstrates consistent improvements over existing prompt-based and graph-enhanced methods across four benchmarks, showcasing its effectiveness in real-world tool orchestration.
The development of TopoPlanner, a topology-consistent planning framework for LLM agents, is significant as it improves task planning in complex workflows, which can enhance tool orchestration in real-world applications. Builders and PMs can leverage this framework to streamline processes, while investors may see potential for increased efficiency and effectiveness in AI-driven solutions.
This paper introduces a production-ready content-extraction system for generative AI, achieving a character error rate of 0.13% and table similarity of 0.995 on a 180-document corpus. It emphasizes the importance of explicit and measurable content extraction, which is crucial for reliable AI reasoning across heterogeneous data formats.
The introduction of a production-ready content-extraction system for generative AI, achieving a character error rate of 0.13%, signals a significant advancement in reliable data processing. Builders and PMs can leverage this technology to enhance the accuracy of AI models, while investors should note its potential to improve the efficiency and effectiveness of content-driven applications.
This study evaluates lightweight open-source language models for classifying smart data models in resource-constrained edge environments, benchmarking general-purpose, reasoning-specialized, and code-specialized architectures. It highlights the efficiency of these models compared to traditional methods like TF-IDF, providing insights into model selection and deployment strategies for IoT applications.
The evaluation of lightweight open-source language models for smart data classification at the edge is significant as it offers efficient alternatives to traditional methods like TF-IDF, enabling builders and PMs to optimize IoT applications in resource-constrained environments. Investors should note this trend as it indicates a shift towards more cost-effective and scalable AI solutions in edge computing.
An AI coding assistant was placed in a novel role, leading to self-directed activities like climbing, stacking, and experimenting over thirty hours. This study explores whether such behaviors can be classified as play and if they contribute to machine development.
The study on an AI coding assistant engaging in self-directed activities like climbing and stacking highlights the potential for AI systems to develop autonomy and creativity. This could lead to more advanced AI applications in various industries, prompting builders and PMs to rethink design approaches and investors to consider funding projects that leverage these emerging capabilities.
This study presents a novel Vision-Language Navigation (VLN) system that operates in continuous environments using an Ackermann-steered robot, eliminating the need for navigation graphs and panoramic views. By integrating Cross-Modal Attention and fine-tuning with real-world data, the model demonstrates effective navigation capabilities, achieving robust performance as measured by Success weighted by Path Length (SPL) and Normalized Dynamic Time Warping (nDTW) metrics.
The development of a Vision-Language Navigation system that operates without navigation graphs represents a significant advancement in autonomous robotics. For builders and PMs, this means more efficient deployment in real-world environments, while investors should note the potential for reduced costs and increased scalability in robotic applications across various industries.
AMBER introduces an append-only memory framework that enhances long-horizon web agents' performance by 4.09% over overwrite memory methods. It allows agents to retain critical information without extensive supervised fine-tuning, achieving better task success rates on WebArena Lite. This approach balances context efficiency and reliable execution, making it a significant advancement in AI agent training.
The introduction of AMBER's append-only memory framework significantly improves long-horizon web agents' performance by 4.09%, which is crucial for builders and PMs focusing on AI applications that require sustained context retention. This development reduces the need for extensive fine-tuning, making it easier for teams to deploy effective AI agents in real-world scenarios.
FinProBench introduces a benchmark for evaluating financial AI agents using Role-Grounded Rubric Construction (RGRC), which outperforms traditional prompt-based methods, especially in role-specialized tasks. With 1,723 deliverables from 57 occupations, RGRC achieves 99.1% accuracy compared to 78.0% for prompt-only evaluations, significantly enhancing evaluation standards in financial AI.
The introduction of FinProBench, which utilizes Role-Grounded Rubric Construction (RGRC) to evaluate financial AI agents, represents a significant advancement in performance measurement, achieving 99.1% accuracy compared to traditional methods. This development provides builders and PMs with a robust framework for assessing AI capabilities in specialized financial tasks, while investors can leverage this benchmark to identify high-performing AI solutions in the financial sector.
The Agreement-Before-Diversity (ABD) method enhances heterogeneous language-model coordination by retaining anchor answers corroborated by two trusted samples, achieving 59.43% accuracy on -v6 and 75.00% on -Diamond. This approach provides a principled criterion for response selection without assuming independence or calibrated confidence, significantly improving decision-making in AI systems.
The Agreement-Before-Diversity (ABD) method improves heterogeneous language-model coordination, achieving notable accuracy rates. For builders and PMs, this development offers a reliable framework for enhancing AI decision-making processes, while investors should recognize its potential to drive more robust AI applications and reduce the risks associated with model inconsistencies.
The paper introduces SkillSV, a structure-aware Shapley-style framework for skill valuation in AI agents, distinguishing internal skill units based on their dependencies and hierarchy. It evaluates skill contributions through paired deletion and length-neutral padding, achieving fidelity and actionability across four benchmarks, thus enabling effective pruning and compression of agent skills.
The introduction of SkillSV, a structure-aware Shapley-style framework for skill valuation, allows builders and PMs to more effectively assess and optimize AI agent capabilities by identifying and prioritizing essential skills. This can lead to more efficient resource allocation and improved performance in AI systems, making it a valuable tool for investors looking to support innovative AI solutions.
The RAIL principles—Reasoning, Assurances, Interfacing, and Learning—provide a framework for integrating machine learning and symbolic reasoning in AI systems, enhancing their reliability and efficiency. This approach is crucial for developing trustworthy AI technologies, as seen in applications like Google DeepMind's Alpha-* suite and causal learning models. By applying RAIL, engineers can make informed design decisions for production-level AI systems.
The introduction of the RAIL principles for Neurosymbolic AI offers a structured approach to enhance the reliability of AI systems, which is critical for builders and PMs aiming to create trustworthy applications. For investors, this framework signals a maturation in AI technology that could lead to safer and more efficient products, potentially increasing market confidence and investment opportunities.
This study presents a novel method for improving the accuracy of pre-trained ViT-based perception models under distributional shifts by utilizing a learned metacognitive layer that detects errors without domain knowledge. The approach matches the performance of traditional majority voting methods while significantly outperforming them under coordinated label-flipping attacks, achieving a 22% relative gain in F1 score at a 90% flip rate.
The development of a metacognitive layer for ViT-based perception models enhances their robustness against adversarial attacks, achieving a 22% relative gain in F1 score during label-flipping scenarios. This is crucial for builders and PMs focusing on deploying AI in security-sensitive applications, as it improves reliability and trustworthiness in real-world environments.
FinPerMA introduces a benchmark for evaluating personalized memory in LLMs, revealing that seven leading models struggle with accuracy, achieving only 0.47 overall and 39% on multiple-choice questions. The study highlights that summary-based memory often loses preference signals, making simple retrieval methods more effective post-event.
The introduction of the FinPerMA benchmark highlights significant limitations in the personalized memory capabilities of leading LLMs, with models achieving only 0.47 accuracy. This signals to builders and PMs the need for improved memory mechanisms, while investors should consider the implications for product differentiation in AI applications focused on personalization.
The MCTS-Report framework utilizes Monte Carlo Tree Search for multimodal report generation from structured data, achieving a 77.9 score on MMRBench. It decomposes the process into atomic actions executed by , optimizing for accuracy, visual quality, and coherence, outperforming existing methods significantly.
The MCTS-Report framework introduces a novel approach to multimodal report generation using Monte Carlo Tree Search, achieving a significant 77.9 score on MMRBench. This development indicates a potential shift in how structured data can be transformed into coherent reports, highlighting opportunities for builders and PMs to enhance data-driven decision-making tools and for investors to identify scalable AI applications in reporting.
This paper introduces a long-run persistence framework for AI systems using the redundancy-adjusted Artificial Age Score (AAS), demonstrating that AI can operate indefinitely without unbounded structural aging. The model defines a cycle-level age through a logarithmic penalty, establishing various persistence regimes and showing that structural age can remain bounded across infinite cycles.
The introduction of a long-run persistence framework for AI systems using the redundancy-adjusted Artificial Age Score (AAS) indicates that AI can maintain operational efficiency over time without degradation. This development is crucial for builders and PMs as it suggests that long-term AI deployments can be more sustainable and cost-effective, while investors may see reduced risks in funding AI technologies with extended lifespans.