https://vercel.com/blog
DeepSignal tracks AI updates from Vercel AI, filtering research and product signals into plain-English summaries, signal scores and source-linked article pages.
Current topics: Tooling, Open Source, AI Assistant, AI Coding, Agent · Companies: Vercel, Claude, DeepSeek, AWS
High-signal updates

Ling 3.0 Tiny from ANT Group is now available on AI Gateway, free until 8:00 AM PT on August 14. This MOE model features 7.9B parameters, a 256K token context window, and is designed for responsive agents and multi-turn conversations.
The release of Ling 3.0 Tiny by ANT Group on AI Gateway introduces a highly capable MOE model with 7.9B parameters and a 256K token context window, enabling builders and PMs to develop more sophisticated conversational agents. This could lead to enhanced user engagement and retention, making it a significant opportunity for investors looking to back innovative AI solutions.

Vercel has launched Agent Plugins 1.0.0, a vendor-neutral standard for packaging AI agent skills and servers, allowing for easier integration across clients like ChatGPT and GitHub Copilot. This format simplifies the discovery and loading of components while maintaining client-specific flexibility through a namespaced extension mechanism.
Vercel's launch of Agent Plugins 1.0.0 introduces a vendor-neutral standard for AI agent skills, which simplifies integration across platforms like ChatGPT and GitHub Copilot. This development allows builders to create more interoperable AI solutions, enabling PMs to streamline product features and offering investors insight into a growing ecosystem of adaptable AI technologies.

Vercel AI Gateway now generates OpenTelemetry traces for each request, allowing Pro and Enterprise teams to send data to OTLP/HTTP endpoints via Vercel Drains. Each trace includes detailed metrics like token usage and request duration, costing $0.05 per 1,000 traces plus standard data transfer fees.
The introduction of OpenTelemetry traces via Vercel Drains allows Pro and Enterprise teams to gain deep insights into their AI application's performance, including token usage and request duration, which can inform optimizations and cost management. This development is crucial for builders and PMs aiming to enhance user experience and for investors assessing the operational efficiency of AI products.

AI Gateway is now available on AWS Marketplace, allowing teams to streamline AI model procurement through their AWS accounts. It offers a single API for multiple models with built-in reliability and governance features, while maintaining provider pricing without markup.
The availability of AI Gateway on AWS Marketplace simplifies the procurement of AI models for builders and PMs by providing a unified API with governance features, enhancing operational efficiency. For investors, this development signals a growing market for streamlined AI solutions, potentially increasing the attractiveness of companies leveraging this infrastructure.

Vercel AI has launched the new v0 API, enabling programmatic access to its app-building agent, allowing users to create apps from prompts, manage isolated workspaces, and deploy seamlessly to Vercel. The API is currently in public beta and supports synchronous, asynchronous, and streaming responses.
The launch of Vercel AI's new v0 API allows builders and PMs to programmatically create and manage apps from prompts, significantly streamlining the app development process. For investors, this indicates a growing trend towards automation in software development, potentially increasing the efficiency and scalability of projects built on the Vercel platform.

Meta's Muse Spark 1.2 is now available on Vercel AI Gateway, enhancing code generation, debugging, and developer workflows. It supports long-horizon tasks like generating entire repositories and iterative code improvements, with a unified API for tracking usage and costs.
The release of Meta's Muse Spark 1.2 on Vercel AI Gateway significantly enhances code generation and debugging capabilities, allowing developers to automate long-horizon tasks like generating entire repositories. This improvement can streamline workflows and reduce development time, making it a valuable tool for builders and PMs focused on efficiency and cost management in software projects.

Vercel has introduced full Sandbox egress firewall features on its Hobby plan, enabling network isolation for free-tier users. This allows developers to control outbound requests securely, attaching secrets to requests without exposing them in code, thus enhancing security for untrusted or AI-generated code.
Vercel's introduction of full Sandbox egress firewall features on the Hobby plan allows developers to securely manage outbound requests, enhancing the security of applications that utilize untrusted or AI-generated code. This development is significant for builders and PMs as it lowers the barrier for implementing robust security practices, while investors can see increased adoption of Vercel’s platform among developers focused on secure coding practices.

Vercel AI introduces skill packs on skills.sh, allowing users to bundle multiple agent skills into a single shareable URL. Teams can standardize skills across projects by creating packs from community skills or local sources, with easy installation and updates via command line.
Vercel AI's introduction of skill packs on skills.sh allows teams to standardize and share multiple agent skills through a single URL, streamlining collaboration and enhancing project efficiency. This development is significant for builders and PMs as it simplifies skill management and integration, potentially reducing development time and improving product consistency.

DeepSeek v4 Flash is currently available at a 90% discount for Vercel Pro customers on AI Gateway until August 11. Users can access this deal by routing requests through Novita, with fallback options to other providers at standard rates if necessary.
The 90% discount on DeepSeek V4 Flash for Vercel Pro customers through AI Gateway represents a significant cost-saving opportunity for builders and PMs looking to integrate advanced AI capabilities into their applications. For investors, this development indicates a competitive pricing strategy that could drive user adoption and market share for AI solutions.

Qwen 3.8 Max, with 2.4 trillion parameters and a 1 million token context window, is now available on Vercel AI Gateway. This model excels in text-only and vision-language tasks, making it ideal for software engineering and visual work, such as transforming screenshots into functional pages and video captioning. Users can access it via the model playground and integrate it with coding agents.
The availability of Qwen 3.8 Max on Vercel AI Gateway, featuring 2.4 trillion parameters and a 1 million token context window, significantly enhances capabilities for builders and PMs in developing advanced applications that require complex text and visual processing. This development allows for more sophisticated automation in software engineering tasks, such as converting screenshots into functional code, which can streamline workflows and reduce development time.

Vercel AI Gateway now allows budget management at team and project levels, in addition to individual API keys. Users can set spend limits, track usage, and receive alerts when nearing budget thresholds, ensuring better financial control over API requests.
Vercel AI Gateway's new feature for team and project budget management allows builders and PMs to set spend limits and track usage, enhancing financial control over API costs. This development is crucial for investors as it indicates a shift towards more responsible and scalable usage of AI resources, potentially leading to improved ROI and reduced financial risk.

DeepSeek V4 Flash now operates with updated weights on AI Gateway, achieving a score of 82.7 on , a significant increase of 25.8 points. Users can access these new weights automatically without changing the model ID or code, with other providers expected to follow next week.
The update of DeepSeek V4 Flash with new weights on AI Gateway, achieving a score of 82.7 on Terminal-Bench, indicates a significant performance boost that builders and PMs can leverage for improved application efficiency. For investors, this development suggests a competitive edge in AI capabilities, potentially leading to better market positioning and returns.

Vercel AI's AI Gateway now features a dedicated Logs page that displays all requests with detailed metrics such as cost, token counts, and routing information. Users can filter and search requests by various parameters, export data as CSV or JSON, and drill down into request details, including fallback paths and provider attempts.
The introduction of a dedicated Logs page for Vercel AI's AI Gateway allows builders and PMs to monitor and analyze request metrics in real-time, enhancing debugging and optimization processes. This feature can lead to better resource management and cost efficiency, making it a valuable tool for investors assessing the platform's operational effectiveness.

The Chat SDK for Microsoft Teams now supports reactions and ephemeral messages, allowing bots to interact more dynamically. Bots can add/remove reactions and send targeted messages visible only to specific users, enhancing user engagement and privacy.
The addition of reactions and ephemeral messages in the Chat SDK for Microsoft Teams enables developers to create more interactive and personalized bot experiences. This enhancement can lead to increased user engagement and satisfaction, making it a valuable feature for PMs and investors focusing on improving communication tools in enterprise settings.

Laguna S 2.1 from Poolside now offers 10x more capacity on AI Gateway for both paid and free versions, enhancing support for high-volume coding tasks. Users can configure the model via the AI SDK and connect coding agents to the Gateway for improved performance and tracking.
The Laguna S 2.1's 10x increased capacity on AI Gateway significantly enhances its utility for developers handling high-volume coding tasks, allowing for better performance and tracking. This development signals a shift towards more scalable AI solutions, which could attract investment and drive product innovation in the coding space.

Vercel AI has launched MiniMax H3 on its AI Gateway, enabling 2K video generation from text prompts, images, and multimodal references. The model supports various aspect ratios and durations, allowing for versatile video creation tailored to user specifications.
The launch of MiniMax H3 on Vercel AI Gateway allows for efficient 2K video generation from diverse inputs, which can significantly enhance content creation workflows for builders and PMs. Investors should note this development as it opens new monetization avenues in the growing AI-driven media landscape.

Inkling Small from Thinking Machines is now available on AI Gateway, offering performance similar to larger models at a quarter of the size. It excels in reasoning, coding, and image processing while allowing users to control the trade-off between quality, cost, and latency. This model is also compatible with Zero Data Retention for enhanced privacy.
The release of Inkling Small by Thinking Machines on AI Gateway offers builders and PMs a compact model that balances performance and resource efficiency, enabling faster deployment in applications requiring reasoning and image processing. For investors, this development signals a growing trend towards smaller, privacy-focused AI solutions that can reduce operational costs while maintaining competitive capabilities.

AI Gateway has announced significant updates for GPT-5.6 models: Luna sees an 80% price reduction to $0.2 per million tokens, Terra is reduced by 20% to $2 per million tokens, and Sol's fast mode is now 2.5x faster without a price change. These adjustments apply to both short and long context pricing, benefiting users without requiring code changes.
The significant price reduction for GPT-5.6 models, particularly Luna's 80% cut to $0.2 per million tokens, enables builders and PMs to reduce operational costs and scale AI applications more affordably. Investors should note that these enhancements could lead to increased adoption and competitive advantages for businesses leveraging these models.

AI Gateway introduces a unified fast mode in beta, allowing users to request faster processing for models like 'anthropic/claude-opus-5' by simply setting speed to 'fast'. This mode enhances performance at a higher cost per token, with fallback to standard speed when necessary.
The introduction of AI Gateway's unified fast mode allows builders and PMs to optimize model performance by easily toggling between speed settings, which can enhance user experience for time-sensitive applications. For investors, this development indicates a growing demand for flexible AI solutions that can cater to varying processing needs, potentially leading to increased adoption and revenue opportunities.

Grok Voice Think Fast 2.0 by xAI is now available on AI Gateway, featuring enhanced reasoning, transcription accuracy, and real-time conversation capabilities. This speech-to-speech model processes queries without latency and requires fewer reasoning tokens, improving response times. It performs well even in challenging audio conditions, making it suitable for various applications.
The release of Grok Voice Think Fast 2.0 on AI Gateway significantly enhances speech-to-speech capabilities with improved reasoning and transcription accuracy. This advancement allows builders and PMs to integrate faster and more reliable voice interactions into their applications, while investors can recognize the potential for increased adoption in sectors requiring robust audio processing.

AI Gateway now enables regional inference, allowing users to pin requests to either the US or EU. This simplifies compliance for teams by ensuring data residency, with requests failing if no provider can serve the selected region. The feature incurs a potential cost increase of about 10% above standard rates.
The introduction of regional inference on AI Gateway allows builders and PMs to ensure compliance with data residency regulations in the US and EU, reducing legal risks. However, the potential 10% cost increase may impact budget planning for projects reliant on this feature, making it crucial for investors to assess financial implications.

Sandstone achieved a remarkable 40x revenue growth in just 147 days by leveraging Vercel's AI SDK and secure compute capabilities, automating legal workflows and enhancing responsiveness. With over 1,000 legal requests managed daily, their innovative platform integrates seamlessly across multiple systems, allowing lawyers to access all necessary information instantly.
Sandstone's 40x revenue growth in 147 days, driven by Vercel's AI SDK, highlights the potential of integrating AI to automate workflows in legal tech. This signals to builders and PMs the importance of leveraging advanced tools for scalability, while investors should note the rapid market validation of AI-driven solutions in traditional industries.

OpenAI's DeepsecBench evaluates model performance in identifying cybersecurity vulnerabilities, revealing GPT-5.6 Sol as the top performer with a score of 35.58 at a cost of $55.98. The benchmark highlights the growing efficiency of various models, making comprehensive scanning more accessible and cost-effective for developers.
The introduction of DeepsecBench, which highlights GPT-5.6 Sol as the leading model for identifying cybersecurity vulnerabilities, signals a significant advancement in automated security assessments. For builders and PMs, this means enhanced tools for vulnerability detection can streamline development processes, while investors should note the potential for reduced costs and increased efficiency in cybersecurity solutions.

AI Gateway now supports WebSocket for the OpenAI Responses API, enabling persistent connections that enhance performance by up to 40% for agentic rollouts with multiple tool calls. Developers can send incremental inputs with previous response IDs instead of full context, streamlining interactions.
The introduction of WebSocket support for the OpenAI Responses API on AI Gateway allows developers to create more efficient applications by maintaining persistent connections, which can enhance performance by up to 40%. This development is crucial for builders and PMs as it streamlines interactions and reduces latency during agentic rollouts, ultimately leading to better user experiences and faster deployment times.

Moonshot AI's Kimi K3 and Kimi K3 Fast are now available via US-based providers on AI Gateway, supporting Zero Data Retention (ZDR). Kimi K3 Fast offers lower latency at a ~50% higher cost per token, while the AI Gateway ensures automatic failover and optimal throughput across multiple providers.
The availability of Kimi K3 and Kimi K3 Fast on AI Gateway with Zero Data Retention (ZDR) is significant for builders and PMs as it allows for enhanced data privacy and lower latency options for applications. Investors should note the competitive pricing structure, as Kimi K3 Fast offers faster performance at a premium, indicating a growing market for high-performance AI solutions.

Vercel AI introduces Claude Managed Agents with Chat SDK, enabling seamless chat interfaces across platforms like Slack and WhatsApp. The solution features token-by-token streaming, a live activity feed, and eliminates the need for a separate database, making it easy to deploy and manage conversations.
Vercel AI's introduction of Claude Managed Agents with Chat SDK allows builders and PMs to easily integrate advanced chat functionalities into existing platforms like Slack and WhatsApp without the overhead of a separate database. This streamlining reduces development time and costs, making it an attractive option for investors looking to support scalable communication solutions.

Claude Opus 5 by Anthropic is now available on AI Gateway, enhancing long-horizon coding with improved multi-file handling and effective team coordination. It operates efficiently at low to medium effort levels, offering robust visual analysis and cybersecurity features, while supporting Zero Data Retention and fast mode options.
The release of Claude Opus 5 on AI Gateway enhances long-horizon coding capabilities, which is crucial for builders and PMs managing complex projects. Its improved multi-file handling and cybersecurity features signal a shift towards more efficient team collaboration and secure coding practices, making it a valuable tool for investors looking to support innovative software development solutions.

Ling 3.0 Flash, a 124B parameter Mixture-of-Experts model from Ant Group, is now available for free on AI Gateway for three weeks. It features a 256K token context window and is optimized for high-frequency workflows and multi-step agent interactions.
The release of Ling 3.0 Flash, a 124B parameter Mixture-of-Experts model with a 256K token context window, offers builders and PMs a powerful tool for developing applications that require extensive context and complex interactions. This could significantly enhance user experience in high-frequency workflows, making it a valuable asset for investors looking to support innovative AI solutions.

Vercel now enables WebSocket connections for Python applications, supporting real-time features in frameworks like FastAPI, Django, and Flask. This allows bidirectional communication, enhancing interactive AI streaming and collaboration capabilities for developers deploying on Vercel Functions.
Vercel's introduction of WebSocket support for Python Functions enables real-time bidirectional communication in frameworks like FastAPI, Django, and Flask. This development is significant for builders and PMs as it enhances the interactivity of AI applications, allowing for more dynamic user experiences and potentially increasing user engagement and retention.

Vercel MCP now enables direct code deployment to projects, allowing AI assistants to ship builds and return shareable URLs seamlessly. Users can connect to supported clients like Claude or Cursor and initiate deployments using the deploy_to_vercel tool, which automates project creation and dependency management.
Vercel's MCP now allows direct code deployment through AI assistants, streamlining the development process by automating project creation and dependency management. This development enables builders and PMs to enhance productivity and reduce time-to-market, while investors can recognize the potential for increased efficiency in software delivery and the competitive edge it provides.