https://www.together.ai/blog
DeepSignal tracks AI updates from Together AI, filtering research and product signals into plain-English summaries, signal scores and source-linked article pages.
Current topics: Infrastructure, AI Startup, Inference, LLM, Open Source · Companies: Claude, DeepSeek, NVIDIA
High-signal updates

ThunderAgent enhances synthetic data generation by achieving 2.5× higher single-node throughput and 2.4× speedup across 8 nodes, effectively mitigating KV cache thrashing in agentic inference workflows. This system is crucial for training like Claude Code and Codex that require high concurrency in multi-turn tasks.
The development of ThunderAgent, which achieves 2.5× higher throughput for synthetic data generation, is significant for builders and PMs as it allows for more efficient training of large language models like Claude Code and Codex. This efficiency can lead to reduced costs and faster time-to-market for AI applications that rely on high concurrency in multi-turn tasks.

Together AI partners with Moonshot AI to launch Kimi K3, a 2.8T parameter sparse MoE model, providing developers immediate access to cutting-edge open models with zero data retention. This collaboration enhances production-ready infrastructure and allows for custom training, ensuring high performance and scalability for various applications.
The partnership between Together AI and Moonshot AI to launch the Kimi K3 model, a 2.8T parameter sparse MoE model, provides developers with immediate access to advanced AI models without data retention concerns. This enhances the infrastructure for custom training, allowing builders and PMs to scale applications efficiently while attracting investors looking for innovative, production-ready solutions.

Together AI's Dedicated Model Inference architecture integrates endpoints, deployments, and configs, enabling capacity-aware traffic routing for efficient model inference. This setup supports advanced features like A/B testing and zero-downtime changes, ensuring optimal resource utilization and traffic management.
Together AI's Dedicated Model Inference architecture allows for efficient model deployment with capacity-aware traffic routing, which is crucial for builders and PMs to optimize resource utilization and manage traffic effectively. This development also supports advanced features like A/B testing and zero-downtime changes, making it a valuable tool for investors looking to back scalable AI solutions.

Kimi K3 outperforms GPT-5.6 Sol in pass@4 (89.4% vs 85.8%) and cost per rollout ($4.65 vs $8.37), while Sol leads in pass@1 (72.7% vs 68.5%). Both models excel in different programming languages, making them complementary for task routing.
The comparison between Kimi K3 and GPT-5.6 Sol highlights the strengths of each model, with Kimi K3 offering better cost efficiency and higher performance in multi-task scenarios. Builders and PMs can leverage this information to optimize task routing and resource allocation, while investors may identify opportunities in integrating complementary AI solutions for enhanced productivity.

Kimi K3, an from Moonshot AI, offers competitive performance at $4.65 per rollout, just behind Claude Fable 5's $13.41, achieving 68.5% pass@1 versus Fable's 69.9%. With superior pass@2 and pass@4 scores, Kimi K3 is a cost-effective choice for coding tasks.
The competitive performance of Kimi K3 at $4.65 per rollout compared to Claude Fable 5's $13.41 highlights a cost-effective option for coding tasks. Builders and PMs should consider integrating Kimi K3 to optimize budget while maintaining high performance, making it an attractive choice for investors focused on cost efficiency in AI solutions.

Together AI's new inference platform enables complete control over open-weight models, optimizing performance and cost while allowing for rapid deployment and iteration. With support for various model sizes, it significantly reduces startup times, achieving up to 4× faster warm starts for large models like Qwen 3 and Llama 3.3, enhancing user experience and operational efficiency.
Together AI's new inference platform allows for rapid deployment and optimization of open-weight models, achieving up to 4× faster warm starts for large models like Qwen 3 and Llama 3.3. This development is crucial for builders and PMs as it enhances user experience and operational efficiency, while investors should note the potential for reduced costs and improved performance in AI applications.

Together AI and Y Combinator have launched the first dedicated GPU cluster for YC's AI-native startups, addressing the compute bottleneck that many face. Startups can now access flexible, cost-effective GPU resources without long-term commitments, enabling rapid scaling and innovation.
The launch of a dedicated GPU cluster by Together AI and Y Combinator provides AI startups with flexible and cost-effective computing resources, addressing a critical bottleneck in scaling their products. This development enables builders and PMs to innovate faster and allows investors to support companies with reduced operational constraints.

Together AI emphasizes that achieving 99.9% uptime for inference requires robust architecture and continuous traffic management across multiple facilities, as failures can occur at various layers including compute, network, and storage. Their experience with clients like Cursor and Decagon highlights the complexities of maintaining reliability in GPU-based systems, especially during outages.
Together AI's emphasis on achieving 99.9% uptime for inference highlights the critical need for robust architecture and traffic management in AI systems. For builders and PMs, this signals the importance of investing in reliable infrastructure to avoid costly outages, while investors should recognize the potential for companies that can deliver high availability in GPU-based applications.

Thinking Machines Lab's new , Inkling, is now available on Together AI, offering token-efficient reasoning and support for text, image, and audio inputs. With 975B parameters and strong performance across various benchmarks, it allows developers to adjust inference effort for optimized task execution.
The launch of Inkling, a multimodal model by Thinking Machines Lab on Together AI, offers builders and PMs a powerful tool for developing applications that require efficient reasoning across text, image, and audio inputs. Its 975B parameters and adjustable inference effort can significantly enhance performance while optimizing resource usage, making it attractive for investors looking at scalable AI solutions.
Together AI has enhanced its GPU Clusters with passive health checks and auto node repair, improving reliability and operational control for large-scale training and inference. The new features, including Slinky 1.0, aim to minimize downtime and streamline cluster management as organizations scale.
Together AI's introduction of passive health checks and auto node repair in their GPU Clusters enhances reliability and operational control, which is crucial for builders and PMs managing large-scale AI projects. This development minimizes downtime, allowing teams to focus on innovation rather than maintenance, making it an attractive proposition for investors looking at scalable AI infrastructure.

Together AI introduces Provisioned Throughput, offering guaranteed inference capacity for MiniMax M3 and GLM-5.2 at $0.05 per PTU per minute, achieving costs up to 90% lower than Claude Opus 4.8. This new model provides predictable pricing and a 99% uptime SLA, catering to companies transitioning to open weight models for production workloads.
Together AI's introduction of Provisioned Throughput for MiniMax M3 and GLM-5.2 at a significantly lower cost provides builders and PMs with a reliable and affordable option for scaling AI applications. This predictable pricing model and 99% uptime SLA enable companies to confidently transition to open weight models, reducing operational risks and costs.
.png)
Together AI has secured $800M in Series C funding to promote open-source AI, highlighting the unsustainable economics of proprietary models. The investment aims to drive innovation and accessibility in AI technologies, addressing the limitations of closed systems.
Together AI's $800M Series C funding signals a significant shift towards open-source AI, emphasizing the increasing challenges of proprietary models. For builders and PMs, this development suggests a growing ecosystem for collaborative innovation, while investors should note the potential for more sustainable and accessible AI solutions in the market.

Together AI presents nine papers at ICML 2026, showcasing advancements from frontier agents like DSGym and ThunderAgent to kernel optimizations, emphasizing a holistic approach across the AI stack. Noteworthy achievements include a 3.6x increase in agent throughput and state-of-the-art discoveries in various fields using open models.
Together AI's presentation of nine papers at ICML 2026 highlights significant advancements in agent throughput, specifically a 3.6x increase, which indicates improved efficiency for AI applications. This development signals to builders and PMs the potential for enhanced performance in deploying AI solutions, while investors may see opportunities in technologies that leverage these optimizations for competitive advantage.

ParallelKernelBench evaluates LLMs' ability to generate efficient multi-GPU CUDA kernels across 87 workloads. While the best model manages to solve less than a third of the tasks effectively, some generated kernels outperform existing public implementations, highlighting the potential for improvement in LLM capabilities.
The evaluation of LLMs in generating efficient multi-GPU CUDA kernels reveals that while current models struggle, some outputs show promise by outperforming existing implementations. This indicates a potential area for investment and development in AI-driven programming tools, which could significantly enhance productivity in high-performance computing applications.

Kimi K2.7 Code generated 12 landing pages at a cost 94% lower than Claude Fable 5, with comparable performance. This significant cost reduction highlights Kimi's efficiency in landing page creation, impacting businesses seeking budget-friendly solutions.
The development of Kimi K2.7 Code, which generates landing pages at 94% lower cost than Claude Fable 5 while maintaining performance, is crucial for builders and PMs focused on cost efficiency. This advancement allows businesses to allocate resources more effectively, potentially increasing ROI and enabling faster iteration on marketing strategies.
Together AI has achieved ISO 27001:2022 certification, validating its Information Security Management System for secure AI workloads. This certification enhances governance around sensitive data, access control, and incident response, ensuring robust protection for customers using its global platform.
Together AI's achievement of ISO 27001:2022 certification signifies a commitment to robust information security for AI workloads, which is crucial for builders and PMs focusing on enterprise solutions. For investors, this certification enhances the company's credibility and marketability, indicating a lower risk profile in handling sensitive data.
MiniMax's M3 model introduces a 1M-token context and multimodal capabilities, optimized for efficient inference with a 9x speedup in prefill and 15x in decoding, supported by Together AI's cloud infrastructure.
The introduction of MiniMax's M3 model with 1M-token context and multimodal capabilities allows builders and PMs to create more complex and contextually aware applications, significantly enhancing user experience. For investors, the 9x speedup in prefill and 15x in decoding represents a critical advancement in AI efficiency, indicating potential for higher returns in scalable AI solutions.
Together AI developed the fastest speech-to-text stack, achieving 20 hours of transcription in under 10 seconds using NVIDIA's Parakeet-TDT 0.6B v3 and OpenAI's Whisper Large v3. Key optimizations included TensorRT for encoder efficiency and GPU-based decoding, resulting in a 2-3x faster performance.
Together AI's development of the world's fastest speech-to-text stack, capable of transcribing 20 hours of audio in under 10 seconds, signals a significant leap in real-time transcription capabilities. This advancement can enhance applications in accessibility, customer service, and content creation, making it a critical consideration for builders and PMs focused on integrating efficient speech recognition into their products.
Together AI's new benchmark for coding agents reveals that its Together Inference Engine achieves 31% higher TPS than TensorRT-, maintaining under 1s TTFT at 625 TPM per GPU. This performance is crucial for handling high concurrency and long context requests in production environments.
Together AI's new benchmark for coding agents, demonstrating a 31% higher transactions per second (TPS) than TensorRT-LLM, is significant for builders and PMs as it indicates improved performance for high-concurrency applications. For investors, this development suggests a competitive edge in the market, potentially leading to increased adoption and revenue opportunities in AI-driven solutions.
Violin is an open-source video translation tool by Together AI, utilizing Whisper V3 for ASR, Deepseek V4 Pro for translation, and Cartesia’s Sonic 3 for TTS, enabling high-quality multilingual video accessibility. It features an interactive chat assistant for user queries and is designed for content creators and developers alike.
Violin's open-source video translation tool leverages advanced AI technologies to enhance multilingual accessibility for video content. This development signals a growing demand for tools that enable creators to reach broader audiences, presenting opportunities for builders and PMs to integrate similar capabilities into their products, while investors can identify potential market growth in the video translation sector.
Together AI's Voice Finder tool allows developers to quickly search over 600 voices across multiple TTS models, including MiniMax and Deepgram. By using prompts or audio samples, users can find suitable voices based on 15+ metadata attributes, streamlining the process of selecting the right voice for applications like fintech support or meditation guides.
Together AI's Voice Finder tool enables developers to efficiently select from over 600 TTS voices using metadata attributes, significantly reducing the time spent on voice selection for applications like fintech and meditation. This development is crucial for builders and PMs looking to enhance user experience through personalized voice interactions, while investors can see potential for increased adoption of voice technology in various sectors.
DeepSeek-V4 transforms million-token context into a serving-systems challenge, as explored by Together AI on NVIDIA HGX B200. Key innovations include compressed KV layouts, prefix caching, and optimized kernel maturity for efficient long-context inference workloads.
The development of DeepSeek-V4, which addresses million-token context as a serving-systems challenge, is significant for builders and PMs as it highlights the need for optimized infrastructure in AI applications. Investors should note that innovations like compressed KV layouts and prefix caching can lead to more efficient long-context inference, potentially enhancing the performance and scalability of AI systems.
Deploy any Hugging Face model effortlessly with Goose and Together's Dedicated Container Inference. This solution allows users to run models in a production-grade GPU environment with just one prompt, eliminating setup complexities and enabling immediate deployment on release day.
The launch of Goose and Together's Dedicated Container Inference allows builders and PMs to deploy Hugging Face models with minimal setup, streamlining the path from development to production. This efficiency can significantly reduce time-to-market for AI applications, making it a crucial development for investors looking for scalable solutions in the AI space.
As AI transitions from research to production, the focus for AI-native teams is shifting towards efficient and reliable model deployment at scale. This involves overcoming challenges related to resource management and performance optimization to ensure models operate effectively in real-world applications.
The shift towards efficient inference at scale is critical for builders and PMs as it directly impacts the feasibility of deploying AI models in real-world applications, requiring advancements in resource management and performance optimization. For investors, this development signals potential growth opportunities in companies that can successfully navigate these challenges and deliver scalable AI solutions.
Together AI swiftly mitigated the Copy Fail vulnerability (CVE-2026-31431) by disabling the algif_aead interface across its infrastructure, preventing potential privilege escalation and cross-tenant risks in AI workloads. This proactive measure ensured minimal operational impact while maintaining security in shared kernel environments.
The swift mitigation of the Copy Fail vulnerability (CVE-2026-31431) by Together AI highlights the importance of proactive security measures in AI infrastructure. Builders and PMs should prioritize security in shared environments to prevent privilege escalation risks, while investors should recognize that robust security practices can enhance the overall reliability and trustworthiness of AI systems.
Together AI partners with Adaption to integrate Together Fine-Tuning into Adaptive Data, enhancing model training efficiency. This collaboration enables users to optimize datasets and achieve an average 82% increase in data quality, facilitating the deployment of fine-tuned models on Together AI's infrastructure.
The partnership between Together AI and Adaption to integrate Together Fine-Tuning into Adaptive Data significantly enhances model training efficiency, achieving an average 82% increase in data quality. This development is crucial for builders and PMs as it streamlines the deployment of fine-tuned models, ultimately reducing time-to-market and improving product performance.
DeepSeek-V4 Pro is now available on Together AI, offering 1.6T-parameter MoE with 512K context for serverless inference. It supports three reasoning modes and reduces costs for repeated long-context queries, with input pricing at $2.10 per million tokens.
The release of DeepSeek-V4 Pro on Together AI, featuring 1.6T parameters and 512K context for serverless inference, significantly lowers costs for applications requiring long-context queries. This development allows builders and PMs to create more efficient AI solutions while investors can recognize potential cost savings and scalability in AI deployment.
Together AI has launched the NVIDIA Nemotron 3 Nano Omni, a model that combines video, audio, images, and language reasoning. This model, utilizing a hybrid Mamba-Transformer architecture, allows developers to build agentic applications with high efficiency and low latency, streamlining deployment from prototype to production without infrastructure management.
The launch of the NVIDIA Nemotron 3 Nano Omni by Together AI enables developers to create multimodal AI applications with improved efficiency and reduced latency. This development allows builders and PMs to streamline their deployment processes, while investors can recognize the potential for scalable solutions in the rapidly evolving AI landscape.
Together AI's Distribution-aware Speculative Decoding (DAS) framework accelerates RL rollouts by over 50% without altering model outputs, addressing the rollout bottleneck in reinforcement learning. This improvement is crucial for large models like DeepSeek-R1, which experience significant delays during the rollout phase, consuming 70% of total training time.
Together AI's Distribution-aware Speculative Decoding (DAS) framework significantly speeds up reinforcement learning rollouts by over 50%, directly addressing a major bottleneck in training efficiency. This development is crucial for builders and PMs working on large models, as it reduces training time and costs, while investors should note its potential to enhance product competitiveness and accelerate deployment timelines.
Multi-tenant GPU clusters enable AI-native teams to share resources efficiently while maintaining isolation, preventing idle capacity and ensuring predictable access. This architecture supports pooled economics without chaos, allowing teams to operate as if they have dedicated clusters.
The development of multi-tenant GPU clusters allows AI-native teams to optimize resource usage while ensuring isolation and predictable access. This architecture can significantly reduce operational costs and improve efficiency, making it a crucial consideration for builders, PMs, and investors focused on scalable AI solutions.