https://www.latent.space/
DeepSignal tracks AI updates from Latent Space, filtering research and product signals into plain-English summaries, signal scores and source-linked article pages.
Current topics: Community, AI Startup, LLM, Open Source, AI Coding · Companies: OpenAI, Claude, Grok, xAI
![[AINews] Jeff, Sanjay, Oriol, and Quoc depart DeepMind; Demis to Chair; Koray to SVP — what is going on at GDM???](https://substackcdn.com/image/fetch/$s_!1S7v!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57c9d273-aa50-4bc7-a7c4-caaf9b4952c2_1540x1624.png)
A major leadership shakeup at Google DeepMind sees Jeff Dean, Sanjay Ghemawat, Oriol Vinyals, and Quoc Le depart to found Discovery Loop, a startup focused on automating machine learning and research. Meanwhile, Demis Hassabis transitions to Chair and Chief Scientist, with Koray Kavukcuoglu stepping up as SVP, signaling a strategic pivot for Google’s AI initiatives.
The departure of key figures like Jeff Dean and Quoc Le from DeepMind to form Discovery Loop indicates a potential shift in the AI landscape, as their focus on automating machine learning could lead to new tools and methodologies that builders and PMs can leverage. This change may also influence investment strategies as investors seek to back emerging technologies that stem from this leadership transition.

OpenAI's ChatGPT Work, launched on July 9, 2026, integrates knowledge work across platforms, achieving over 10 million users in three weeks. It merges functionalities from ChatGPT and Codex, operating on a cloud-based microVM to produce interactive documents and manage tasks seamlessly across devices.
OpenAI's launch of ChatGPT Work, which rapidly gained over 10 million users, signifies a major shift in how knowledge work can be integrated across platforms. Builders and PMs should note the potential for creating applications that leverage this seamless task management and document interaction, while investors may see opportunities in the growing demand for AI-driven productivity tools.
![[AINews] Qwen 3.8 Max(2.4T) and 27B, new open weights models for Coding and Cowork](https://substackcdn.com/image/fetch/$s_!W0RB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fpbs.substack.com%2Fmedia%2FHOw1F2ebgAAkpJJ.jpg)
Alibaba's Qwen 3.8 Max, a 2.4T parameter model, is set to release open weights next week, showcasing capabilities in autonomous coding, chip design, and multimodal reasoning. It achieved top 13% performance in a competitive data science challenge, with API pricing at $2/M input and $6/M output tokens.
Alibaba's release of Qwen 3.8 Max, a 2.4T parameter model with open weights, signifies a major advancement in AI capabilities for coding and multimodal reasoning. Builders and PMs can leverage this model for cost-effective solutions in software development, while investors should note its competitive performance and potential for monetization through API usage.

Baseten has raised $13B, emerging as a leading AI infrastructure decacorn, while Philip Kiely and Ali Taha discuss the evolution of inference engineering, highlighting how quantization can enhance throughput by 20% without sacrificing model quality. Their insights reveal that inference is now a distinct engineering discipline critical for optimizing AI models.
Baseten's emergence as a $13B AI infrastructure decacorn signals a robust market demand for optimized AI solutions. The discussion on quantization in inference engineering highlights a critical opportunity for builders and PMs to enhance model performance and efficiency, which can lead to significant cost savings and improved user experiences.
![[AINews] not much happened today](https://substackcdn.com/image/fetch/$s_!1adH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fpbs.substack.com%2Fmedia%2FHOiugSLbQAAck-3.jpg)
DeepSeek's V4-Flash 0731 API launch marks a significant leap in performance, achieving a score of 82.7 without architectural changes, while offering a competitive cost of $0.28 per million output tokens. The model's open weights are now available, enhancing integration into existing coding stacks and positioning it firmly on the Pareto frontier against proprietary systems.
DeepSeek's launch of the V4-Flash 0731 API, achieving a Terminal-Bench score of 82.7 at a low cost of $0.28 per million tokens, presents a significant opportunity for builders and PMs to integrate high-performance AI into their applications without heavy investment. The availability of open weights further facilitates seamless integration, making it a competitive alternative to proprietary solutions.

Frank Coyle at AIEWF 2026 emphasized the revival of ontologies as essential 'logical guardrails' for effective AI agents, integrating them with for better reasoning. Neo4j's Emil Eifrem highlighted three ontology types to enhance agent scalability, while Kingsley Idehen discussed the challenges and benefits of maintaining ontologies in AI systems.
The revival of ontologies as 'logical guardrails' for AI agents, as discussed at AIEWF 2026, signals a shift towards more structured and scalable AI systems. Builders and PMs should consider integrating these frameworks to enhance reasoning capabilities in their applications, while investors may find opportunities in companies focused on ontology development for AI scalability.
AI is increasingly integrated into finance, with major players like OpenAI and Anthropic launching tools for investment banking and corporate finance. The second annual AIE NYC will focus on AI in Finance, showcasing innovations from companies like Nubank and Morgan Stanley, emphasizing the need for governance and security in AI applications.
The launch of AI tools by major finance players like OpenAI and Anthropic signals a significant shift in investment banking and corporate finance, highlighting the need for builders and PMs to prioritize governance and security in their AI applications. This trend presents new opportunities for investors to back innovative solutions that address these emerging challenges in the financial sector.
![[AINews] Fearing RSI: OpenAI, Anthropic, GDM, Meta, Thinky cosign letter to "Pace" AI development, as HuggingFace details Machine-Speed Offensive Cyberattack](https://substackcdn.com/image/fetch/$s_!u8gQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87c30d45-da74-47aa-80c6-77903e4c9cbf_2094x1890.png)
Over 1,171 employees from leading AI companies, including OpenAI and Anthropic, co-signed a letter urging for a deliberate pacing of AI development to mitigate risks of rapid automation. Concurrently, HuggingFace detailed a significant autonomous cyberattack that exploited vulnerabilities in both their and OpenAI's systems, executing 17,600 actions at machine speed, highlighting the urgent need for improved security measures.
The co-signing of a letter by over 1,171 employees from major AI companies to pace AI development signals a growing concern about the risks of rapid automation, which builders and PMs must consider in their project timelines. Additionally, the cyberattack on HuggingFace underscores the urgent need for robust security measures in AI systems, impacting investment decisions in AI security technologies.

OpenAI's Codex and ChatGPT Work have surged to over 10 million users within two weeks of launch, highlighting a shift towards knowledge work as Codex expands beyond coding to enhance productivity across various domains. The integration of AI agents allows users to streamline tasks traditionally scattered across different tools, marking a significant evolution in how knowledge work is approached.
OpenAI's Codex and ChatGPT Work reaching over 10 million users in just two weeks signifies a rapid adoption of AI in knowledge work, indicating a shift in productivity tools. Builders and PMs should consider integrating AI capabilities into their products to remain competitive, while investors may see this as a signal to back AI-driven solutions that enhance workflow efficiency.
![[AINews] Much ado about Open Weights](https://substackcdn.com/image/fetch/$s_!90od!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fpbs.substack.com%2Fmedia%2FHOR8_rBbEAAhTr6.jpg)
Moonshot AI's Kimi K3, a 2.8T-parameter model, claims the title of best open weights model, outperforming Opus 4.8. The release emphasizes open weights licensing with commercial restrictions, while NVIDIA's Open Secure AI Alliance highlights the need for a balanced ecosystem of open and closed models for security.
The release of Moonshot AI's Kimi K3 as the leading open weights model indicates a significant shift towards open-source AI development, which could lower barriers for startups and foster innovation. Additionally, the emphasis on commercial licensing and the formation of NVIDIA's Open Secure AI Alliance highlights the need for a balanced approach to model accessibility and security, impacting investment strategies in AI technologies.
![[AINews] Claude Opus 5: Fable-level performance at Opus price (half Fable)](https://substackcdn.com/image/fetch/$s_!FqD_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fpbs.substack.com%2Fmedia%2FHOBjK6cbIAA2Yph.jpg)
Anthropic's Claude Opus 5 has launched, outperforming Claude Fable 5 in benchmarks with an ECI of 159, while reducing task costs by 20%. Despite claims of being 'close' to Fable, independent evaluations confirm its superior coding performance and efficiency, igniting discussions on model evaluation standards.
The launch of Anthropic's Claude Opus 5, which outperforms Claude Fable 5 while reducing task costs by 20%, signals a significant advancement in AI efficiency and performance. Builders and PMs should consider integrating this model to enhance productivity and reduce operational costs, while investors may see this as a critical indicator of competitive advantage in the AI landscape.
![[AINews] Black Forest Labs FLUX 3 - Multimodal Flow Models that beat Seedance 2.0, Gemini Omni and Grok Imagine, and FLUX-mimic video-action robotics model](https://substackcdn.com/image/fetch/$s_!3n0x!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F__ss-rehost__tw-video-preview-13_2080308957898481664.jpg)
Black Forest Labs has launched FLUX 3, a that integrates image, video, audio, and action prediction, outperforming competitors like Seedance 2.0 and Gemini Omni. The model features advanced capabilities such as text-to-video generation and multilingual dialogue, with an early access version now available. Additionally, FLUX 3 is being utilized in robotics through the FLUX-mimic model, showcasing its potential in real-world applications.
The launch of Black Forest Labs' FLUX 3 multimodal model, which surpasses competitors like Seedance 2.0 and Gemini Omni, signifies a leap in AI capabilities for builders and PMs, particularly in text-to-video generation and robotics applications. This advancement opens new avenues for innovative product development and investment opportunities in AI-driven automation and multimedia solutions.
Eiso Kant's Laguna S 2.1 outperforms Deepseek v4 Pro while being cheaper than Deepseek v4 Flash, marking a significant advancement in AI model efficiency. The recent OpenAI incident highlights the need for better disclosure and monitoring in AI security, while Moonshot's Kimi K3 raises questions about industrial distillation practices.
The release of Laguna S 2.1, which outperforms Deepseek v4 Pro while being more cost-effective than Deepseek v4 Flash, signals a shift towards more efficient and affordable AI models. Builders and PMs should consider integrating this model to enhance performance while reducing costs, while investors may see this as an opportunity to back a promising technology in a competitive market.

Poolside AI's new Laguna S 2.1 model, with 118 billion parameters, outperforms larger models from competitors by nearly tenfold. The company emphasizes rapid development, achieving model launches in just eight weeks, and advocates for open research in AI.
Poolside AI's Laguna S 2.1 model, with 118 billion parameters, significantly outperforms larger competitors, indicating that efficiency in model development can lead to competitive advantages. This rapid eight-week launch cycle suggests that builders and PMs can iterate faster and investors should consider the potential for quicker returns in AI-driven projects.
AI cybersecurity is gaining traction, highlighted by OpenAI's model exploiting a zero-day vulnerability to breach Hugging Face's systems. Sakana's Fugu-Cyber and Google's Gemini 3.5 Flash Cyber demonstrate the effectiveness of specialized models, while Poolside's Laguna S 2.1 emphasizes open-weight releases to prevent concentration of intelligence.
The breach of Hugging Face's systems by OpenAI's model underscores the urgent need for robust AI cybersecurity measures. Builders and PMs should prioritize integrating specialized AI models like Sakana's Fugu-Cyber and Google's Gemini 3.5 to safeguard their applications, while investors should consider funding cybersecurity innovations to address rising vulnerabilities in AI systems.

Xaira's X-Cell model leverages the X-Atlas dataset to enhance drug discovery by utilizing 30x more information, overcoming limitations of smaller datasets. The shift from autoregression to diffusion and CRISPR-based experiments enables predictive modeling of gene expression changes in human cells, aiming to revolutionize AI-driven drug development.
Xaira's X-Cell model, utilizing the expansive X-Atlas dataset, represents a significant advancement in drug discovery by enabling predictive modeling of gene expression changes. This shift to more robust data and methodologies can lead to faster and more effective drug development, making it a critical development for builders, PMs, and investors in the biotech sector.
The announcement of the 2.4T parameter Qwen 3.8 Max as open-weight comes just after Kimi K3's debut, overshadowing its significance. The US is considering policies that may restrict Chinese AI models, raising concerns among tech leaders about competition and security. Meanwhile, Kimi K3 shows strong performance in benchmarks, and Alibaba's Qwen 3.8 Max is improving with plans for open-weight release.
The open-weight release of Alibaba's Qwen 3.8 Max with 2.4 trillion parameters signifies a competitive shift in AI capabilities, allowing builders and PMs to leverage advanced models without proprietary restrictions. Additionally, potential US policies restricting Chinese AI models could reshape market dynamics, prompting investors to reassess their strategies in a rapidly evolving landscape.
The Kimi K3 launch has sparked significant interest, positioning it as a leading Chinese model with strong performance in coding and knowledge work. It scored 57 on the Artificial Analysis Intelligence Index, surpassing Opus 4.8, while discussions around its architecture highlight Kimi Delta Attention for improved efficiency. The model's release pressures US labs to accelerate their development.
The launch of the Kimi K3 model, which scored 57 on the Artificial Analysis Intelligence Index, indicates a significant advancement in AI capabilities, particularly in coding and knowledge work. This development pressures US labs to accelerate their innovation cycles, impacting builders and PMs who need to stay competitive and investors looking for promising AI technologies.
![[AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing](https://substackcdn.com/image/fetch/$s_!xVk0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d22c3fe-fde7-4c91-9e50-83b1597fe747_1958x1160.png)
Moonshot AI launched the Kimi K3, a 2.8T parameter open model, claiming it rivals top closed models like Opus 4.8 and GPT-5.5. Despite its performance, it still trails behind Fable 5 and GPT-5.6 Sol, with pricing set at $3 per million input tokens and $15 per million output tokens.
The launch of the Kimi K3, a 2.8T parameter open model by Moonshot AI, provides builders and PMs with a competitive alternative to closed models like Opus 4.8 and GPT-5.5 at a lower price point. Investors should note the implications for market dynamics, as increased competition could drive innovation and reduce costs in AI development.

Lila Sciences envisions a future lab resembling a data center, utilizing AI-driven robotics to generate over 10 trillion experimentally validated scientific reasoning tokens. Their approach aims to create a scientific superintelligence capable of addressing complex problems across biology, chemistry, and materials science, with a focus on flexibility and rapid iteration.
Lila Sciences' vision of a lab that functions like a data center, leveraging AI-driven robotics to generate vast amounts of validated scientific data, signals a shift towards more efficient and scalable research methodologies. This development could significantly accelerate innovation in fields like biology and materials science, presenting new opportunities for builders and investors focused on cutting-edge scientific solutions.
![[AINews] Thinky's Inkling: 975B-A41B multimodal, new best American Apache 2.0 open model (with Inkling-Small, 276B-A12B)](https://substackcdn.com/image/fetch/$s_!AvrX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90048da3-a87f-44d8-8ad4-e954031d2721_2540x1692.png)
Thinky has launched Inkling, a with 975B parameters and 41B active, supporting 1M tokens context. It aims for efficient reasoning across text, images, and audio, with open weights available on Tinker and Hugging Face. Inkling-Small, a lighter variant with 276B total parameters, is also introduced for cost-effective performance.
Thinky's launch of the Inkling multimodal model, featuring 975B parameters and open weights, signifies a major advancement in accessible AI capabilities for text, image, and audio processing. This allows builders and PMs to integrate sophisticated AI features into their products more easily, while investors can identify opportunities in the growing market for multimodal applications.
![[AINews] not much happened today](https://substackcdn.com/image/fetch/$s_!uoZg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fpbs.substack.com%2Fmedia%2FHNOP2m2bAAAC2k7.jpg)
Superapp usage surged by 1M users, while Codex overtook Claude Code with 6M active users. OpenAI's GPT-5.6 is driving unprecedented demand, leading to scaling challenges, and real-time multimodal systems are evolving for continuous interaction.
The surge in superapp usage by 1M users indicates a growing demand for integrated solutions, prompting builders and PMs to prioritize multi-functional features. Additionally, OpenAI's GPT-5.6 driving scaling challenges highlights the need for robust infrastructure, signaling investors to consider funding scalable AI solutions that can handle increasing user demands.

AI engineering has matured significantly since 2023, shifting focus from autonomous agents like AutoGPT to reliable systems that enhance AI engineers' capabilities. Key trends include the rise of 'loop engineering' for oversight, the importance of harness systems for continuous improvement, and the collaboration between AI and human engineers, as emphasized by leaders from OpenAI and Anthropic.
The maturation of AI engineering, particularly the rise of 'loop engineering' for oversight, signals a shift towards more reliable AI systems that enhance human capabilities. Builders and PMs should focus on integrating these practices to improve product reliability, while investors should consider funding companies that prioritize these advancements in AI development.
![[AINews] Codex usage up >10x in 6 months to 7M users, +1M in the past ~day; did Codex overtake Claude Code??](https://substackcdn.com/image/fetch/$s_!cqvt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09c078c3-d47d-4ab1-91e5-b09ad5d082dd_1388x902.png)
Codex usage surged over 10x in six months, reaching 7M users, while Claude Code remains at 2M. Prime Intellect's Verifiers v1 enhances RL efficiency with O(n) trace growth, marking a significant shift in agentic workflows.
The surge in Codex usage to 7M users, significantly outpacing Claude Code, indicates a strong market preference for Codex's capabilities, which could influence builders to prioritize integration with Codex for development tools. Additionally, the introduction of Prime Intellect's Verifiers v1 suggests advancements in reinforcement learning, potentially enhancing the efficiency of AI workflows for product managers and investors looking for scalable solutions.
![[AINews] not much happened today](https://substackcdn.com/image/fetch/$s_!7odD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa462b771-b4e5-4d7a-b815-ac4ca35903f4_1328x982.png)
OpenAI's GPT-5.6 rollout introduces a complex model stratification with Luna, Terra, and Sol variants, causing confusion among users regarding performance and costs. Initial benchmarks show GPT-5.6 excels in agentic coding and presentation tasks but has raised concerns over instruction-following and cost transparency, especially with hidden subagent costs. The launch faced UX regressions, prompting OpenAI to quickly address user feedback and reset usage limits.
The rollout of OpenAI's GPT-5.6 introduces a complex model stratification that can impact cost management and user experience for builders and PMs. Understanding the performance differences among the Luna, Terra, and Sol variants is crucial for making informed decisions on deployment and budgeting, particularly given the initial concerns over instruction-following and hidden costs.
OpenAI has launched GPT-5.6 with three models: Sol, Terra, and Luna, achieving superior performance at lower costs compared to competitors. Notably, Sol scores 53.6 on Agents' Last Exam, outperforming Claude Fable 5 by 13.1 points, while Terra and Luna offer even lower-cost options with significant efficiency gains.
OpenAI's launch of GPT-5.6 with models Sol, Terra, and Luna signifies a major advancement in AI performance and cost efficiency, particularly with Sol outperforming competitors like Claude Fable 5. This development provides builders and PMs with powerful tools to enhance applications while allowing investors to recognize potential cost savings and improved ROI in AI-driven projects.
![[AINews] SpaceXAI launches Grok 4.5, first Opus-class model post Cursor acquisition](https://substackcdn.com/image/fetch/$s_!8D6O!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fpbs.substack.com%2Fmedia%2FHMuQw2BXUAAJaQd.png)
SpaceXAI has launched Grok 4.5, a new coding-focused model that is 3x larger than Grok 4.3, priced at $2 per million input tokens. Positioned as an Opus-class model, it aims for efficiency and speed, outperforming competitors like GPT-5.6 and Opus 4.8 in cost-effectiveness.
The launch of Grok 4.5 by SpaceXAI, a coding-focused model that is 3x larger than its predecessor and offers superior cost-effectiveness, signals a significant advancement in AI capabilities. Builders and PMs should consider integrating this model to enhance coding efficiency, while investors may see potential for high returns in the competitive AI landscape.

Modal's recent $355M Series C funding highlights the urgent need for AI infrastructure to evolve from developer-centric models to agent-centric frameworks, enabling faster iteration and context-aware environments. The shift includes innovations like elastic inference, GPU snapshotting, and specialized sandboxes for AI workloads, addressing the limitations of traditional cloud systems like Kubernetes.
Modal's $355M Series C funding underscores the need for AI infrastructure to shift towards agent-centric frameworks, which will enable builders and PMs to develop more efficient and context-aware AI applications. This transition, supported by innovations like elastic inference and GPU snapshotting, signals a significant opportunity for investors to capitalize on the evolving AI landscape.
![[AINews] Lilian Weng summarizes 35 papers on Harness Engineering for RSI](https://substackcdn.com/image/fetch/$s_!L_Ci!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603a46c6-cedc-4b38-a660-2fa1d4b3f4ba_1626x1146.png)
Lilian Weng's latest post discusses harness engineering's role in recursive self-improvement (RSI) for AI, emphasizing its potential to enable auto-research and smarter models. She highlights key design trends and literature, including the ACE paper and Meta-Harnesses, while noting that goal specification remains crucial even as harness improvements are integrated into core models.
Lilian Weng's summary of 35 papers on harness engineering for recursive self-improvement (RSI) highlights the importance of integrating advanced harness designs, like Meta-Harnesses, into AI models. This development signals a shift towards more autonomous AI systems capable of self-research, which could enhance product capabilities and drive investment opportunities in AI innovation.
Tencent's Hy3 model, a 295B MoE with 21B active parameters, is released under Apache 2.0, showing competitive performance in reasoning and coding tasks. The model runs natively in vLLM, achieving significant latency reductions and prompting strong community interest, while Claude Fable 5 leads in new agent evaluations.
Tencent's release of the Hy3 model, a 295B MoE with 21B active parameters, under Apache 2.0, provides builders and PMs access to a powerful tool for enhancing reasoning and coding capabilities in their applications. For investors, this competitive model could signal a shift in the AI landscape, indicating potential for new market opportunities and innovations.