
Expedia Uses AI Driven Service Telemetry Analyzer to Accelerate Incident Investigation
Quick Answer
Expedia has launched the Service Telemetry Analyzer (STAR), an AI-assisted platform that streamlines incident investigations by analyzing service telemetry and generating structured root cause assessments.
Quick Take
By integrating operational metrics with , STAR aims to minimize time to know (TTK) and time to recover (TTR) while maintaining human oversight in decision-making.
Key Points
- STAR integrates with Datadog to analyze operational metrics from Kubernetes and JVM applications.
- The platform uses a deterministic workflow with prompt chaining for structured analysis.
- Expedia transitioned to a Celery-based architecture for asynchronous processing of telemetry data.
- STAR has been utilized for production incident investigations and Kubernetes troubleshooting.
- Future enhancements include service dependency information and conversational interfaces.
DeepSignal Analysis
What happened
Expedia Group has launched the Service Telemetry Analyzer (STAR), an AI-assisted platform designed to aid engineers in investigating production incidents. By analyzing service telemetry and generating structured root cause assessments, STAR aims to reduce the time engineers spend on identifying service degradation sources while ensuring human oversight in decision-making.
Key evidence
- STAR integrates operational metrics with large language models (LLMs) and follows a deterministic workflow, which includes collecting telemetry, analyzing it with domain-specific prompts, and summarizing findings into a final report.
- The platform is built as a FastAPI application that connects with Datadog for service metrics retrieval and employs a Celery-based asynchronous architecture to handle multiple analysis tasks concurrently.
- Expedia reports that STAR has been utilized for various purposes, including production incident investigations and Kubernetes troubleshooting, with engineers reviewing generated findings before taking action.
Why it matters
The introduction of STAR reflects a growing trend in the tech industry to leverage AI for operational efficiency. By streamlining incident investigations, Expedia aims to enhance its response times and overall service reliability. However, the emphasis on human oversight suggests a cautious approach to AI integration, prioritizing accuracy and accountability in critical operational processes.
📖 Reader Mode
~3 min readExpedia Group has introduced Service Telemetry Analyzer (STAR), an internal AI-assisted observability platform that helps engineers investigate production incidents by analyzing service telemetry and generating structured root cause assessments. The system combines operational metrics with large language models (LLMs) through predefined diagnostic workflows, aiming to reduce the time engineers spend identifying the source of service degradation while keeping humans responsible for validation and decision-making.
Rather than adopting autonomous AI agents, the platform follows a deterministic workflow in which telemetry is collected, analyzed using domain-specific prompts, consolidated into intermediate findings, and summarized into a final report containing potential root causes and recommended next steps.
Describing the project's goal, the team wrote,
Our objective with this service was to minimize the time to know (TTK) and time to recover (TTR).

STAR Architecture (Source: Expedia Blog Post)
The platform is implemented as a FastAPI application that integrates with Datadog to retrieve service metrics and an internal generative AI gateway that manages authentication and access to LLM providers. The workflow uses prompt chaining, allowing multiple specialized analyses to be performed before producing a consolidated diagnosis. Expedia said the current implementation does not use capabilities such as function calling, retrieval-augmented generation (RAG), memory, or autonomous tool use, instead relying on predefined workflows to generate consistent analyses.
STAR focuses on standardized infrastructure telemetry collected from Kubernetes-based services and JVM applications. The system analyzes metrics including request throughput, latency, HTTP, gRPC, and GraphQL error rates, CPU and memory utilization, container restart events, Kubernetes readiness and liveness probe failures, Java heap utilization, and garbage collection activity. Expedia explains that infrastructure metrics provide a consistent view across services developed using different programming languages and frameworks.
As the platform evolved, Expedia replaced FastAPI background tasks with a Celery-based asynchronous architecture using Redis as both the message broker and result backend. The engineering team states that most processing consists of I/O bound operations involving telemetry retrieval and LLM requests. The asynchronous execution model allows STAR to process multiple analysis tasks concurrently while accommodating rate limits imposed by Datadog and the company's internal AI gateway.
The company reports that STAR has been used to support production incident investigations, post-incident analysis, Kubernetes troubleshooting, and JVM memory diagnostics. Engineers review the generated findings before acting on recommendations, making the platform an assistive tool rather than an autonomous operational system.
The blog also describes ongoing work to expand the platform. Planned enhancements include incorporating service dependency information, additional operational metadata, Model Context Protocol (MCP) based tool integrations, and conversational interfaces. Expedia is also evaluating STAR as part of its chaos engineering practices to help analyze the results of controlled failure experiments. Prompt management, tracing, and evaluation are currently supported through Langfuse, while system performance is assessed using subject matter expert reviews and user feedback.
About the Author
Leela Kumili
Show moreShow less
— Originally published at infoq.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from InfoQ AI, ML & Data Engineering
See more →Google Cloud Workbench Notebooks Extension Connects VS Code to Google Cloud's Jupyter Notebooks
The Google Cloud Workbench Notebooks extension for VS Code allows developers to seamlessly connect their local IDE to managed Jupyter notebook environments on Google Cloud, enhancing ML workflow efficiency. This integration eliminates context switching, enabling smooth transitions from local experimentation to high-performance cloud computing.

