
Yelp Unifies ML Model Training with Training Orchestrator
Quick Answer
Yelp's new Training Orchestrator replaces fragmented ML training scripts with a unified, configuration-driven framework, enhancing reproducibility and efficiency across teams.
Quick Take
The DAG-based model allows for local runs, improved testing, and streamlined orchestration, addressing common issues in large ML platforms.
Key Points
- Training Orchestrator uses Pydantic-based configurations for improved validation and efficiency.
- The framework allows for local runs and unit testing, enhancing development speed.
- Automatic logging of configurations as MLflow artifacts enables easy reruns.
- Yelp plans to implement model evaluation and lineage tracking in future updates.
- This centralization trend is seen across the industry, improving reproducibility and testability.
DeepSignal Analysis
What happened
Yelp has introduced the Training Orchestrator, an internal framework designed to unify machine learning model training. This system replaces fragmented Spark training scripts with a configuration-driven, Directed Acyclic Graph (DAG) model, which enhances reproducibility and efficiency across teams by addressing issues like duplicated code and inconsistent configurations.
Key evidence
- The Training Orchestrator replaces individual team Spark training scripts, addressing problems like duplicated code and inconsistent configurations.
- Yelp's previous training setup was monolithic and tightly coupled to Spark, causing slow iteration and poor testability.
- The new framework allows for local runs and unit testing, which improves the development process and reduces integration test times.
Why it matters
This development reflects a broader industry trend toward centralized machine learning infrastructure, which aims to improve reproducibility and efficiency. By addressing common issues faced by multiple teams, Yelp's Training Orchestrator could serve as a model for other organizations facing similar challenges. The ability to run tests locally and validate configurations quickly may lead to faster iterations and more reliable deployments.
📖 Reader Mode
~4 min readYelp has launched Training Orchestrator. This new internal framework replaces individual team Spark training scripts. Now, it uses a configuration-driven, DAG-based execution model. The move tackles a common problem in large ML platforms. Each applied ML team created its own orchestration. This led to duplicated code, inconsistent configurations, and fragile custom monitoring.
Training at Yelp previously ran as monolithic scripts tightly coupled to Spark cluster and job-launcher execution. That coupling caused no local runs, poor testability, and slow iteration. A small code change needed submitting a Spark job. Then, you had to wait for containers and clusters to start up before finding out if there was a failure. Validation logic was spread out in scripts. This made it difficult to reproduce past runs in different environments. There were issues with inconsistent settings and missing provenance.
Training Orchestrator separates what to run from how it runs. Pipelines use Pydantic-based configuration objects. These include an orchestrator config, an MLflow run config, and various step configs. All are validated when created, so no compute is wasted. The orchestrator uses declared step dependencies to create a Directed Acyclic Graph. It then runs the steps in topological order. A shared Spark/MLflow context is added, allowing the same step definitions to run unchanged, whether locally, in Jupyter, or in production. Each step links a settings class to a user-defined function. It uses an input/output schema to show what it takes and gives back. For example, a PrepareDataStep takes in a dataset and outputs a prepared dataset. This setup lets teams easily change preprocessing logic. They can run different model versions simultaneously. Moreover, they can reuse functions across pipelines without changing orchestration code.

Schema validation quickly catches configuration mismatches. It flags issues like incorrect step input/output pairings and out-of-range parameters in seconds. This is much faster than finding them hours into a running Spark job. Developers can run full pipelines locally with sample data since steps are separate from infrastructure. They can also write unit tests for each step. Plus, caching input for data-loading steps reduces integration test time on smaller data sets. Every run's complete configuration is logged automatically as an MLflow artifact, making reruns replayable from that single record. Slack notifications and MLflow run tracking, previously hand-rolled per team, now require only a few configuration parameters.
The Training Orchestrator works with Yelp's ML platform. It connects feature stores, a shared neural network, a gradient-boosted-tree library, and MLflow for tracking and deployment. The Core Machine Learning Team sees it as the vital orchestration layer they needed. It is an internal framework, not an open-source release. Yelp plans to add clear steps for model evaluation, comparison, and explanation. It will also implement lineage tracking at the orchestration layer. This way, it will spread to current pipelines without needing migration.
The main pattern involves Pydantic-validated, declarative step configs that drive a DAG execution engine. This pattern separates training logic from cluster runtime. It also enables local runs and unit testing. This approach can be used beyond Yelp's Spark-specific stack. The tradeoff involves an initial investment in schema and step-type design. It also requires migrating existing scripts into the settings-and-function contract. Yelp views this as a way to centralise the cost, so teams don’t have to pay it repeatedly.
Yelp's effort reflects a wider trend in the industry. Many companies with several ML teams are moving towards a centralised, clear training infrastructure. Netflix has also moved forward with Metaflow. They recently added a configuration object. This lets teams manage flow behaviour clearly across the many ML pipelines they maintain. Uber's Michelangelo platform was created to tackle the fragmentation issue that Yelp mentions. It helps teams using different tools by providing a clear path from prototyping to production.
These efforts show a shared conclusion from various engineering groups: as more models and training teams emerge, ad hoc orchestration becomes a bottleneck. Centralising this process, despite upfront costs, improves reproducibility, testability, and developer speed.
About the Author
Claudio Masolo
Show moreShow less
— Originally published at infoq.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from InfoQ AI, ML & Data Engineering
See more →Google Cloud Workbench Notebooks Extension Connects VS Code to Google Cloud's Jupyter Notebooks
The Google Cloud Workbench Notebooks extension for VS Code allows developers to seamlessly connect their local IDE to managed Jupyter notebook environments on Google Cloud, enhancing ML workflow efficiency. This integration eliminates context switching, enabling smooth transitions from local experimentation to high-performance cloud computing.

