Training Orchestrator: Unifying Model Training at Yelp
At Yelp, we train many machine learning models on different schedules. Applied machine learning teams all have their own set of Spark-based training batches, scripts, and configurations. Over time, these diverged, leading to duplicated code, subtle inconsistencies, and a growing maintenance burden. Yelp’s Core Machine Learning Team has developed excellent tooling across our ML ecosystem over the years: feature stores for reproducible data, a unified training library for neural networks and gradient-boosted trees, seamless Spark integration, and MLflow services for model tracking and deployment. But there was still one key piece missing right in the middle: a standardized way to...
Score breakdown
- Technical depth35
- Practical value30
- Originality35
- Writing quality30
- Source reputation80
- Recency7
- External engagement0
- On-site engagement0
More like this
The So-fine Real-time ML Paradigm
Kafka App? There’s a Skill for That
Introducing Roast: Structured AI workflows made easy (2025) - Shopify
RoCE networks for distributed AI training at scale
Engradar shows a summary and links to the original article. The full article is hosted on engineeringblog.yelp.com.