GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model
Meta’s Generative Ads Recommendation Model (GEM), the foundation model behind ads recommendations across Instagram and Facebook, now trains at LLM scale on several thousand of the latest-generation GPUs. This post goes into the details on how we achieved: doubling end-to-end (E2E) training efficiency to 20–25% Model FLOPs Utilization (MFU) while scaling training FLOPs 4x in [...] Read More... The post GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model appeared first on Engineering at Meta.
Score breakdown
- Technical depth59
- Practical value90
- Originality100
- Writing quality100
- Source reputation88
- Recency28
- External engagement0
- On-site engagement0
More like this
Pinterest Engineering59
Achieving Near-Linear Training Scalability for Pinterest’s Foundation Models
Pinterest Engineering·unknown·11 min read
Etsy — Code as Craft64
Making Ads Count: Using MMoE and Auxiliary Tasks to Better Connect Buyers & Sellers
Amanda Steigman·unknown·11 min read
Yelp Engineering32
Training Orchestrator: Unifying Model Training at Yelp
Ying Wang and Nathan Sponberg, Software Engineer
·unknown·1 min read
LinkedIn Engineering55
Faster than Light: Optimizing Generative Recommender Training Efficiency at LinkedIn
unknown·14 min read
Engradar shows a summary and links to the original article. The full article is hosted on engineering.fb.com.