Principal ML Engineer - Large Scale Training Performance Optimization
Questa offerta fa per te?
Crea il mio CV Crea il tuo CV e scopri la tua percentuale di corrispondenza con questa posizione — e con tutte le altre.
La posizione
Lead distributed training of large-scale models across multi-GPU systems to convergence.
Design and optimize end-to-end training pipelines, leveraging data, tensor, pipeline, and expert parallelism (including ZeRO) to scale out.
Implement algorithmic and kernel-level optimizations to improve training throughput and efficiency.
Stay current with the latest training methods and contribute changes to open-source projects.
Collaborate across teams and stakeholders to influence the direction of the AI platform while communicating results clearly.
Requires a master's or PhD in CS/AI with hybrid (partial on-site) work in San Jose, CA or Bellevue, WA.
Design and optimize end-to-end training pipelines, leveraging data, tensor, pipeline, and expert parallelism (including ZeRO) to scale out.
Implement algorithmic and kernel-level optimizations to improve training throughput and efficiency.
Stay current with the latest training methods and contribute changes to open-source projects.
Collaborate across teams and stakeholders to influence the direction of the AI platform while communicating results clearly.
Requires a master's or PhD in CS/AI with hybrid (partial on-site) work in San Jose, CA or Bellevue, WA.
Vedi l'annuncio completo
Mansioni, profilo, competenze e vantaggi — crea il tuo account gratuito.
Hai già un account? Accedi
Offerte simili
Altre posizioni che potrebbero interessarti.
? Intern, Machine Learning Engineer ? Senior Data Scientist - Product Analytics (Payments/Ecommerce) ? Data Scientist Graduate (TikTok-Product-Data Science)-2026 Start (PhD) ? [HackerRank] Software Engineer Graduate (On-Device AI) - 2025 Start (BS/MS) ? Senior Director, Risk Data Science ? Data Scientist Graduate (TikTok-Product-Data Science)-2026 Start (BS/MS)
Lavoro da remotoPartial
CittàSan Jose, Stati Uniti