Staff Machine Learning Engineer, ML Performance & Optimization
Questa offerta fa per te?
Crea il mio CV Crea il tuo CV e scopri la tua percentuale di corrispondenza con questa posizione — e con tutte le altre.
La posizione
Take ownership as a Staff ML Engineer focused on ML performance and optimization across multi-platform deployments.
You will optimize neural architectures and systems for high performance on GPU/TPU hardware, including onboard and simulation platforms.
Develop post-training techniques like quantization and kernel-level optimizations to reduce latency and memory footprint for real-time constraints.
Experiment with new architectures (sparse models) and decoding strategies (speculative decoding) to boost inference speed.
Enhance training efficiency for large models and fine-tuning in data-heavy pipelines, while collaborating with ML infra, hardware, and research teams.
This hybrid role emphasizes hands-on optimization and cross-team collaboration to scale production-grade models.
You will optimize neural architectures and systems for high performance on GPU/TPU hardware, including onboard and simulation platforms.
Develop post-training techniques like quantization and kernel-level optimizations to reduce latency and memory footprint for real-time constraints.
Experiment with new architectures (sparse models) and decoding strategies (speculative decoding) to boost inference speed.
Enhance training efficiency for large models and fine-tuning in data-heavy pipelines, while collaborating with ML infra, hardware, and research teams.
This hybrid role emphasizes hands-on optimization and cross-team collaboration to scale production-grade models.
Vedi l'annuncio completo
Mansioni, profilo, competenze e vantaggi — crea il tuo account gratuito.
Hai già un account? Accedi
Offerte simili
Altre posizioni che potrebbero interessarti.
Lavoro da remotoPartial
CittàSan Francisco, Stati Uniti