Staff Machine Learning Engineer, ML Performance & Optimization
¿Es esta oferta para usted?
Crear mi CV Cree su CV y descubra su porcentaje de coincidencia con este puesto — y con todos los demás.
El puesto
Take ownership as a Staff ML Engineer focused on ML performance and optimization across multi-platform deployments.
You will optimize neural architectures and systems for high performance on GPU/TPU hardware, including onboard and simulation platforms.
Develop post-training techniques like quantization and kernel-level optimizations to reduce latency and memory footprint for real-time constraints.
Experiment with new architectures (sparse models) and decoding strategies (speculative decoding) to boost inference speed.
Enhance training efficiency for large models and fine-tuning in data-heavy pipelines, while collaborating with ML infra, hardware, and research teams.
This hybrid role emphasizes hands-on optimization and cross-team collaboration to scale production-grade models.
You will optimize neural architectures and systems for high performance on GPU/TPU hardware, including onboard and simulation platforms.
Develop post-training techniques like quantization and kernel-level optimizations to reduce latency and memory footprint for real-time constraints.
Experiment with new architectures (sparse models) and decoding strategies (speculative decoding) to boost inference speed.
Enhance training efficiency for large models and fine-tuning in data-heavy pipelines, while collaborating with ML infra, hardware, and research teams.
This hybrid role emphasizes hands-on optimization and cross-team collaboration to scale production-grade models.
Ver la oferta completa
Funciones, perfil, competencias y ventajas — crea tu cuenta gratis.
¿Ya tiene una cuenta? Iniciar sesión
Ofertas similares
Otros puestos que podrían encajar.
TeletrabajoPartial
CiudadSan Francisco, Estados Unidos