Large Machine Learning Model Optimization Engineer
Is this job for you?
Build my CV Build your CV and see how well you match this role — and every other one.
The role
Lead on-device optimization of large language models and diffusion models to deliver real-time, low-latency experiences.
Drive model compression strategies including quantization, pruning, and distillation for on-device deployment.
Implement and optimize hardware-aware pipelines and ML compilers for efficient inference.
Collaborate with hardware, software, and ML teams to align hardware-software co-design and improve on-device experiences.
Contribute to research publications and share findings with cross-functional teams.
Requirements include Python software engineering, experience with large-scale ML models, and strong collaboration and communication skills.
Drive model compression strategies including quantization, pruning, and distillation for on-device deployment.
Implement and optimize hardware-aware pipelines and ML compilers for efficient inference.
Collaborate with hardware, software, and ML teams to align hardware-software co-design and improve on-device experiences.
Contribute to research publications and share findings with cross-functional teams.
Requirements include Python software engineering, experience with large-scale ML models, and strong collaboration and communication skills.
See the full job post
Responsibilities, requirements, skills and benefits — create your free account.
or
Already have an account?
Log inSimilar openings
Other roles that could suit you.