Machine Learning Operations Lead
Is this job for you?
Build my CV Build your CV and see how well you match this role — and every other one.
The role
Lead the design, delivery, and operation of production ML inference and fine-tuning services across serverless and multi-cluster deployments.
Own availability and performance SLAs, drive incident response, and conduct postmortems to prevent recurrence.
Build and scale testing, deployment, configuration management, and monitoring practices in collaboration with Infra SREs.
Define and enforce configuration best practices for inference engines (vLLM, tvLLM, Pulsar) to prevent runtime issues.
Develop self-serve tooling and internal developer platforms to reduce operational toil for ML engineers and customers.
Lead, mentor, and grow an MLOps team while partnering with infrastructure and ML engineering to improve reliability and cost efficiency.
Own availability and performance SLAs, drive incident response, and conduct postmortems to prevent recurrence.
Build and scale testing, deployment, configuration management, and monitoring practices in collaboration with Infra SREs.
Define and enforce configuration best practices for inference engines (vLLM, tvLLM, Pulsar) to prevent runtime issues.
Develop self-serve tooling and internal developer platforms to reduce operational toil for ML engineers and customers.
Lead, mentor, and grow an MLOps team while partnering with infrastructure and ML engineering to improve reliability and cost efficiency.
See the full job post
Responsibilities, requirements, skills and benefits — create your free account.
or
Already have an account?
Log inSimilar openings
Other roles that could suit you.
Remote workPartial
CitySan Francisco, United States