Machine Learning Platform Engineer
Is this job for you?
Build my CV Build your CV and see how well you match this role — and every other one.
The role
Build and operate a scalable container-based inference platform to enable custom ML models, optimizing autoscaling and reducing cold starts for end-to-end performance.
Work on multi-cluster orchestration, predictive autoscaling, deployment APIs, inference worker SDKs, and CLI tools.
Analyze and improve robustness and scalability of distributed systems, APIs, databases, and infrastructure.
Partner with product teams to translate requirements into scalable solutions and write clear, well-tested software and IaC.
Lead design and code reviews, create developer documentation, and develop testing strategies for robustness and fault tolerance.
Ideal candidates have 5+ years in large-scale distributed systems, strong OS concepts, proficiency in Python/Golang/Rust/C++, Kubernetes, Terraform, and ML bottlenecks awareness.
Work on multi-cluster orchestration, predictive autoscaling, deployment APIs, inference worker SDKs, and CLI tools.
Analyze and improve robustness and scalability of distributed systems, APIs, databases, and infrastructure.
Partner with product teams to translate requirements into scalable solutions and write clear, well-tested software and IaC.
Lead design and code reviews, create developer documentation, and develop testing strategies for robustness and fault tolerance.
Ideal candidates have 5+ years in large-scale distributed systems, strong OS concepts, proficiency in Python/Golang/Rust/C++, Kubernetes, Terraform, and ML bottlenecks awareness.
See the full job post
Responsibilities, requirements, skills and benefits — create your free account.
or
Already have an account?
Log inSimilar openings
Other roles that could suit you.
Remote workno
CitySan Francisco, United States