Machine Learning Platform Engineer
职位介绍
Build and operate a scalable container-based inference platform to enable custom ML models, optimizing autoscaling and reducing cold starts for end-to-end performance.
Work on multi-cluster orchestration, predictive autoscaling, deployment APIs, inference worker SDKs, and CLI tools.
Analyze and improve robustness and scalability of distributed systems, APIs, databases, and infrastructure.
Partner with product teams to translate requirements into scalable solutions and write clear, well-tested software and IaC.
Lead design and code reviews, create developer documentation, and develop testing strategies for robustness and fault tolerance.
Ideal candidates have 5+ years in large-scale distributed systems, strong OS concepts, proficiency in Python/Golang/Rust/C++, Kubernetes, Terraform, and ML bottlenecks awareness.
Work on multi-cluster orchestration, predictive autoscaling, deployment APIs, inference worker SDKs, and CLI tools.
Analyze and improve robustness and scalability of distributed systems, APIs, databases, and infrastructure.
Partner with product teams to translate requirements into scalable solutions and write clear, well-tested software and IaC.
Lead design and code reviews, create developer documentation, and develop testing strategies for robustness and fault tolerance.
Ideal candidates have 5+ years in large-scale distributed systems, strong OS concepts, proficiency in Python/Golang/Rust/C++, Kubernetes, Terraform, and ML bottlenecks awareness.
查看完整职位
工作职责、任职要求、技能与福利 — 免费创建账号即可查看。
或
已有账户?
登录相似职位
其他可能适合您的职位。
远程办公no
城市San Francisco, 美国