Machine Learning Platform Engineer
仕事内容
Build and operate a scalable container-based inference platform to enable custom ML models, optimizing autoscaling and reducing cold starts for end-to-end performance.
Work on multi-cluster orchestration, predictive autoscaling, deployment APIs, inference worker SDKs, and CLI tools.
Analyze and improve robustness and scalability of distributed systems, APIs, databases, and infrastructure.
Partner with product teams to translate requirements into scalable solutions and write clear, well-tested software and IaC.
Lead design and code reviews, create developer documentation, and develop testing strategies for robustness and fault tolerance.
Ideal candidates have 5+ years in large-scale distributed systems, strong OS concepts, proficiency in Python/Golang/Rust/C++, Kubernetes, Terraform, and ML bottlenecks awareness.
Work on multi-cluster orchestration, predictive autoscaling, deployment APIs, inference worker SDKs, and CLI tools.
Analyze and improve robustness and scalability of distributed systems, APIs, databases, and infrastructure.
Partner with product teams to translate requirements into scalable solutions and write clear, well-tested software and IaC.
Lead design and code reviews, create developer documentation, and develop testing strategies for robustness and fault tolerance.
Ideal candidates have 5+ years in large-scale distributed systems, strong OS concepts, proficiency in Python/Golang/Rust/C++, Kubernetes, Terraform, and ML bottlenecks awareness.
求人の全文を見る
業務内容、求める人物像、スキル、待遇 — 無料アカウントの作成で閲覧できます。
または
すでにアカウントをお持ちですか?
ログイン類似の求人
あなたに合いそうな他の職種。
リモートワークno
勤務地San Francisco, アメリカ合衆国