Machine Learning Infrastructure Engineer
仕事内容
The role focuses on designing, building, and maintaining scalable ML training and serving infrastructure to accelerate research and product development.
You will develop tooling to diagnose cluster issues and hardware failures, monitor deployments, and manage experiments.
A core priority is maximizing GPU allocation and utilization for both training and serving workloads.
Provide infrastructure support to ML teams, troubleshoot performance bottlenecks, and coordinate with platform engineers.
Required hands-on experience with cloud platforms, Kubernetes, and GPU-enabled environments, plus experience with PyTorch, TensorFlow, or JAX.
This role demands 4+ years of ML infra experience and a proactive, collaborative mindset in a fast-paced research setting.
You will develop tooling to diagnose cluster issues and hardware failures, monitor deployments, and manage experiments.
A core priority is maximizing GPU allocation and utilization for both training and serving workloads.
Provide infrastructure support to ML teams, troubleshoot performance bottlenecks, and coordinate with platform engineers.
Required hands-on experience with cloud platforms, Kubernetes, and GPU-enabled environments, plus experience with PyTorch, TensorFlow, or JAX.
This role demands 4+ years of ML infra experience and a proactive, collaborative mindset in a fast-paced research setting.
求人の全文を見る
業務内容、求める人物像、スキル、待遇 — 無料アカウントの作成で閲覧できます。
または
すでにアカウントをお持ちですか?
ログイン類似の求人
あなたに合いそうな他の職種。
? Machine Learning Intern Fall 2026 (Toronto) ? Machine Learning Engineer, Behavior Internship ? Machine Learning Engineer (L3) ? Senior C++ Programmer - Machine Learning Content Creation Technology Group ? Senior C++ Programmer - Machine Learning - Content Creation Technology Group ? Team Lead - Machine Learning Engineer
リモートワークno
勤務地Redwood City, United States