Machine Learning Infrastructure Engineer
职位介绍
The role focuses on designing, building, and maintaining scalable ML training and serving infrastructure to accelerate research and product development.
You will develop tooling to diagnose cluster issues and hardware failures, monitor deployments, and manage experiments.
A core priority is maximizing GPU allocation and utilization for both training and serving workloads.
Provide infrastructure support to ML teams, troubleshoot performance bottlenecks, and coordinate with platform engineers.
Required hands-on experience with cloud platforms, Kubernetes, and GPU-enabled environments, plus experience with PyTorch, TensorFlow, or JAX.
This role demands 4+ years of ML infra experience and a proactive, collaborative mindset in a fast-paced research setting.
You will develop tooling to diagnose cluster issues and hardware failures, monitor deployments, and manage experiments.
A core priority is maximizing GPU allocation and utilization for both training and serving workloads.
Provide infrastructure support to ML teams, troubleshoot performance bottlenecks, and coordinate with platform engineers.
Required hands-on experience with cloud platforms, Kubernetes, and GPU-enabled environments, plus experience with PyTorch, TensorFlow, or JAX.
This role demands 4+ years of ML infra experience and a proactive, collaborative mindset in a fast-paced research setting.
查看完整职位
工作职责、任职要求、技能与福利 — 免费创建账号即可查看。
或
已有账户?
登录相似职位
其他可能适合您的职位。
? Machine Learning Intern Fall 2026 (Toronto) ? Machine Learning Engineer, Behavior Internship ? Machine Learning Engineer (L3) ? Senior C++ Programmer - Machine Learning Content Creation Technology Group ? Senior C++ Programmer - Machine Learning - Content Creation Technology Group ? Team Lead - Machine Learning Engineer
远程办公no
城市Redwood City, United States