跳到正文

Principal ML Engineer - Large Scale Training Performance Optimization

这个职位适合你吗?

创建简历,即可看到你与这个职位——以及其他所有职位——的匹配度。

创建我的简历

职位介绍

Lead distributed training of large-scale models across multi-GPU systems to convergence.
Design and optimize end-to-end training pipelines, leveraging data, tensor, pipeline, and expert parallelism (including ZeRO) to scale out.
Implement algorithmic and kernel-level optimizations to improve training throughput and efficiency.
Stay current with the latest training methods and contribute changes to open-source projects.
Collaborate across teams and stakeholders to influence the direction of the AI platform while communicating results clearly.
Requires a master's or PhD in CS/AI with hybrid (partial on-site) work in San Jose, CA or Bellevue, WA.

查看完整职位

工作职责、任职要求、技能与福利 — 免费创建账号即可查看。

至少 8 个字符,需包含一个大写字母和一个数字。

已有账户? 登录

相似职位

其他可能适合您的职位。

查看全部 →

您的城市

职位和企业将按该国家筛选。

推荐

所有国家 66