Skip to content

Principal ML Engineer - Large Scale Training Performance Optimization

Is this job for you?

Build your CV and see how well you match this role — and every other one.

Build my CV

The role

Lead distributed training of large-scale models across multi-GPU systems to convergence.
Design and optimize end-to-end training pipelines, leveraging data, tensor, pipeline, and expert parallelism (including ZeRO) to scale out.
Implement algorithmic and kernel-level optimizations to improve training throughput and efficiency.
Stay current with the latest training methods and contribute changes to open-source projects.
Collaborate across teams and stakeholders to influence the direction of the AI platform while communicating results clearly.
Requires a master's or PhD in CS/AI with hybrid (partial on-site) work in San Jose, CA or Bellevue, WA.

See the full job post

Responsibilities, requirements, skills and benefits — create your free account.

At least 8 characters, one uppercase letter and one digit.

Already have an account? Log in

Similar openings

Other roles that could suit you.

See all →

Your location

Jobs and companies will be filtered on this country.

Suggested

All countries 65