Naar de inhoud

Principal ML Engineer - Large Scale Training Performance Optimization

Is deze vacature iets voor u?

Maak uw cv en ontdek uw matchpercentage met deze functie — en met alle andere.

Mijn cv maken

De functie

Lead distributed training of large-scale models across multi-GPU systems to convergence.
Design and optimize end-to-end training pipelines, leveraging data, tensor, pipeline, and expert parallelism (including ZeRO) to scale out.
Implement algorithmic and kernel-level optimizations to improve training throughput and efficiency.
Stay current with the latest training methods and contribute changes to open-source projects.
Collaborate across teams and stakeholders to influence the direction of the AI platform while communicating results clearly.
Requires a master's or PhD in CS/AI with hybrid (partial on-site) work in San Jose, CA or Bellevue, WA.

Bekijk de volledige vacature

Taken, profiel, vaardigheden en voordelen — maak gratis een account aan.

Minimaal 8 tekens, waaronder een hoofdletter en een cijfer.

Al een account? Inloggen

Vergelijkbare vacatures

Andere functies die kunnen passen.

Alles bekijken →

Uw locatie

Vacatures en bedrijven worden op dit land gefilterd.

Voorgesteld

Alle landen 68