Zum Inhalt springen

Principal ML Engineer - Large Scale Training Performance Optimization

Ist diese Stelle etwas für Sie?

Erstellen Sie Ihren Lebenslauf und entdecken Sie Ihre Übereinstimmung mit dieser Stelle — und mit allen anderen.

Lebenslauf erstellen

Die Stelle

Lead distributed training of large-scale models across multi-GPU systems to convergence.
Design and optimize end-to-end training pipelines, leveraging data, tensor, pipeline, and expert parallelism (including ZeRO) to scale out.
Implement algorithmic and kernel-level optimizations to improve training throughput and efficiency.
Stay current with the latest training methods and contribute changes to open-source projects.
Collaborate across teams and stakeholders to influence the direction of the AI platform while communicating results clearly.
Requires a master's or PhD in CS/AI with hybrid (partial on-site) work in San Jose, CA or Bellevue, WA.

Die vollständige Anzeige sehen

Aufgaben, Anforderungen, Kompetenzen und Vorteile — mit Ihrem kostenlosen Konto.

Mindestens 8 Zeichen, davon ein Großbuchstabe und eine Ziffer.

Bereits ein Konto? Anmelden

Ähnliche Stellen

Weitere Positionen, die passen könnten.

Alle ansehen →

Ihr Ort

Stellen und Unternehmen werden nach diesem Land gefiltert.

Vorschläge

Alle Länder 66