Zum Inhalt springen

Principal Machine Learning Engineer, Distributed vLLM Inference

Ist diese Stelle etwas für Sie?

Erstellen Sie Ihren Lebenslauf und entdecken Sie Ihre Übereinstimmung mit dieser Stelle — und mit allen anderen.

Lebenslauf erstellen

Die Stelle

Develop and maintain distributed inference infrastructure leveraging Kubernetes APIs, operators, and the Gateway Inference Extension API for scalable LLM deployments.
Build system components in Go and/or Rust to integrate with the vLLM project and manage distributed workloads.
Design and implement KV cache-aware routing and scoring to optimize memory usage and request distribution at scale.
Enhance resource utilization, fault tolerance, and stability of the inference stack.
Contribute to the design, development, and testing of various inference optimization algorithms; actively participate in technical design discussions.
Provide mentorship and conduct code reviews to engineers, fostering a culture of continuous learning and innovation.

Die vollständige Anzeige sehen

Aufgaben, Anforderungen, Kompetenzen und Vorteile — mit Ihrem kostenlosen Konto.

Mindestens 6 Zeichen. Je länger, desto sicherer.
oder

Bereits ein Konto?

Ähnliche Stellen

Weitere Positionen, die passen könnten.

Alle ansehen →

Ihr Ort

Stellen und Unternehmen werden nach diesem Land gefiltert.

Vorschläge

Alle Länder 66