Principal Machine Learning Engineer, Distributed vLLM Inference
Ist diese Stelle etwas für Sie?
Lebenslauf erstellen Erstellen Sie Ihren Lebenslauf und entdecken Sie Ihre Übereinstimmung mit dieser Stelle — und mit allen anderen.
Die Stelle
Develop and maintain distributed inference infrastructure leveraging Kubernetes APIs, operators, and the Gateway Inference Extension API for scalable LLM deployments.
Build system components in Go and/or Rust to integrate with the vLLM project and manage distributed workloads.
Design and implement KV cache-aware routing and scoring to optimize memory usage and request distribution at scale.
Enhance resource utilization, fault tolerance, and stability of the inference stack.
Contribute to the design, development, and testing of various inference optimization algorithms; actively participate in technical design discussions.
Provide mentorship and conduct code reviews to engineers, fostering a culture of continuous learning and innovation.
Build system components in Go and/or Rust to integrate with the vLLM project and manage distributed workloads.
Design and implement KV cache-aware routing and scoring to optimize memory usage and request distribution at scale.
Enhance resource utilization, fault tolerance, and stability of the inference stack.
Contribute to the design, development, and testing of various inference optimization algorithms; actively participate in technical design discussions.
Provide mentorship and conduct code reviews to engineers, fostering a culture of continuous learning and innovation.
Die vollständige Anzeige sehen
Aufgaben, Anforderungen, Kompetenzen und Vorteile — mit Ihrem kostenlosen Konto.
oder
Bereits ein Konto?
AnmeldenÄhnliche Stellen
Weitere Positionen, die passen könnten.
StadtBoston, Vereinigte Staaten