Vai al contenuto

Principal Machine Learning Engineer, Distributed vLLM Inference

Questa offerta fa per te?

Crea il tuo CV e scopri la tua percentuale di corrispondenza con questa posizione — e con tutte le altre.

Crea il mio CV

La posizione

Develop and maintain distributed inference infrastructure leveraging Kubernetes APIs, operators, and the Gateway Inference Extension API for scalable LLM deployments.
Build system components in Go and/or Rust to integrate with the vLLM project and manage distributed workloads.
Design and implement KV cache-aware routing and scoring to optimize memory usage and request distribution at scale.
Enhance resource utilization, fault tolerance, and stability of the inference stack.
Contribute to the design, development, and testing of various inference optimization algorithms; actively participate in technical design discussions.
Provide mentorship and conduct code reviews to engineers, fostering a culture of continuous learning and innovation.

Vedi l'annuncio completo

Mansioni, profilo, competenze e vantaggi — crea il tuo account gratuito.

Almeno 6 caratteri. Più è lunga, più è sicura.
o

Hai già un account?

Offerte simili

Altre posizioni che potrebbero interessarti.

Vedi tutto →

La tua località

Offerte e aziende saranno filtrate su questo paese.

Suggeriti

Tutti i paesi 68