Aller au contenu

Principal Machine Learning Engineer, Distributed vLLM Inference

Cette offre est-elle pour vous ?

Créez votre CV et découvrez votre pourcentage de correspondance avec ce poste — et avec tous les autres.

Créer mon CV

Le poste

Develop and maintain distributed inference infrastructure leveraging Kubernetes APIs, operators, and the Gateway Inference Extension API for scalable LLM deployments.
Build system components in Go and/or Rust to integrate with the vLLM project and manage distributed workloads.
Design and implement KV cache-aware routing and scoring to optimize memory usage and request distribution at scale.
Enhance resource utilization, fault tolerance, and stability of the inference stack.
Contribute to the design, development, and testing of various inference optimization algorithms; actively participate in technical design discussions.
Provide mentorship and conduct code reviews to engineers, fostering a culture of continuous learning and innovation.

Voir l'offre en entier

Missions, profil recherché, compétences et avantages — en créant votre compte gratuitement.

6 caractères minimum. Plus il est long, plus il est sûr.
ou

Déjà un compte ?

Offres similaires

D'autres postes qui pourraient vous convenir.

Voir tout →

Votre lieu

Les offres et entreprises seront filtrées sur ce pays.

Suggérés

Tous les pays 65