Naar de inhoud

Principal Machine Learning Engineer, Distributed vLLM Inference

Is deze vacature iets voor u?

Maak uw cv en ontdek uw matchpercentage met deze functie — en met alle andere.

Mijn cv maken

De functie

Develop and maintain distributed inference infrastructure leveraging Kubernetes APIs, operators, and the Gateway Inference Extension API for scalable LLM deployments.
Build system components in Go and/or Rust to integrate with the vLLM project and manage distributed workloads.
Design and implement KV cache-aware routing and scoring to optimize memory usage and request distribution at scale.
Enhance resource utilization, fault tolerance, and stability of the inference stack.
Contribute to the design, development, and testing of various inference optimization algorithms; actively participate in technical design discussions.
Provide mentorship and conduct code reviews to engineers, fostering a culture of continuous learning and innovation.

Bekijk de volledige vacature

Taken, profiel, vaardigheden en voordelen — maak gratis een account aan.

Minimaal 6 tekens. Hoe langer, hoe veiliger.
of

Al een account?

Vergelijkbare vacatures

Andere functies die kunnen passen.

Alles bekijken →

Uw locatie

Vacatures en bedrijven worden op dit land gefilterd.

Voorgesteld

Alle landen 68