本文へスキップ

Principal Machine Learning Engineer, Distributed vLLM Inference

この求人はあなたに合っていますか?

履歴書を作成すると、この求人との適合度が表示されます。ほかのすべての求人についても同様です。

履歴書を作る

仕事内容

Develop and maintain distributed inference infrastructure leveraging Kubernetes APIs, operators, and the Gateway Inference Extension API for scalable LLM deployments.
Build system components in Go and/or Rust to integrate with the vLLM project and manage distributed workloads.
Design and implement KV cache-aware routing and scoring to optimize memory usage and request distribution at scale.
Enhance resource utilization, fault tolerance, and stability of the inference stack.
Contribute to the design, development, and testing of various inference optimization algorithms; actively participate in technical design discussions.
Provide mentorship and conduct code reviews to engineers, fostering a culture of continuous learning and innovation.

求人の全文を見る

業務内容、求める人物像、スキル、待遇 — 無料アカウントの作成で閲覧できます。

6文字以上。長いほど安全です。
または

すでにアカウントをお持ちですか?

類似の求人

あなたに合いそうな他の職種。

すべて見る →

あなたの地域

求人と企業がこの国で絞り込まれます。

おすすめ

すべての国 69