跳到正文

Principal Machine Learning Engineer, Distributed vLLM Inference

这个职位适合你吗?

创建简历,即可看到你与这个职位——以及其他所有职位——的匹配度。

创建我的简历

职位介绍

Develop and maintain distributed inference infrastructure leveraging Kubernetes APIs, operators, and the Gateway Inference Extension API for scalable LLM deployments.
Build system components in Go and/or Rust to integrate with the vLLM project and manage distributed workloads.
Design and implement KV cache-aware routing and scoring to optimize memory usage and request distribution at scale.
Enhance resource utilization, fault tolerance, and stability of the inference stack.
Contribute to the design, development, and testing of various inference optimization algorithms; actively participate in technical design discussions.
Provide mentorship and conduct code reviews to engineers, fostering a culture of continuous learning and innovation.

查看完整职位

工作职责、任职要求、技能与福利 — 免费创建账号即可查看。

至少 6 个字符。越长越安全。

已有账户?

相似职位

其他可能适合您的职位。

查看全部 →

您的城市

职位和企业将按该国家筛选。

推荐

所有国家 66