ML Engineer - Inference Serving
Is this job for you?
Build my CV Build your CV and see how well you match this role — and every other one.
The role
As an ML Engineer - Inference Serving, you will ship new model architectures by integrating them into our inference engine. Collaborate across research, engineering, and infrastructure to optimize model efficiency and deployment at scale. Build internal tooling to measure, profile, and track the lifetime of inference jobs and workflows. Automate, test, and maintain inference services to ensure maximum uptime and reliability. Design deployment workflows and sophisticated scheduling to optimally leverage GPU resources while meeting SLOs. Develop CI/CD pipelines for processing model checkpoints, platform components, and internal SDKs.
See the full job post
Responsibilities, requirements, skills and benefits — create your free account.
Already have an account? Log in
Remote workno
CityPalo Alto, US