Skip to content

Machine Learning Engineer | Python | PyTorch | Distributed Training | GPU | Hybrid

Is this job for you?

Build your CV and see how well you match this role — and every other one.

Build my CV

The role

Productize and optimize models from research into reliable, performant, and cost-efficient services with clear SLOs.
Scale training across nodes/GPUs (DDP/FSDP/ZeRO, pipeline/tensor parallelism) and own throughput/time-to-train via profiling and optimization.
Implement model-efficiency techniques (quantization, distillation, pruning, KV-cache, Flash Attention) for training and inference without materially degrading quality.
Build and maintain model-serving systems (vLLM/Triton/ONNX/TensorRT/AITemplate) with batching, streaming, caching, and memory management.
Integrate with vector/feature stores and data pipelines (FAISS/Milvus/Pinecone/pgvector; Parquet/Delta) as needed for production.
Define and track performance and cost KPIs; run continuous improvement loops and capacity planning; partner with ML Ops and scientists on handoffs and evaluations.

See the full job post

Responsibilities, requirements, skills and benefits — create your free account.

At least 6 characters. The longer, the safer.
or

Already have an account?

Similar openings

Other roles that could suit you.

See all →

Your location

Jobs and companies will be filtered on this country.

Suggested

All countries 65