Aller au contenu
Whileresume
Classement Créer mon CV Recruter Se connecter

Machine Learning Engineer | Python | PyTorch | Distributed Training | GPU | Hybrid

Cette offre est-elle pour vous ?

Créez votre CV et découvrez votre pourcentage de correspondance avec ce poste — et avec tous les autres.

Créer mon CV
Suivez vos candidatures sur mobile L'application Whileresume, gratuite, sur iPhone et Android.

Le poste

Productize and optimize models from research into reliable, performant, and cost-efficient services with clear SLOs.
Scale training across nodes/GPUs (DDP/FSDP/ZeRO, pipeline/tensor parallelism) and own throughput/time-to-train via profiling and optimization.
Implement model-efficiency techniques (quantization, distillation, pruning, KV-cache, Flash Attention) for training and inference without materially degrading quality.
Build and maintain model-serving systems (vLLM/Triton/ONNX/TensorRT/AITemplate) with batching, streaming, caching, and memory management.
Integrate with vector/feature stores and data pipelines (FAISS/Milvus/Pinecone/pgvector; Parquet/Delta) as needed for production.
Define and track performance and cost KPIs; run continuous improvement loops and capacity planning; partner with ML Ops and scientists on handoffs and evaluations.

Voir l'offre en entier

Missions, profil recherché, compétences et avantages — en créant votre compte gratuitement.

6 caractères minimum. Plus il est long, plus il est sûr.
ou

Déjà un compte ?

Ces offres pourraient vous intéresser

Aucune offre vraiment proche pour l'instant — voici les plus récentes.

Votre lieu

Les offres et entreprises seront filtrées sur ce pays.

Suggérés

Tous les pays 65