Research Engineer, Reward Models Training
Cette offre est-elle pour vous ?
Créer mon CV Créez votre CV et découvrez votre pourcentage de correspondance avec ce poste — et avec tous les autres.
Suivez vos candidatures sur mobile L'application Whileresume, gratuite, sur iPhone et Android.
Le poste
Lead the end-to-end engineering of reward model training, from data ingestion to deployment and evaluation.
Design scalable, reliable training pipelines capable of supporting larger models and multiple data modalities.
Build robust data pipelines for collecting, processing, and integrating human feedback into reward model training.
Optimize training infrastructure for throughput, efficiency, and fault tolerance across distributed systems.
Collaborate with researchers to translate novel reward modeling techniques into production-ready systems.
Develop tooling and monitoring to ensure training quality and accelerate iteration cycles.
Design scalable, reliable training pipelines capable of supporting larger models and multiple data modalities.
Build robust data pipelines for collecting, processing, and integrating human feedback into reward model training.
Optimize training infrastructure for throughput, efficiency, and fault tolerance across distributed systems.
Collaborate with researchers to translate novel reward modeling techniques into production-ready systems.
Develop tooling and monitoring to ensure training quality and accelerate iteration cycles.
Voir l'offre en entier
Missions, profil recherché, compétences et avantages — en créant votre compte gratuitement.
ou
Déjà un compte ?
Se connecterCes offres pourraient vous intéresser
Aucune offre vraiment proche pour l'instant — voici les plus récentes.