Machine Learning Engineer, Reinforcement Learning & Reward Modeling
Is deze vacature iets voor u?
Mijn cv maken Maak uw cv en ontdek uw matchpercentage met deze functie — en met alle andere.
Volg uw sollicitaties op mobiel De gratis Whileresume-app, op iPhone en Android.
De functie
Design and optimize end-to-end pipelines for training reward models and RL agents to be reproducible and high-throughput.
Develop tooling for data processing, annotation, and inference within RL workflows.
Build, refine, and deploy reward models that encode safe, interpretable, and effective driving behaviours.
Integrate reward models with diverse data sources: real-world trajectories, simulation, and synthetic datasets.
Conduct ablations, hyperparameter explorations, and controlled studies to analyse how reward structures and training dynamics affect policy performance.
Diagnose failure modes, iterate on reward objectives, and partner with RL scientists to translate ideas into scalable engineering solutions and testing frameworks.
Develop tooling for data processing, annotation, and inference within RL workflows.
Build, refine, and deploy reward models that encode safe, interpretable, and effective driving behaviours.
Integrate reward models with diverse data sources: real-world trajectories, simulation, and synthetic datasets.
Conduct ablations, hyperparameter explorations, and controlled studies to analyse how reward structures and training dynamics affect policy performance.
Diagnose failure modes, iterate on reward objectives, and partner with RL scientists to translate ideas into scalable engineering solutions and testing frameworks.
Bekijk de volledige vacature
Taken, profiel, vaardigheden en voordelen — maak gratis een account aan.
of
Al een account?
InloggenDeze vacatures zijn misschien iets voor u
Nog geen echt vergelijkbare vacatures — hier zijn de nieuwste.