Research Engineer, Reward Models Platform
Is this job for you?
Build my CV Build your CV and see how well you match this role — and every other one.
The role
Partner with researchers in the Rewards and Fine-Tuning teams to understand workflows and automate high-friction tasks.
Design and build scalable infrastructure to experiment with reward signals, including rubric editors, data analysis tools, and robustness evaluations.
Develop automated quality checks to detect reward hacks and other pathologies during training.
Create tooling to compare reward methodologies (preference models, rubrics, programmatic rewards) and quantify their impact on model behavior.
Build end-to-end pipelines from dataset preparation to evaluation and deployment, with monitoring and observability to surface issues.
Collaborate with scientists to translate research requirements into platform capabilities, balancing speed, reliability, and maintainability.
Design and build scalable infrastructure to experiment with reward signals, including rubric editors, data analysis tools, and robustness evaluations.
Develop automated quality checks to detect reward hacks and other pathologies during training.
Create tooling to compare reward methodologies (preference models, rubrics, programmatic rewards) and quantify their impact on model behavior.
Build end-to-end pipelines from dataset preparation to evaluation and deployment, with monitoring and observability to surface issues.
Collaborate with scientists to translate research requirements into platform capabilities, balancing speed, reliability, and maintainability.
See the full job post
Responsibilities, requirements, skills and benefits — create your free account.
or
Already have an account?
Log inSimilar openings
Other roles that could suit you.