Senior Research Scientist, Reward Models
Is this job for you?
Build my CV Build your CV and see how well you match this role — and every other one.
The role
Lead cutting-edge research on reward model architectures and RLHF training for large language models.
Develop rubric-based grading and evaluation methods to improve interpretability and consistency.
Study reward hacking, specification gaming, and robust alignment strategies.
Design large-scale experiments to assess generalization, robustness, and failure modes.
Collaborate with Finetuning and Alignment Science teams to translate insights into production improvements.
Mentor researchers, publish findings, and grow institutional knowledge around reward modeling.
Develop rubric-based grading and evaluation methods to improve interpretability and consistency.
Study reward hacking, specification gaming, and robust alignment strategies.
Design large-scale experiments to assess generalization, robustness, and failure modes.
Collaborate with Finetuning and Alignment Science teams to translate insights into production improvements.
Mentor researchers, publish findings, and grow institutional knowledge around reward modeling.
See the full job post
Responsibilities, requirements, skills and benefits — create your free account.
or
Already have an account?
Log inSimilar openings
Other roles that could suit you.
Remote workPartial
CitySan Francisco, United States