Research Engineer, AI Safety & Alignment
Is this job for you?
Build my CV Build your CV and see how well you match this role — and every other one.
Track your applications on mobile The free Whileresume app, on iPhone and Android.
The role
Develop and implement novel evaluation methodologies and metrics to assess safety and alignment of large language models.
Research and build techniques for model alignment, value learning, and interpretability.
Conduct adversarial testing to proactively uncover vulnerabilities and failure modes.
Analyze and mitigate biases, toxicity, and other harmful behaviors through RLHF and fine-tuning.
Collaborate with engineering and product teams to translate safety research into scalable, production-ready solutions.
Stay current with AI safety advances and contribute to the academic community through publications and talks.
Research and build techniques for model alignment, value learning, and interpretability.
Conduct adversarial testing to proactively uncover vulnerabilities and failure modes.
Analyze and mitigate biases, toxicity, and other harmful behaviors through RLHF and fine-tuning.
Collaborate with engineering and product teams to translate safety research into scalable, production-ready solutions.
Stay current with AI safety advances and contribute to the academic community through publications and talks.
See the full job post
Responsibilities, requirements, skills and benefits — create your free account.
or
Already have an account?
Log inYou might also like these jobs
No closely matching jobs yet — here are the most recent ones.