Research Engineer, Reward Models Platform
応募状況をスマホで確認 Whileresume の無料アプリ(iPhone・Android)。
仕事内容
Partner with researchers in the Rewards and Fine-Tuning teams to understand workflows and automate high-friction tasks.
Design and build scalable infrastructure to experiment with reward signals, including rubric editors, data analysis tools, and robustness evaluations.
Develop automated quality checks to detect reward hacks and other pathologies during training.
Create tooling to compare reward methodologies (preference models, rubrics, programmatic rewards) and quantify their impact on model behavior.
Build end-to-end pipelines from dataset preparation to evaluation and deployment, with monitoring and observability to surface issues.
Collaborate with scientists to translate research requirements into platform capabilities, balancing speed, reliability, and maintainability.
Design and build scalable infrastructure to experiment with reward signals, including rubric editors, data analysis tools, and robustness evaluations.
Develop automated quality checks to detect reward hacks and other pathologies during training.
Create tooling to compare reward methodologies (preference models, rubrics, programmatic rewards) and quantify their impact on model behavior.
Build end-to-end pipelines from dataset preparation to evaluation and deployment, with monitoring and observability to surface issues.
Collaborate with scientists to translate research requirements into platform capabilities, balancing speed, reliability, and maintainability.
求人の全文を見る
業務内容、求める人物像、スキル、待遇 — 無料アカウントの作成で閲覧できます。
または
すでにアカウントをお持ちですか?
ログインこちらの求人もおすすめです
近い求人はまだありません — 最新の求人をご紹介します。