Ir al contenido
Whileresume
Clasificación Crear mi CV Contratar Iniciar sesión

Senior Research Scientist, Reward Models

¿Es esta oferta para usted?

Cree su CV y descubra su porcentaje de coincidencia con este puesto — y con todos los demás.

Crear mi CV
Sigue tus candidaturas desde el móvil La app gratuita de Whileresume, en iPhone y Android.

El puesto

Lead research on novel reward model architectures and RLHF training approaches for large language models.
Develop and evaluate LLM-based grading and evaluation methods, including rubric-driven approaches that improve consistency and interpretability.
Research techniques to detect, characterize, and mitigate reward hacking and specification gaming.
Design experiments to understand reward model generalization, robustness, and failure modes.
Collaborate with the Finetuning team to translate research insights into improvements for production training pipelines.
Contribute to research publications, blog posts, and internal documentation; mentor other researchers and help build institutional knowledge around reward modeling.

Ver la oferta completa

Funciones, perfil, competencias y ventajas — crea tu cuenta gratis.

6 caracteres como mínimo. Cuanto más largo, más seguro.
o

¿Ya tiene una cuenta?

Estas ofertas podrían interesarte

Ninguna oferta realmente parecida por ahora — estas son las más recientes.

Su ubicación

Las ofertas y empresas se filtrarán por este país.

Sugeridos

Todos los países 66