Ir al contenido

Research Scientist, Frontier, Zurich

¿Es esta oferta para usted?

Cree su CV y descubra su porcentaje de coincidencia con este puesto — y con todos los demás.

Crear mi CV

El puesto

Design and validate novel post-training pipelines (SFT, RLHF, RLAIF) for frontier-class models where no teacher model exists.
Lead research into next-generation Reward Models, exploring architectures, reducing reward hacking, and improving data signal-to-noise ratios.
Develop methods to enhance internal reasoning and Chain-of-Thought capabilities, emphasizing correctness and multi-step problem solving.
Revamp RL paradigms and prompts to maximize performance while maintaining alignment and safety.
Create robust mechanisms to transform user signals into training data, building a flywheel that scales without regression or bias.
Collaborate across teams to apply these recipes to various sizes and modalities, including audio and multimodal tasks.

Ver la oferta completa

Funciones, perfil, competencias y ventajas — crea tu cuenta gratis.

Mínimo 8 caracteres, con una mayúscula y una cifra.

¿Ya tiene una cuenta? Iniciar sesión

Ofertas similares

Otros puestos que podrían encajar.

Ver todo →

Su ubicación

Las ofertas y empresas se filtrarán por este país.

Sugeridos

Todos los países 66