Research Scientist, Frontier, Zurich
Cette offre est-elle pour vous ?
Créer mon CV Créez votre CV et découvrez votre pourcentage de correspondance avec ce poste — et avec tous les autres.
Le poste
Design and validate novel post-training pipelines (SFT, RLHF, RLAIF) for frontier-class models where no teacher model exists.
Lead research into next-generation Reward Models, exploring architectures, reducing reward hacking, and improving data signal-to-noise ratios.
Develop methods to enhance internal reasoning and Chain-of-Thought capabilities, emphasizing correctness and multi-step problem solving.
Revamp RL paradigms and prompts to maximize performance while maintaining alignment and safety.
Create robust mechanisms to transform user signals into training data, building a flywheel that scales without regression or bias.
Collaborate across teams to apply these recipes to various sizes and modalities, including audio and multimodal tasks.
Lead research into next-generation Reward Models, exploring architectures, reducing reward hacking, and improving data signal-to-noise ratios.
Develop methods to enhance internal reasoning and Chain-of-Thought capabilities, emphasizing correctness and multi-step problem solving.
Revamp RL paradigms and prompts to maximize performance while maintaining alignment and safety.
Create robust mechanisms to transform user signals into training data, building a flywheel that scales without regression or bias.
Collaborate across teams to apply these recipes to various sizes and modalities, including audio and multimodal tasks.
Voir l'offre en entier
Missions, profil recherché, compétences et avantages — en créant votre compte gratuitement.
ou
Déjà un compte ?
Se connecterOffres similaires
D'autres postes qui pourraient vous convenir.
Télétravailno
VilleZurich, Suisse