Naar de inhoud

Research Scientist, Frontier, Zurich

Is deze vacature iets voor u?

Maak uw cv en ontdek uw matchpercentage met deze functie — en met alle andere.

Mijn cv maken

De functie

Design and validate novel post-training pipelines (SFT, RLHF, RLAIF) for frontier-class models where no teacher model exists.
Lead research into next-generation Reward Models, exploring architectures, reducing reward hacking, and improving data signal-to-noise ratios.
Develop methods to enhance internal reasoning and Chain-of-Thought capabilities, emphasizing correctness and multi-step problem solving.
Revamp RL paradigms and prompts to maximize performance while maintaining alignment and safety.
Create robust mechanisms to transform user signals into training data, building a flywheel that scales without regression or bias.
Collaborate across teams to apply these recipes to various sizes and modalities, including audio and multimodal tasks.

Bekijk de volledige vacature

Taken, profiel, vaardigheden en voordelen — maak gratis een account aan.

Minimaal 6 tekens. Hoe langer, hoe veiliger.

Al een account? Inloggen

Vergelijkbare vacatures

Andere functies die kunnen passen.

Alles bekijken →

Uw locatie

Vacatures en bedrijven worden op dit land gefilterd.

Voorgesteld

Alle landen 68