Skip to content

Research Scientist, Frontier, Zurich

Is this job for you?

Build your CV and see how well you match this role — and every other one.

Build my CV

The role

Design and validate novel post-training pipelines (SFT, RLHF, RLAIF) for frontier-class models where no teacher model exists.
Lead research into next-generation Reward Models, exploring architectures, reducing reward hacking, and improving data signal-to-noise ratios.
Develop methods to enhance internal reasoning and Chain-of-Thought capabilities, emphasizing correctness and multi-step problem solving.
Revamp RL paradigms and prompts to maximize performance while maintaining alignment and safety.
Create robust mechanisms to transform user signals into training data, building a flywheel that scales without regression or bias.
Collaborate across teams to apply these recipes to various sizes and modalities, including audio and multimodal tasks.

See the full job post

Responsibilities, requirements, skills and benefits — create your free account.

At least 8 characters, one uppercase letter and one digit.

Already have an account? Log in

Similar openings

Other roles that could suit you.

See all →

Your location

Jobs and companies will be filtered on this country.

Suggested

All countries 65