跳到正文

Research Scientist, Frontier, Zurich

这个职位适合你吗?

创建简历,即可看到你与这个职位——以及其他所有职位——的匹配度。

创建我的简历

职位介绍

Design and validate novel post-training pipelines (SFT, RLHF, RLAIF) for frontier-class models where no teacher model exists.
Lead research into next-generation Reward Models, exploring architectures, reducing reward hacking, and improving data signal-to-noise ratios.
Develop methods to enhance internal reasoning and Chain-of-Thought capabilities, emphasizing correctness and multi-step problem solving.
Revamp RL paradigms and prompts to maximize performance while maintaining alignment and safety.
Create robust mechanisms to transform user signals into training data, building a flywheel that scales without regression or bias.
Collaborate across teams to apply these recipes to various sizes and modalities, including audio and multimodal tasks.

查看完整职位

工作职责、任职要求、技能与福利 — 免费创建账号即可查看。

至少 6 个字符。越长越安全。

已有账户?

相似职位

其他可能适合您的职位。

查看全部 →

您的城市

职位和企业将按该国家筛选。

推荐

所有国家 66