跳到正文

(USA) Principal, Data Scientist | Gen AI Vision

这个职位适合你吗?

创建简历,即可看到你与这个职位——以及其他所有职位——的匹配度。

创建我的简历

职位介绍

Lead the design and implementation of evaluation architectures for conversational agents using LLM-as-a-judge, defining robust, scalable metrics.
Own prompt engineering and calibration to achieve high inter-rater reliability and alignment with human judgments.
Drive model distillation and optimization to create cost-effective Judge models balancing accuracy, latency, and budget.
Curate large-scale datasets and Golden Sets with clear annotation instructions to standardize ground truth for subjective tasks.
Collaborate with engineering to embed quality signals into CI/CD pipelines, enabling automated regression testing and monitoring in production.
Perform failure mode analyses (hallucinations, tool misuse, safety violations), extract insights, mentor teams, and advance best practices for evaluation.

查看完整职位

工作职责、任职要求、技能与福利 — 免费创建账号即可查看。

至少 8 个字符,需包含一个大写字母和一个数字。

已有账户? 登录

相似职位

其他可能适合您的职位。

查看全部 →

您的城市

职位和企业将按该国家筛选。

推荐

所有国家 66