(USA) Principal, Data Scientist | Conversational AI
仕事内容
Design and implement state-of-the-art evaluation architectures for conversational agents, leveraging LLM-as-a-judge and hybrid scoring systems.
Develop high-precision prompts and calibration methods to ensure high inter-rater reliability against human judgments.
Lead distillation and optimization of smaller judge models to balance accuracy, latency, and cost.
Curate large-scale dialogue datasets, create Golden Sets, and define clear ground-truth annotation instructions for subjective tasks.
Collaborate with engineering to embed quality signals into CI/CD pipelines, enabling automated regression testing and production monitoring.
Perform failure-mode analyses and drive actionable improvements while mentoring teams and shaping evaluation practices.
Develop high-precision prompts and calibration methods to ensure high inter-rater reliability against human judgments.
Lead distillation and optimization of smaller judge models to balance accuracy, latency, and cost.
Curate large-scale dialogue datasets, create Golden Sets, and define clear ground-truth annotation instructions for subjective tasks.
Collaborate with engineering to embed quality signals into CI/CD pipelines, enabling automated regression testing and production monitoring.
Perform failure-mode analyses and drive actionable improvements while mentoring teams and shaping evaluation practices.
求人の全文を見る
業務内容、求める人物像、スキル、待遇 — 無料アカウントの作成で閲覧できます。
すでにアカウントをお持ちですか? ログイン
類似の求人
あなたに合いそうな他の職種。
リモートワークno
勤務地Sunnyvale, United States