(USA) Principal, Data Scientist | Gen AI Vision
仕事内容
Lead the design and implementation of evaluation architectures for conversational agents using LLM-as-a-judge, defining robust, scalable metrics.
Own prompt engineering and calibration to achieve high inter-rater reliability and alignment with human judgments.
Drive model distillation and optimization to create cost-effective Judge models balancing accuracy, latency, and budget.
Curate large-scale datasets and Golden Sets with clear annotation instructions to standardize ground truth for subjective tasks.
Collaborate with engineering to embed quality signals into CI/CD pipelines, enabling automated regression testing and monitoring in production.
Perform failure mode analyses (hallucinations, tool misuse, safety violations), extract insights, mentor teams, and advance best practices for evaluation.
Own prompt engineering and calibration to achieve high inter-rater reliability and alignment with human judgments.
Drive model distillation and optimization to create cost-effective Judge models balancing accuracy, latency, and budget.
Curate large-scale datasets and Golden Sets with clear annotation instructions to standardize ground truth for subjective tasks.
Collaborate with engineering to embed quality signals into CI/CD pipelines, enabling automated regression testing and monitoring in production.
Perform failure mode analyses (hallucinations, tool misuse, safety violations), extract insights, mentor teams, and advance best practices for evaluation.
求人の全文を見る
業務内容、求める人物像、スキル、待遇 — 無料アカウントの作成で閲覧できます。
または
すでにアカウントをお持ちですか?
ログイン類似の求人
あなたに合いそうな他の職種。
リモートワークno
勤務地Sunnyvale, United States