本文へスキップ

(USA) Principal, Data Scientist | Gen AI Vision

この求人はあなたに合っていますか?

履歴書を作成すると、この求人との適合度が表示されます。ほかのすべての求人についても同様です。

履歴書を作る

仕事内容

Lead the design and implementation of evaluation architectures for conversational agents using LLM-as-a-judge, defining robust, scalable metrics.
Own prompt engineering and calibration to achieve high inter-rater reliability and alignment with human judgments.
Drive model distillation and optimization to create cost-effective Judge models balancing accuracy, latency, and budget.
Curate large-scale datasets and Golden Sets with clear annotation instructions to standardize ground truth for subjective tasks.
Collaborate with engineering to embed quality signals into CI/CD pipelines, enabling automated regression testing and monitoring in production.
Perform failure mode analyses (hallucinations, tool misuse, safety violations), extract insights, mentor teams, and advance best practices for evaluation.

求人の全文を見る

業務内容、求める人物像、スキル、待遇 — 無料アカウントの作成で閲覧できます。

6文字以上。長いほど安全です。
または

すでにアカウントをお持ちですか?

類似の求人

あなたに合いそうな他の職種。

すべて見る →

あなたの地域

求人と企業がこの国で絞り込まれます。

おすすめ

すべての国 69