(USA) Principal, Data Scientist | Conversational AI
Cette offre est-elle pour vous ?
Créer mon CV Créez votre CV et découvrez votre pourcentage de correspondance avec ce poste — et avec tous les autres.
Le poste
Design and implement state-of-the-art evaluation architectures for conversational agents, leveraging LLM-as-a-judge and hybrid scoring systems.
Develop high-precision prompts and calibration methods to ensure high inter-rater reliability against human judgments.
Lead distillation and optimization of smaller judge models to balance accuracy, latency, and cost.
Curate large-scale dialogue datasets, create Golden Sets, and define clear ground-truth annotation instructions for subjective tasks.
Collaborate with engineering to embed quality signals into CI/CD pipelines, enabling automated regression testing and production monitoring.
Perform failure-mode analyses and drive actionable improvements while mentoring teams and shaping evaluation practices.
Develop high-precision prompts and calibration methods to ensure high inter-rater reliability against human judgments.
Lead distillation and optimization of smaller judge models to balance accuracy, latency, and cost.
Curate large-scale dialogue datasets, create Golden Sets, and define clear ground-truth annotation instructions for subjective tasks.
Collaborate with engineering to embed quality signals into CI/CD pipelines, enabling automated regression testing and production monitoring.
Perform failure-mode analyses and drive actionable improvements while mentoring teams and shaping evaluation practices.
Voir l'offre en entier
Missions, profil recherché, compétences et avantages — en créant votre compte gratuitement.
Déjà un compte ? Se connecter
Offres similaires
D'autres postes qui pourraient vous convenir.
Télétravailno
VilleSunnyvale, United States