(USA) Principal, Data Scientist | Conversational AI
¿Es esta oferta para usted?
Crear mi CV Cree su CV y descubra su porcentaje de coincidencia con este puesto — y con todos los demás.
El puesto
Design and implement state-of-the-art evaluation architectures for conversational agents, leveraging LLM-as-a-judge and hybrid scoring systems.
Develop high-precision prompts and calibration methods to ensure high inter-rater reliability against human judgments.
Lead distillation and optimization of smaller judge models to balance accuracy, latency, and cost.
Curate large-scale dialogue datasets, create Golden Sets, and define clear ground-truth annotation instructions for subjective tasks.
Collaborate with engineering to embed quality signals into CI/CD pipelines, enabling automated regression testing and production monitoring.
Perform failure-mode analyses and drive actionable improvements while mentoring teams and shaping evaluation practices.
Develop high-precision prompts and calibration methods to ensure high inter-rater reliability against human judgments.
Lead distillation and optimization of smaller judge models to balance accuracy, latency, and cost.
Curate large-scale dialogue datasets, create Golden Sets, and define clear ground-truth annotation instructions for subjective tasks.
Collaborate with engineering to embed quality signals into CI/CD pipelines, enabling automated regression testing and production monitoring.
Perform failure-mode analyses and drive actionable improvements while mentoring teams and shaping evaluation practices.
Ver la oferta completa
Funciones, perfil, competencias y ventajas — crea tu cuenta gratis.
¿Ya tiene una cuenta? Iniciar sesión
Ofertas similares
Otros puestos que podrían encajar.
Teletrabajono
CiudadSunnyvale, United States