(USA) Principal, Data Scientist | Conversational AI
Is this job for you?
Build my CV Build your CV and see how well you match this role — and every other one.
The role
Design and implement state-of-the-art evaluation architectures for conversational agents, leveraging LLM-as-a-judge and hybrid scoring systems.
Develop high-precision prompts and calibration methods to ensure high inter-rater reliability against human judgments.
Lead distillation and optimization of smaller judge models to balance accuracy, latency, and cost.
Curate large-scale dialogue datasets, create Golden Sets, and define clear ground-truth annotation instructions for subjective tasks.
Collaborate with engineering to embed quality signals into CI/CD pipelines, enabling automated regression testing and production monitoring.
Perform failure-mode analyses and drive actionable improvements while mentoring teams and shaping evaluation practices.
Develop high-precision prompts and calibration methods to ensure high inter-rater reliability against human judgments.
Lead distillation and optimization of smaller judge models to balance accuracy, latency, and cost.
Curate large-scale dialogue datasets, create Golden Sets, and define clear ground-truth annotation instructions for subjective tasks.
Collaborate with engineering to embed quality signals into CI/CD pipelines, enabling automated regression testing and production monitoring.
Perform failure-mode analyses and drive actionable improvements while mentoring teams and shaping evaluation practices.
See the full job post
Responsibilities, requirements, skills and benefits — create your free account.
or
Already have an account?
Log inSimilar openings
Other roles that could suit you.
Remote workno
CitySunnyvale, United States