Skip to content

(USA) Principal, Data Scientist | Gen AI Vision

Is this job for you?

Build your CV and see how well you match this role — and every other one.

Build my CV

The role

Lead the design and implementation of evaluation architectures for conversational agents using LLM-as-a-judge, defining robust, scalable metrics.
Own prompt engineering and calibration to achieve high inter-rater reliability and alignment with human judgments.
Drive model distillation and optimization to create cost-effective Judge models balancing accuracy, latency, and budget.
Curate large-scale datasets and Golden Sets with clear annotation instructions to standardize ground truth for subjective tasks.
Collaborate with engineering to embed quality signals into CI/CD pipelines, enabling automated regression testing and monitoring in production.
Perform failure mode analyses (hallucinations, tool misuse, safety violations), extract insights, mentor teams, and advance best practices for evaluation.

See the full job post

Responsibilities, requirements, skills and benefits — create your free account.

At least 8 characters, one uppercase letter and one digit.

Already have an account? Log in

Similar openings

Other roles that could suit you.

See all →

Your location

Jobs and companies will be filtered on this country.

Suggested

All countries 65