Ir al contenido

(USA) Principal, Data Scientist | Gen AI Vision

¿Es esta oferta para usted?

Cree su CV y descubra su porcentaje de coincidencia con este puesto — y con todos los demás.

Crear mi CV

El puesto

Lead the design and implementation of evaluation architectures for conversational agents using LLM-as-a-judge, defining robust, scalable metrics.
Own prompt engineering and calibration to achieve high inter-rater reliability and alignment with human judgments.
Drive model distillation and optimization to create cost-effective Judge models balancing accuracy, latency, and budget.
Curate large-scale datasets and Golden Sets with clear annotation instructions to standardize ground truth for subjective tasks.
Collaborate with engineering to embed quality signals into CI/CD pipelines, enabling automated regression testing and monitoring in production.
Perform failure mode analyses (hallucinations, tool misuse, safety violations), extract insights, mentor teams, and advance best practices for evaluation.

Ver la oferta completa

Funciones, perfil, competencias y ventajas — crea tu cuenta gratis.

6 caracteres como mínimo. Cuanto más largo, más seguro.
o

¿Ya tiene una cuenta?

Ofertas similares

Otros puestos que podrían encajar.

Ver todo →

Su ubicación

Las ofertas y empresas se filtrarán por este país.

Sugeridos

Todos los países 66