Vai al contenuto

(USA) Principal, Data Scientist | Gen AI Vision

Questa offerta fa per te?

Crea il tuo CV e scopri la tua percentuale di corrispondenza con questa posizione — e con tutte le altre.

Crea il mio CV

La posizione

Lead the design and implementation of evaluation architectures for conversational agents using LLM-as-a-judge, defining robust, scalable metrics.
Own prompt engineering and calibration to achieve high inter-rater reliability and alignment with human judgments.
Drive model distillation and optimization to create cost-effective Judge models balancing accuracy, latency, and budget.
Curate large-scale datasets and Golden Sets with clear annotation instructions to standardize ground truth for subjective tasks.
Collaborate with engineering to embed quality signals into CI/CD pipelines, enabling automated regression testing and monitoring in production.
Perform failure mode analyses (hallucinations, tool misuse, safety violations), extract insights, mentor teams, and advance best practices for evaluation.

Vedi l'annuncio completo

Mansioni, profilo, competenze e vantaggi — crea il tuo account gratuito.

Minimo 8 caratteri, di cui una maiuscola e una cifra.

Hai già un account? Accedi

Offerte simili

Altre posizioni che potrebbero interessarti.

Vedi tutto →

La tua località

Offerte e aziende saranno filtrate su questo paese.

Suggeriti

Tutti i paesi 68