Zum Inhalt springen

(USA) Principal, Data Scientist | Gen AI Vision

Ist diese Stelle etwas für Sie?

Erstellen Sie Ihren Lebenslauf und entdecken Sie Ihre Übereinstimmung mit dieser Stelle — und mit allen anderen.

Lebenslauf erstellen

Die Stelle

Lead the design and implementation of evaluation architectures for conversational agents using LLM-as-a-judge, defining robust, scalable metrics.
Own prompt engineering and calibration to achieve high inter-rater reliability and alignment with human judgments.
Drive model distillation and optimization to create cost-effective Judge models balancing accuracy, latency, and budget.
Curate large-scale datasets and Golden Sets with clear annotation instructions to standardize ground truth for subjective tasks.
Collaborate with engineering to embed quality signals into CI/CD pipelines, enabling automated regression testing and monitoring in production.
Perform failure mode analyses (hallucinations, tool misuse, safety violations), extract insights, mentor teams, and advance best practices for evaluation.

Die vollständige Anzeige sehen

Aufgaben, Anforderungen, Kompetenzen und Vorteile — mit Ihrem kostenlosen Konto.

Mindestens 6 Zeichen. Je länger, desto sicherer.
oder

Bereits ein Konto?

Ähnliche Stellen

Weitere Positionen, die passen könnten.

Alle ansehen →

Ihr Ort

Stellen und Unternehmen werden nach diesem Land gefiltert.

Vorschläge

Alle Länder 66