(USA) Principal, Data Scientist - Gen AI Vision
¿Es esta oferta para usted?
Crear mi CV Cree su CV y descubra su porcentaje de coincidencia con este puesto — y con todos los demás.
El puesto
Own the evaluation architecture for AI agents, designing how success is measured across non-deterministic outputs.
Lead the brain that critiques agents with a mix of LLM-based judging, human benchmarks, and automated pipelines.
Develop evaluation pipelines, calibrate prompts, and distill smaller judge models balancing accuracy, latency, and cost.
Curate large-scale conversation data to build reliable Golden Set datasets and standardized ground truth for subjective tasks.
Collaborate with engineering to integrate quality signals into CI/CD pipelines for automated testing and monitoring.
Drive insights to identify systemic weaknesses, mentor data scientists, and advance best practices in AI evaluation.
Lead the brain that critiques agents with a mix of LLM-based judging, human benchmarks, and automated pipelines.
Develop evaluation pipelines, calibrate prompts, and distill smaller judge models balancing accuracy, latency, and cost.
Curate large-scale conversation data to build reliable Golden Set datasets and standardized ground truth for subjective tasks.
Collaborate with engineering to integrate quality signals into CI/CD pipelines for automated testing and monitoring.
Drive insights to identify systemic weaknesses, mentor data scientists, and advance best practices in AI evaluation.
Ver la oferta completa
Funciones, perfil, competencias y ventajas — crea tu cuenta gratis.
o
¿Ya tiene una cuenta?
Iniciar sesiónOfertas similares
Otros puestos que podrían encajar.
Teletrabajono
CiudadSunnyvale, United States