AI Engineer, Product
¿Es esta oferta para usted?
Crear mi CV Cree su CV y descubra su porcentaje de coincidencia con este puesto — y con todos los demás.
El puesto
Contribute to the product team's LLM evaluation framework by designing reference tests, heuristics, and model-graded checks.
Define and monitor metrics such as task success, helpfulness, safety flags, latency, and cost to drive data-driven decisions.
Plan and execute A/B tests for prompts, models, and system prompts, with clear rollout or rollback recommendations.
Implement end-to-end observability for LLM calls, including structured logging, tracing, dashboards, and alerts.
Manage model release processes with canary/shadow traffic, sign-offs, SLO-based rollback criteria, and regression detection.
Improve core behaviors (memory policies, intent classification, follow-ups, routing, and tool-call reliability) and create reusable eval templates for teams.
Define and monitor metrics such as task success, helpfulness, safety flags, latency, and cost to drive data-driven decisions.
Plan and execute A/B tests for prompts, models, and system prompts, with clear rollout or rollback recommendations.
Implement end-to-end observability for LLM calls, including structured logging, tracing, dashboards, and alerts.
Manage model release processes with canary/shadow traffic, sign-offs, SLO-based rollback criteria, and regression detection.
Improve core behaviors (memory policies, intent classification, follow-ups, routing, and tool-call reliability) and create reusable eval templates for teams.
Ver la oferta completa
Funciones, perfil, competencias y ventajas — crea tu cuenta gratis.
o
¿Ya tiene una cuenta?
Iniciar sesiónOfertas similares
Otros puestos que podrían encajar.
TeletrabajoPartial
CiudadParis and London, France, United Kingdom