AI Engineer, Product
Questa offerta fa per te?
Crea il mio CV Crea il tuo CV e scopri la tua percentuale di corrispondenza con questa posizione — e con tutte le altre.
La posizione
Contribute to the product team's LLM evaluation framework by designing reference tests, heuristics, and model-graded checks.
Define and monitor metrics such as task success, helpfulness, safety flags, latency, and cost to drive data-driven decisions.
Plan and execute A/B tests for prompts, models, and system prompts, with clear rollout or rollback recommendations.
Implement end-to-end observability for LLM calls, including structured logging, tracing, dashboards, and alerts.
Manage model release processes with canary/shadow traffic, sign-offs, SLO-based rollback criteria, and regression detection.
Improve core behaviors (memory policies, intent classification, follow-ups, routing, and tool-call reliability) and create reusable eval templates for teams.
Define and monitor metrics such as task success, helpfulness, safety flags, latency, and cost to drive data-driven decisions.
Plan and execute A/B tests for prompts, models, and system prompts, with clear rollout or rollback recommendations.
Implement end-to-end observability for LLM calls, including structured logging, tracing, dashboards, and alerts.
Manage model release processes with canary/shadow traffic, sign-offs, SLO-based rollback criteria, and regression detection.
Improve core behaviors (memory policies, intent classification, follow-ups, routing, and tool-call reliability) and create reusable eval templates for teams.
Vedi l'annuncio completo
Mansioni, profilo, competenze e vantaggi — crea il tuo account gratuito.
o
Hai già un account?
AccediOfferte simili
Altre posizioni che potrebbero interessarti.
Lavoro da remotoPartial
CittàParis and London, France, United Kingdom