AI Engineer, Product
Is this job for you?
Build my CV Build your CV and see how well you match this role — and every other one.
The role
Contribute to the product team's LLM evaluation framework by designing reference tests, heuristics, and model-graded checks.
Define and monitor metrics such as task success, helpfulness, safety flags, latency, and cost to drive data-driven decisions.
Plan and execute A/B tests for prompts, models, and system prompts, with clear rollout or rollback recommendations.
Implement end-to-end observability for LLM calls, including structured logging, tracing, dashboards, and alerts.
Manage model release processes with canary/shadow traffic, sign-offs, SLO-based rollback criteria, and regression detection.
Improve core behaviors (memory policies, intent classification, follow-ups, routing, and tool-call reliability) and create reusable eval templates for teams.
Define and monitor metrics such as task success, helpfulness, safety flags, latency, and cost to drive data-driven decisions.
Plan and execute A/B tests for prompts, models, and system prompts, with clear rollout or rollback recommendations.
Implement end-to-end observability for LLM calls, including structured logging, tracing, dashboards, and alerts.
Manage model release processes with canary/shadow traffic, sign-offs, SLO-based rollback criteria, and regression detection.
Improve core behaviors (memory policies, intent classification, follow-ups, routing, and tool-call reliability) and create reusable eval templates for teams.
See the full job post
Responsibilities, requirements, skills and benefits — create your free account.
or
Already have an account?
Log inSimilar openings
Other roles that could suit you.
Remote workPartial
CityParis and London, France, United Kingdom