Senior Machine Learning Engineer - Evaluation
Is this job for you?
Build my CV Build your CV and see how well you match this role — and every other one.
The role
Lead the design and implementation of offline evaluation pipelines for embodied AI models, building scalable and interpretable benchmarks that cover action, perception, and language-grounded reasoning. Design metrics across vision, language, and driving tasks to robustly evaluate model behavior and inform deployment readiness. Drive human annotation workflows, including task design, QA, and coordination with internal teams and external partners, to ensure high-quality ground-truth data. Collaborate with science, datasets, and infrastructure teams to align evaluation with product goals and system safety. Analyze offline metrics and correlate them with online performance to guide model selection and risk assessment. Demonstrate strong Python software engineering and data processing skills to deliver scalable, reproducible evaluation tooling.
See the full job post
Responsibilities, requirements, skills and benefits — create your free account.
or
Already have an account?
Log inSimilar openings
Other roles that could suit you.