Skip to content

Senior Machine Learning Engineer - Evaluation

Is this job for you?

Build your CV and see how well you match this role — and every other one.

Build my CV

The role

Lead the design and implementation of offline evaluation pipelines for embodied AI models, building scalable and interpretable benchmarks that cover action, perception, and language-grounded reasoning. Design metrics across vision, language, and driving tasks to robustly evaluate model behavior and inform deployment readiness. Drive human annotation workflows, including task design, QA, and coordination with internal teams and external partners, to ensure high-quality ground-truth data. Collaborate with science, datasets, and infrastructure teams to align evaluation with product goals and system safety. Analyze offline metrics and correlate them with online performance to guide model selection and risk assessment. Demonstrate strong Python software engineering and data processing skills to deliver scalable, reproducible evaluation tooling.

See the full job post

Responsibilities, requirements, skills and benefits — create your free account.

At least 6 characters. The longer, the safer.
or

Already have an account?

Similar openings

Other roles that could suit you.

See all →

Your location

Jobs and companies will be filtered on this country.

Suggested

All countries 65