Skip to content

Senior Software Engineer (MLOps) – Annotation & Evaluation

Is this job for you?

Build your CV and see how well you match this role — and every other one.

Build my CV

The role

Design and build systems for automated evaluation of AI models, including LLMs and agents, using production-like telemetry and realistic scenarios.
Lead the development of benchmark suites, evaluation pipelines, and model comparison tools with integrated trust & safety metrics.
Build and maintain integrations with labeling systems (e.g., Label Studio) and coordinate with external/internal annotation workflows.
Collaborate with Applied AI and Bits AI teams to enable fast iteration, reproducible experiments, and interpretable evaluations.
Develop data pipelines that feed metrics, results, and alerts into our observability stack to monitor model behavior at scale.
Promote safe deployment practices via bias checks, hallucination detection, and human-in-the-loop review.

See the full job post

Responsibilities, requirements, skills and benefits — create your free account.

At least 8 characters, one uppercase letter and one digit.

Already have an account? Log in

Similar openings

Other roles that could suit you.

See all →

Your location

Jobs and companies will be filtered on this country.

Suggested

All countries 65