AI Evaluation Engineer
Is this job for you?
Build my CV Build your CV and see how well you match this role — and every other one.
Track your applications on mobile The free Whileresume app, on iPhone and Android.
The role
Mindrift connects specialists with project-based AI opportunities in testing, evaluating, and improving AI systems. The role involves building datasets to evaluate AI coding agents, creating realistic developer environments, and crafting challenging tasks and evaluation criteria. Candidates will review agent solutions, analyse failures, and refine tasks based on feedback, requiring a deep understanding of where models fail. Experience in software development with Python, JavaScript/TypeScript, Docker, and databases is essential. The position requires English proficiency of B2+ and offers flexible hours at competitive rates. This role is not about data labelling, prompt engineering, or writing code from scratch, but rather guiding and evaluating AI-generated code.
See the full job post
Responsibilities, requirements, skills and benefits — create your free account.
or
Already have an account?
Log inYou might also like these jobs
No closely matching jobs yet — here are the most recent ones.