Machine Learning Data Engineer, Replica Pipelines
Is this job for you?
Build my CV Build your CV and see how well you match this role — and every other one.
The role
As a Machine Learning Data Engineer for Replica Pipelines, you will design and scale data ingestion and processing pipelines for ML workflows.
You will normalize and validate customer and synthetic data to produce reliable, structured feeds for training, evaluation, and production.
Define data standards by creating schemas, validation checks, and quality metrics across Replica datasets.
Build curation tooling for dataset filtering, versioning, and annotation support to enable reproducible data pipelines.
Collaborate with ML engineers to understand data needs and optimize delivery for experimentation and deployment.
Required skills include Python, handling large datasets, 3D vision concepts, and familiarity with cloud and MLOps practices.
You will normalize and validate customer and synthetic data to produce reliable, structured feeds for training, evaluation, and production.
Define data standards by creating schemas, validation checks, and quality metrics across Replica datasets.
Build curation tooling for dataset filtering, versioning, and annotation support to enable reproducible data pipelines.
Collaborate with ML engineers to understand data needs and optimize delivery for experimentation and deployment.
Required skills include Python, handling large datasets, 3D vision concepts, and familiarity with cloud and MLOps practices.
See the full job post
Responsibilities, requirements, skills and benefits — create your free account.
Already have an account? Log in
Similar openings
Other roles that could suit you.
Remote workPartial
CityVancouver, Canada