Data Engineer
Is this job for you?
Build my CV Build your CV and see how well you match this role — and every other one.
The role
Architect and build the data infrastructure powering crawling billions of pages, embedding model training, and real-time search.
Design lakehouse architectures using Delta Lake, Iceberg, or Hudi and decide when to use each.
Build and operate large-scale distributed data processing pipelines and streaming systems (Kafka, Flink).
Work with Ray, Spark, or ClickHouse in production to enable fast analytics and indexing pipelines.
Maintain a relentless focus on reliability to ensure systems rarely require on-call support.
Scale data infrastructure to hundreds of petabytes and own the end-to-end data layer for embedding training and real-time indexing.
Design lakehouse architectures using Delta Lake, Iceberg, or Hudi and decide when to use each.
Build and operate large-scale distributed data processing pipelines and streaming systems (Kafka, Flink).
Work with Ray, Spark, or ClickHouse in production to enable fast analytics and indexing pipelines.
Maintain a relentless focus on reliability to ensure systems rarely require on-call support.
Scale data infrastructure to hundreds of petabytes and own the end-to-end data layer for embedding training and real-time indexing.
See the full job post
Responsibilities, requirements, skills and benefits — create your free account.
or
Already have an account?
Log inSimilar openings
Other roles that could suit you.