Data Engineer
Is this job for you?
Build my CV Build your CV and see how well you match this role — and every other one.
The role
Architect and build the data infrastructure powering crawling, embedding training, and real-time search.
Design scalable lakehouse architectures (Delta Lake, Iceberg, Hudi) and large distributed pipelines.
Develop streaming systems (Kafka, Flink) and production-grade data processing with Ray, Spark, or ClickHouse.
Prioritize reliability and uptime, delivering systems that rarely page at 3am.
Leverage GPU-accelerated processing (RAPIDS, cuDF) and vector-native storage formats (Lance) where applicable.
Own the data layer for embedding training and indexing workflows at multi-petabyte scales.
Design scalable lakehouse architectures (Delta Lake, Iceberg, Hudi) and large distributed pipelines.
Develop streaming systems (Kafka, Flink) and production-grade data processing with Ray, Spark, or ClickHouse.
Prioritize reliability and uptime, delivering systems that rarely page at 3am.
Leverage GPU-accelerated processing (RAPIDS, cuDF) and vector-native storage formats (Lance) where applicable.
Own the data layer for embedding training and indexing workflows at multi-petabyte scales.
See the full job post
Responsibilities, requirements, skills and benefits — create your free account.
or
Already have an account?
Log inSimilar openings
Other roles that could suit you.