Data Engineer
Is deze vacature iets voor u?
Mijn cv maken Maak uw cv en ontdek uw matchpercentage met deze functie — en met alle andere.
De functie
Design, build, and operate batch data pipelines that move data from external vendors, internal systems, and public sources into our S3-based data lake and downstream services.
Work across AWS Glue, EMR (Spark), Athena/Hive, and Airflow to ensure data is accurate, well-modeled, and readily accessible for analytics, APIs, and ML workloads.
Own end-to-end data flows from ingestion and transformation to quality checks, monitoring, and performance tuning.
Develop pipelines using Python and Spark (PySpark) to load data into Hive/Athena, model datasets for analytics, APIs, Elasticsearch indexes, and ML models.
Implement robust data quality checks, monitoring, and alerting, including schema validation, freshness/volume checks, and anomaly detection.
Collaborate with data scientists, ML engineers, and application engineers; contribute to internal tooling, documentation, and best practices to standardize data work.
Work across AWS Glue, EMR (Spark), Athena/Hive, and Airflow to ensure data is accurate, well-modeled, and readily accessible for analytics, APIs, and ML workloads.
Own end-to-end data flows from ingestion and transformation to quality checks, monitoring, and performance tuning.
Develop pipelines using Python and Spark (PySpark) to load data into Hive/Athena, model datasets for analytics, APIs, Elasticsearch indexes, and ML models.
Implement robust data quality checks, monitoring, and alerting, including schema validation, freshness/volume checks, and anomaly detection.
Collaborate with data scientists, ML engineers, and application engineers; contribute to internal tooling, documentation, and best practices to standardize data work.
Bekijk de volledige vacature
Taken, profiel, vaardigheden en voordelen — maak gratis een account aan.
of
Al een account?
InloggenVergelijkbare vacatures
Andere functies die kunnen passen.