Data Engineer
¿Es esta oferta para usted?
Crear mi CV Cree su CV y descubra su porcentaje de coincidencia con este puesto — y con todos los demás.
El puesto
Design, build, and operate batch data pipelines that move data from external vendors, internal systems, and public sources into our S3-based data lake and downstream services.
Work across AWS Glue, EMR (Spark), Athena/Hive, and Airflow to ensure data is accurate, well-modeled, and readily accessible for analytics, APIs, and ML workloads.
Own end-to-end data flows from ingestion and transformation to quality checks, monitoring, and performance tuning.
Develop pipelines using Python and Spark (PySpark) to load data into Hive/Athena, model datasets for analytics, APIs, Elasticsearch indexes, and ML models.
Implement robust data quality checks, monitoring, and alerting, including schema validation, freshness/volume checks, and anomaly detection.
Collaborate with data scientists, ML engineers, and application engineers; contribute to internal tooling, documentation, and best practices to standardize data work.
Work across AWS Glue, EMR (Spark), Athena/Hive, and Airflow to ensure data is accurate, well-modeled, and readily accessible for analytics, APIs, and ML workloads.
Own end-to-end data flows from ingestion and transformation to quality checks, monitoring, and performance tuning.
Develop pipelines using Python and Spark (PySpark) to load data into Hive/Athena, model datasets for analytics, APIs, Elasticsearch indexes, and ML models.
Implement robust data quality checks, monitoring, and alerting, including schema validation, freshness/volume checks, and anomaly detection.
Collaborate with data scientists, ML engineers, and application engineers; contribute to internal tooling, documentation, and best practices to standardize data work.
Ver la oferta completa
Funciones, perfil, competencias y ventajas — crea tu cuenta gratis.
o
¿Ya tiene una cuenta?
Iniciar sesiónOfertas similares
Otros puestos que podrían encajar.