Data Engineer
Questa offerta fa per te?
Crea il mio CV Crea il tuo CV e scopri la tua percentuale di corrispondenza con questa posizione — e con tutte le altre.
La posizione
Design, build, and operate batch data pipelines that move data from external vendors, internal systems, and public sources into our S3-based data lake and downstream services.
Work across AWS Glue, EMR (Spark), Athena/Hive, and Airflow to ensure data is accurate, well-modeled, and readily accessible for analytics, APIs, and ML workloads.
Own end-to-end data flows from ingestion and transformation to quality checks, monitoring, and performance tuning.
Develop pipelines using Python and Spark (PySpark) to load data into Hive/Athena, model datasets for analytics, APIs, Elasticsearch indexes, and ML models.
Implement robust data quality checks, monitoring, and alerting, including schema validation, freshness/volume checks, and anomaly detection.
Collaborate with data scientists, ML engineers, and application engineers; contribute to internal tooling, documentation, and best practices to standardize data work.
Work across AWS Glue, EMR (Spark), Athena/Hive, and Airflow to ensure data is accurate, well-modeled, and readily accessible for analytics, APIs, and ML workloads.
Own end-to-end data flows from ingestion and transformation to quality checks, monitoring, and performance tuning.
Develop pipelines using Python and Spark (PySpark) to load data into Hive/Athena, model datasets for analytics, APIs, Elasticsearch indexes, and ML models.
Implement robust data quality checks, monitoring, and alerting, including schema validation, freshness/volume checks, and anomaly detection.
Collaborate with data scientists, ML engineers, and application engineers; contribute to internal tooling, documentation, and best practices to standardize data work.
Vedi l'annuncio completo
Mansioni, profilo, competenze e vantaggi — crea il tuo account gratuito.
o
Hai già un account?
AccediOfferte simili
Altre posizioni che potrebbero interessarti.