Data Engineer
职位介绍
Design, build, and operate batch data pipelines using Python, Spark (EMR), and AWS Glue to load data from S3, RDS, and external sources into Hive/Athena tables.
Model datasets in the S3/Hive data lake to support analytics, API use cases, Elasticsearch indexes, and ML models.
Implement and run Airflow (MWAA) workflows with dependency management, scheduling, retries, and alerting via Slack.
Build robust data quality checks and validation, with monitoring and quick issue surface.
Optimize jobs for cost and performance through partitioning, file formats, and efficient resource usage.
Collaborate with data scientists, ML engineers, and application engineers to design schemas and pipelines that serve multiple use cases and contribute to tooling and best practices.
Model datasets in the S3/Hive data lake to support analytics, API use cases, Elasticsearch indexes, and ML models.
Implement and run Airflow (MWAA) workflows with dependency management, scheduling, retries, and alerting via Slack.
Build robust data quality checks and validation, with monitoring and quick issue surface.
Optimize jobs for cost and performance through partitioning, file formats, and efficient resource usage.
Collaborate with data scientists, ML engineers, and application engineers to design schemas and pipelines that serve multiple use cases and contribute to tooling and best practices.
查看完整职位
工作职责、任职要求、技能与福利 — 免费创建账号即可查看。
或
已有账户?
登录相似职位
其他可能适合您的职位。