An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Jobtailor is seeking a seasoned Data Engineer to design, build, and optimize scalable data pipelines on AWS, leveraging Glue, S3, Athena, MWAA, and Step Functions. You will contribute to a robust data platform by implementing Medallion Architecture and Iceberg tables, ensuring data quality, lineage, and governance.
With 5-8 years of data engineering experience, strong PySpark and SQL skills, and IaC with Terraform, you will collaborate with architects, analysts, and business teams to modernize
Design, develop, and maintain scalable data pipelines using AWS Glue (PySpark), Amazon S3, AWS Step Functions, and Athena
Build robust ETL/ELT solutions for batch and near real-time data processing
Develop reusable PySpark frameworks and data transformation components
Work with architects and business stakeholders to implement data platform requirements
Participate in migration and modernization initiatives from on-premises data platforms to AWS cloud environments
Implement and support enterprise data lakes using Medallion Architecture
Develop and maintain Apache Iceberg tables for storage, schema evolution, and incremental processing
Ensure data quality, lineage, reconciliation, and auditability across the data platform
Contribute to data modeling and optimization for analytical workloads
Develop and maintain Apache Airflow (MWAA) workflows and DAGs
Automate data movement, validation, monitoring, and notification processes
Implement retry mechanisms, dependency management, and failure handling within workflows
Integrate Airflow with AWS Glue, S3, Athena, and downstream applications
Implement AWS security best practices, including IAM roles, KMS encryption, and secrets management
Support infrastructure provisioning and deployment using Terraform
Collaborate with DevOps, infrastructure, and security teams to maintain secure and reliable cloud environments
Assist in configuring VPC endpoints, networking connectivity, and service integrations
Monitor and troubleshoot production data pipelines and workflows
Perform performance tuning of AWS Glue jobs and Spark workloads
Implement logging, monitoring, and alerting using Amazon CloudWatch
Participate in incident resolution, root cause analysis, and continuous improvement initiatives
Ensure adherence to enterprise operational and governance standards
Collaborate with data architects, analysts, application teams, and business users
Participate in code reviews, design discussions, and technical documentation
Mentor junior engineers and share AWS data engineering best practices
Stay current with emerging AWS services and modern data engineering trends
Demonstrates expertise in designing and maintaining scalable data pipelines using AWS technologies, including AWS Glue and S3, while ensuring data quality and compliance with enterprise standards. Proficient in developing ETL/ELT solutions and collaborating with cross-functional teams to optimize data processing and architecture.