Develop and Maintain Data Integration Solutions:
- Design and implement data integration workflows using AWS Glue/EMR,AWS MWAA(Airflow), Lambda, Redshift
- Demonstrate proficiency in Pyspark, Apache Spark and Python for data processing large datasets
- Ensure data is accurately and efficiently extracted, transformed, and loaded into target systems.
Ensure Data Quality and Integrity:
- Validate and cleanse data to maintain high data quality.
- Ensure data quality and integrity by implementing monitoring, validation, and error handling mechanisms within data pipeline.
Optimize Data Integration Processes:
- Enhance the performance, optimization of data workflows to meet SLAs, scalability of data integration processes and cost-efficiency on AWS cloud infrastructure.
- Identify and resolve performance bottlenecks, fine-tuning queries, and optimizing data processing to enhance Redshift's performance
- Regularly review and refine integration processes to improve efficiency.
- Support Business Intelligence and Analytics:
- Translate business requirements to technical specifications and coded data pipelines
- Ensure timely availability of integrated data for business intelligence and analytics.
- Collaborate with data analysts and business stakeholders to meet their data requirements.
Maintain Documentation and Compliance:
- Document all data integration processes, workflows, and technical & system specifications.
- Ensure compliance with data governance policies, industry standards, and regulatory requirements.
WHAT WILL THIS PERSON BE WORKING ON
- 5+ years of experience in data engineering, database design, ETL processes, and data warehousing.
- 3+ years of experience with AWS tools and technologies (S3, EMR, Glue, Athena, RedShift, RDS, Spectrum and Airflow)
- 2+ years of experience with CI/CD tools.
- Strong knowledge of data storage and processing technologies, including databases and data lakes based distributed computing frameworks (e.g., Hadoop, Spark).
- 3+ in programming languages such as Python, Java, or Scala.
- Nice to have Informatica Cloud tool experience (IDMC)
- Nice to have Agentic AI experience with Amazon Kiro
Primary Skills: AWS EMR, GLUE, AIRFLOW, ICEBERG, REDSHIFT, RDS, IDMC, AMAZON KIRO, and CI/CD
Candidate should design and Develop Data Pipelines and support them.
Ensure Data Quality and Integrity:
- Validate and cleanse data to maintain high data quality.
- Ensure data quality and integrity by implementing monitoring, validation, and error handling mechanisms within data pipelines
Optimize Data Integration Processes:
- Enhance the performance and optimization of data workflows to meet SLAs, scalability of data integration processes, and cost-efficiency on AWS cloud infrastructure.
- Identify and resolve performance bottlenecks, fine-tuning queries, and optimizing data processing to enhance Redshift's performance
- Regularly review and refine integration processes to improve efficiency.
Support Business Intelligence and Analytics:
- Translate business requirements to technical specifications and coded data pipelines
- Ensure timely availability of integrated data for business intelligence and analytics.
Collaborate with data analysts and business stakeholders to meet their data requirements