The Data Engineer plays a key role in designing and implementing data integration solutions that align with business objectives.
This role involves developing data pipelines, performing data manipulation and mining, and managing large-scale data processing using Python, PySpark, and AWS services.
The engineer collaborates with cross-functional teams to deliver scalable and reliable data solutions that support enterprise-wide initiatives.
Key Responsibilities
- Provide scoping, estimation, planning, design, development, and support services for data projects.
- Create detailed technical design documentation.
- Design, configure, deploy, and maintain custom ETL infrastructure.
- Develop and maintain ETL pipelines for data processing, transformation, and loading into large data domains.
- Collaborate with business analysts to understand and prioritize user requirements.
- Develop, test, and implement application code following software development lifecycle standards.
- Conduct quality analysis, defect tracking, and classification.
- Monitor project progress and resolve or escalation issues as needed.
- Coordinate with vendors and contractors for specific systems or projects.
- Stay updated on industry best practices and emerging technologies.
Required Qualifications
- Hands‑on expertise in Python and PySpark.
- Experience with AWS services such as Glue, Lambda, MSK (Kafka), S3, Step Functions, RDS, EKS.
- Proficiency in relational databases including Postgres, SQL Server, Oracle, Sybase.
- Strong SQL programming skills including performance tuning, stored procedures, views, and triggers.
- Experience in UNIX scripting and Oracle SQL/PL‑SQL.
- Familiarity with REST APIs and workload automation tools like Control‑M or Autosys.
- Knowledge of CI/CD tools such as Bitbucket, GitHub, Jenkins.
- Experience with Agile/SCRUM methodologies.
Preferred Qualifications
- Experience with other ETL tools such as DataStage, Informatica, or Pentaho.
- Knowledge of Master Data Management (MDM), Data Warehousing, and Data Analytics.