Job Summary
We are seeking a motivated and skilled Data Engineer with 3–4 years of experience in building scalable data pipelines, ETL processes, and cloud-based data solutions. The ideal candidate should have hands-on expertise in Databricks, Python, SQL, and Apache Airflow, and experience supporting data migration initiatives from on-premises platforms to cloud.
The role involves designing, developing, optimizing, and maintaining data pipelines that support analytics, reporting, and business intelligence requirements.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT pipelines using Python, PySpark, Databricks, and SQL.
- Build reusable data ingestion, transformation, and data quality frameworks.
- Develop batch and near real-time data processing solutions.
- Perform data cleansing, transformation, validation, and enrichment activities.
- Databricks Development: Build and optimize Spark/PySpark applications in Azure Databricks.
- Work with Delta Lake, data partitioning, caching, and performance tuning techniques.
- Create and maintain notebooks, jobs, workflows, and reusable libraries.
- Implement medallion architecture (Bronze, Silver, Gold layers).
- Workflow Orchestration: Develop and maintain Apache Airflow DAGs for scheduling and orchestration.
- Monitor pipeline execution and troubleshoot failures; implement alerting, logging, and retry mechanisms.
- Participate in migration of on-premise data assets and ETL jobs to cloud platforms.
- Analyze legacy ETL workflows and redesign them using cloud-native services.
- Support data validation, reconciliation, and migration testing activities.
- Assist in cutover and production deployment activities.
- SQL & Data Modeling: Develop complex SQL queries, stored procedures, and performance optimization.
- Design dimensional and normalized data models.
- Support data warehousing and reporting requirements.
- Collaborate with business analysts, architects, and stakeholders to understand data requirements.
- Participate in Agile ceremonies, sprint planning, and code reviews.
- Ensure adherence to security, governance, and data quality standards.
Required Technical Skills
- Python / PySpark — 3+ Years
- SQL — 3+ Years
- Apache Airflow — 2+ Years
- ETL / ELT Development — 3+ Years
- Cloud Migration Projects — 1+ Years
Required Qualifications
- Bachelor\'s degree in Computer Science, Information Technology, Engineering, or related field.
- 3–4 years of experience in Data Engineering.
- Strong understanding of Data Warehousing concepts.
- Experience working with large-scale structured and semi-structured datasets.
- Good knowledge of cloud-based analytics platforms.
Preferred Skills
- Data Quality and Data Governance frameworks
- Exposure to CI/CD pipelines
- Strong analytical and problem-solving skills
- Excellent communication and stakeholder management
- Ability to work independently and within Agile teams
- Strong debugging and performance tuning capabilities
- Focus on quality, automation, and continuous improvement
Nice to Have
- Experience with enterprise-scale cloud migration programs
- Exposure to healthcare, finance, telecom, or retail data domains