Candidate should have 6+ years of experience in Data Engineering with strong hands-on experience in designing, developing, and maintaining scalable data pipelines and data processing solutions.
The candidate should have strong proficiency in Python and SQL and hands-on experience with ETL/ELT processes, Apache Spark/PySpark, data transformation, data integration, and data warehousing.
Experience with cloud data platforms such as Microsoft Azure or AWS is required. Candidates with hands-on experience in Azure Data Factory, Databricks, Snowflake, AWS Glue, S3, Redshift, or similar cloud data services will be preferred.
The candidate should have experience working with large datasets, batch and/or real-time data processing, data quality, performance optimization, and pipeline monitoring.
Good knowledge of Apache Airflow, Kafka, REST APIs, Git, CI/CD, and Agile methodologies will be an added advantage.
The candidate should possess strong analytical, problem-solving, debugging, communication, and stakeholder-management skills.
Roles & Responsibilities:
- Design, develop, and maintain scalable data pipelines for batch and real-time data processing.
- Develop and optimize ETL/ELT workflows to extract, transform, and load data from multiple sources.
- Write efficient and complex SQL queries for data extraction, transformation, validation, and analysis.
- Develop data processing applications using Python and PySpark/Apache Spark.
- Build and maintain data pipelines using cloud and distributed data technologies.
- Work with cloud platforms such as Microsoft Azure or AWS for data ingestion, storage, processing, and analytics.
- Develop and manage data pipelines using tools such as Azure Data Factory, Databricks, AWS Glue, or equivalent technologies.
- Design and implement scalable data warehouse and data lake solutions.
- Perform data cleansing, transformation, validation, and quality checks to ensure data accuracy and consistency.
- Troubleshoot data pipeline failures, identify root causes, and implement corrective solutions.
- Optimize SQL queries, Spark jobs, and data pipelines for performance and scalability.
- Integrate data from databases, APIs, files, and other structured and unstructured data sources.
- Implement data pipeline monitoring, logging, error handling, and alerting mechanisms.
- Work with Apache Airflow or similar workflow orchestration tools to schedule and manage data pipelines.
- Work with Kafka or other streaming technologies for real-time data ingestion, where required.
- Follow Git, CI/CD, Agile, and SDLC practices for development and deployment.
- Collaborate with Data Scientists, BI Developers, Software Engineers, Architects, and business stakeholders.
- Prepare technical documentation for data pipelines, data models, integrations, and deployment processes.
- Participate in code reviews, testing, production deployment, and ongoing application support.
Preferred Skills:
- Strong experience in Python and SQL.
- Hands-on experience with PySpark / Apache Spark.
- Experience in ETL/ELT and data pipeline development.
- Good knowledge of Azure or AWS cloud platforms.
- Experience with Databricks, Azure Data Factory, or Snowflake.
- Knowledge of data warehousing and data lake architecture.
- Experience with Apache Airflow or similar orchestration tools.
- Knowledge of Kafka and real-time data processing is an advantage.
- Experience with Git and CI/CD practices.
- Good understanding of data quality, validation, and performance optimization.
- Strong debugging, analytical, and problem-solving skills.
- Good communication and ability to work with cross-functional teams.