Job Responsibilities :
SQL Development :
- Write medium to complex SQL queries to support data extraction, transformation, and analysis.
- Optimize SQL queries for performance and maintainability.
PySpark/Spark Development :
- Develop, test, and deploy data processing applications using PySpark or Spark with Scala.
- Implement and maintain ETL/ELT data pipelines, ensuring efficient data processing and integration.
ETL/ Data Engineering Pipeline Design :
- Design and implement robust ETL pipelines to support data ingestion, transformation, and loading.
- Collaborate with data architects and engineers to ensure seamless data flow and integration across systems.
Cloud Technology Utilization :
- Leverage cloud platforms (AWS, Azure, Google Cloud, etc.) to build scalable and efficient data solutions.
- Work with cloud based data storage solutions (e.g., S3, Google Cloud Storage) and compute environments.
Debugging and Issue Resolution :
- Debug data processing issues and resolve challenges independently as an individual contributor.
- Troubleshoot and optimize existing pipelines to improve reliability and performance.
Continuous Integration and Continuous Deployment (CI/CD) :
- Implement and manage CI/CD pipelines for automated testing, deployment, and monitoring.
- Ensure code quality and best practices through version control (e.g., Git) and CI/CD tools.
Desired Skills :
Data Modeling :
- Design and maintain logical and physical data models to support business requirements.
- Ensure data models are optimized for performance, scalability, and ease of use.
Airflow :
- Use Apache Airflow to orchestrate and schedule complex data workflows.
- Develop and maintain DAGs for task automation and monitoring.
Kafka :
- Utilize Apache Kafka for building real-time data streaming pipelines.
- Integrate Kafka with other systems for event -driven data processing.
Big Data Technologies :
- Experience with big data processing frameworks such as Hive, EMR, and Databricks.
- Familiarity with data storage solutions like Snowflake for analytics and reporting.
Data Lakes and Data Warehousing :
- Work with data lake solutions and data warehouses for storing and managing large datasets.
- Implement data lake architectures for scalable data ingestion and processing.
Technical Knowledge and Skills Required :
- Strong hands-on experience with PySpark or Spark with Scala for big data processing.
- Proficient in writing complex SQL queries and optimizing them for performance.
- Understanding of ETL/ELT design principles and best practices.
- Basic to intermediate knowledge of CI/CD processes and tools.
- Familiarity with cloud platforms (AWS, Azure, Google Cloud) and their data-related services.
- Ability to work independently and resolve issues effectively.
Soft Skills Required :
- Strong analytical and problem-solving skills.
- Excellent communication and collaboration abilities.
- Ability to manage multiple tasks and projects efficiently.
- A proactive attitude towards learning new technologies and improving existing skills.
Work Experience :
- 5+ years of relevant experience in data engineering, focusing on big data processing and cloud technologies.
- Experience working with data engineering tools and frameworks like Airflow, Kafka, Hive, EMR, Databricks, and Snowflake.
Education :
- Bachelor's degree in Computer Science, Information Technology, Data Science, or a related field.
- A Master's degree is a plus.
Location: Anywhere in /Multiple Locations
- Delhi / NCR,Bangalore/Bengaluru,Hyderabad/Secunderabad,Chennai,Pune,Kolkata,Ahmedabad,Mumbai