Data Engineer - ETL/PySpark

Forward Eye Technologies

Pune District

On-site

INR 1,200,000 - 2,400,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Forward Eye Technologies is seeking a data engineer specializing in big data technologies to design and implement robust data pipelines. The role focuses on SQL development, PySpark/Spark, and ETL/ELT workflows across cloud environments.

Candidates should have strong experience with Airflow, Kafka, Hive, EMR, Databricks, and Snowflake, and be capable of delivering scalable data solutions while collaborating with architects and engineers.

Qualifications

  • Strong hands-on experience with PySpark or Spark with Scala for big data processing.
  • Proficient in writing complex SQL queries and optimizing them for performance.
  • Understanding of ETL/ELT design principles and best practices.
  • Basic to intermediate knowledge of CI/CD processes and tools.
  • Familiarity with cloud platforms (AWS, Azure, Google Cloud) and their data-related services.
  • Ability to work independently and resolve issues effectively.

Responsibilities

  • Write medium to complex SQL queries for data extraction and analysis.
  • Optimize SQL queries for performance and maintainability.
  • Develop, test, and deploy PySpark/Spark data processing applications.
  • Implement and maintain ETL/ELT data pipelines for efficient processing.
  • Design robust ETL pipelines for ingestion, transformation and loading.
  • Collaborate with data architects to ensure seamless data flow.
  • Leverage cloud platforms (AWS, Azure, Google Cloud) for scalable solutions.
  • Work with cloud storage (S3, Google Cloud Storage) and compute environments.
  • Debug data processing issues and resolve them independently.
  • Troubleshoot and optimize existing pipelines for reliability.
  • Implement and manage CI/CD pipelines for testing and deployment.
  • Maintain code quality with Git and CI/CD best practices.
  • Use Apache Airflow to orchestrate workflows and DAGs.
  • Develop DAGs for task automation and monitoring.
  • Utilize Apache Kafka for real-time data streams.
  • Integrate Kafka with other systems for event-driven processing.
  • Experience with Hive, EMR, Databricks for big data.
  • Familiarity with Snowflake for analytics and reporting.
  • Work with data lakes and data warehouses for large datasets.

Skills

PySpark
Spark with Scala
SQL
ETL/ELT
CI/CD
AWS
Azure
Google Cloud
Airflow
Kafka
Hive
EMR
Databricks
Snowflake
Data Modeling
Data Lakes
Data Warehousing

Education

Bachelor's degree in CS/IT/Data Science
Master's degree (plus)

Tools

Airflow
Kafka
Hive
EMR
Databricks
Snowflake
S3

Job description

Job Responsibilities :
SQL Development :
  • Write medium to complex SQL queries to support data extraction, transformation, and analysis.
  • Optimize SQL queries for performance and maintainability.
PySpark/Spark Development :
  • Develop, test, and deploy data processing applications using PySpark or Spark with Scala.
  • Implement and maintain ETL/ELT data pipelines, ensuring efficient data processing and integration.
ETL/ Data Engineering Pipeline Design :
  • Design and implement robust ETL pipelines to support data ingestion, transformation, and loading.
  • Collaborate with data architects and engineers to ensure seamless data flow and integration across systems.
Cloud Technology Utilization :
  • Leverage cloud platforms (AWS, Azure, Google Cloud, etc.) to build scalable and efficient data solutions.
  • Work with cloud based data storage solutions (e.g., S3, Google Cloud Storage) and compute environments.
Debugging and Issue Resolution :
  • Debug data processing issues and resolve challenges independently as an individual contributor.
  • Troubleshoot and optimize existing pipelines to improve reliability and performance.
Continuous Integration and Continuous Deployment (CI/CD) :
  • Implement and manage CI/CD pipelines for automated testing, deployment, and monitoring.
  • Ensure code quality and best practices through version control (e.g., Git) and CI/CD tools.
Desired Skills :
Data Modeling :
  • Design and maintain logical and physical data models to support business requirements.
  • Ensure data models are optimized for performance, scalability, and ease of use.
Airflow :
  • Use Apache Airflow to orchestrate and schedule complex data workflows.
  • Develop and maintain DAGs for task automation and monitoring.
Kafka :
  • Utilize Apache Kafka for building real-time data streaming pipelines.
  • Integrate Kafka with other systems for event -driven data processing.
Big Data Technologies :
  • Experience with big data processing frameworks such as Hive, EMR, and Databricks.
  • Familiarity with data storage solutions like Snowflake for analytics and reporting.
Data Lakes and Data Warehousing :
  • Work with data lake solutions and data warehouses for storing and managing large datasets.
  • Implement data lake architectures for scalable data ingestion and processing.
Technical Knowledge and Skills Required :
  • Strong hands-on experience with PySpark or Spark with Scala for big data processing.
  • Proficient in writing complex SQL queries and optimizing them for performance.
  • Understanding of ETL/ELT design principles and best practices.
  • Basic to intermediate knowledge of CI/CD processes and tools.
  • Familiarity with cloud platforms (AWS, Azure, Google Cloud) and their data-related services.
  • Ability to work independently and resolve issues effectively.
Soft Skills Required :
  • Strong analytical and problem-solving skills.
  • Excellent communication and collaboration abilities.
  • Ability to manage multiple tasks and projects efficiently.
  • A proactive attitude towards learning new technologies and improving existing skills.
Work Experience :
  • 5+ years of relevant experience in data engineering, focusing on big data processing and cloud technologies.
  • Experience working with data engineering tools and frameworks like Airflow, Kafka, Hive, EMR, Databricks, and Snowflake.
Education :
  • Bachelor's degree in Computer Science, Information Technology, Data Science, or a related field.
  • A Master's degree is a plus.

Location: Anywhere in /Multiple Locations

  • Delhi / NCR,Bangalore/Bengaluru,Hyderabad/Secunderabad,Chennai,Pune,Kolkata,Ahmedabad,Mumbai
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer - ETL/PySpark
Data Engineer - ETL/PySpark

Forward Eye Technologies • Dadri

On-site
INR 1,800,000 - 2,400,000
Data Engineer - SQL/PySpark
Data Engineer - SQL/PySpark

Forward Eye Technologies • Dadri

Hybrid
INR 1,400,000 - 2,000,000
Data Engineer - ETL
Data Engineer - ETL

Forward Eye Technologies • Pune District

On-site
INR 2,000,000 - 3,200,000
Data Engineer - SQL/PySpark
Data Engineer - SQL/PySpark

Forward Eye Technologies • Pune District

On-site
INR 1,500,000 - 2,100,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Maharashtra

On-site
INR 800,000 - 1,200,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Mumbai

On-site
INR 1,000,000 - 1,500,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Kolkata District

On-site
INR 1,000,000 - 1,500,000
Data Engineer
Data Engineer

Qcentrio • Hyderabad

On-site
INR 1,800,000 - 3,200,000
Data Engineer
Data Engineer

HGS • Hyderabad

On-site
INR 1,200,000 - 2,100,000
Data Engineer
Data Engineer

Advance Career Solutions • Pune District, Chennai District, Bengaluru

Hybrid
INR 1,200,000 - 2,800,000