Data Engineer - SQL/PySpark

Forward Eye Technologies

Mumbai, Bengaluru, New Delhi

On-site

INR 1,800,000 - 3,200,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Forward Eye Technologies in Mumbai is seeking a Data Engineer with strong SQL, PySpark, and cloud experience to design and optimize data pipelines for performance, scalability, and reliability. You will work closely with data engineers, data scientists, and stakeholders to develop, maintain, and optimize data infrastructure.

You will design ETL processes, implement Spark-based transformations, and ensure data pipelines meet analytics and BI requirements.

Qualifications

  • 5+ years of data engineering or related field.
  • Strong SQL and PySpark skills for large-scale data.
  • Experience with ETL and data pipeline design.
  • Basic to intermediate knowledge of CI/CD processes.
  • Experience with cloud platforms (AWS/Azure/GCP).

Responsibilities

  • Write medium-complexity SQL queries to extract, transform, and load data efficiently.
  • Optimize SQL queries for performance and scalability.
  • Collaborate with team members to understand data requirements and translate them into actionable SQL scripts.
  • Develop and maintain data processing scripts using PySpark or Spark with Scala.
  • Implement data transformation and ETL processes using Spark to handle large-scale data.
  • Optimize Spark jobs for performance and troubleshoot processing issues.
  • Design and implement ETL/data pipelines to support BI and analytics needs.
  • Ensure data pipelines are scalable, reliable, and maintainable.
  • Collaborate with the team to understand data flow requirements and implement solutions.
  • Work with cloud platforms to build and manage data solutions and ensure security and compliance.
  • Utilize cloud services for data integration, orchestration, and processing.
  • Identify, diagnose, and resolve data-related issues; monitor pipelines for reliability.

Skills

SQL proficiency
PySpark / Spark
ETL design
CI/CD basics
Cloud technologies
Data modeling
Problem-solving
Communication
Data pipelines

Education

Bachelor's degree in CS/IT or related

Tools

Apache Airflow
Kafka
AWS EMR / Databricks
Snowflake
Hive

Job description

We are seeking a skilled Data Engineer with strong experience in SQL, PySpark, and cloud technologies to join our dynamic team. The ideal candidate will have a solid background in designing and implementing data engineering pipelines, with a focus on performance, scalability, and reliability. You will work closely with other data engineers, data scientists, and stakeholders to develop, maintain, and optimize data infrastructure.

Key Responsibilities :
SQL Development :
  • Write medium-complexity SQL queries to extract, transform, and load data efficiently.
  • Optimize SQL queries for performance and scalability.
  • Collaborate with team members to understand data requirements and translate them into actionable SQL scripts.
PySpark/Spark Development :
  • Develop and maintain data processing scripts using PySpark or Spark with Scala.
  • Implement data transformation and ETL processes using Spark to handle large-scale data.
  • Optimize Spark jobs for performance and troubleshoot any issues that arise during processing.
Data Engineering Pipeline Design :
  • Design and implement ETL/data engineering pipelines to support business intelligence and analytics needs.
  • Ensure data pipelines are scalable, reliable, and maintainable.
  • Collaborate with the team to understand data flow requirements and implement appropriate solutions.
Cloud Technology Implementation :
  • Work with various cloud platforms (AWS, Azure, GCP) to build and manage data solutions.
  • Implement cloud-based data storage and processing solutions, ensuring security and compliance.
  • Utilize cloud services for data integration, orchestration, and processing.
Issue Resolution and Debugging :
  • Identify, diagnose, and resolve data-related issues and bugs independently.
  • Perform root cause analysis and implement fixes for any data discrepancies or processing failures.
  • Monitor data pipelines and systems to ensure continuous availability and reliability.
CI/CD Integration :
  • Implement and maintain basic to intermediate level CI/CD processes for data engineering projects.
  • Collaborate with DevOps teams to ensure smooth deployment and integration of data solutions.
  • Automate testing and deployment processes to improve development efficiency.
Desired Skills :
Data Modeling :
  • Experience in designing data models to support analytical and business requirements.
Airflow :
  • Experience with Apache Airflow for workflow scheduling and orchestration.
Kafka :
  • Knowledge of Kafka for real-time data streaming and integration.
EMR/Databricks :
  • Experience with AWS EMR or Databricks for big data processing and analytics.
Snowflake :
  • Familiarity with Snowflake for cloud-based data warehousing solutions.
Hive :
  • Understanding of Hive for querying and managing large datasets.
Qualifications :
  • Experience : 5+ years of experience in data engineering or related fields.
Technical Skills :
  • Strong proficiency in SQL and writing medium-complexity queries.
  • Hands-on experience with PySpark or Spark with Scala.
  • Understanding of ETL processes and data engineering pipeline design.
  • Basic to intermediate knowledge of CI/CD processes.
  • Experience with cloud technologies (AWS, Azure, GCP).
Problem-Solving :
  • Ability to troubleshoot and resolve data issues independently.
Communication :
  • Strong verbal and written communication skills, with the ability to collaborate effectively with technical and non-technical stakeholders.
Education :
  • Bachelor's degree in Computer Science, Information Technology, or related field.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer - SQL/PySpark
Data Engineer - SQL/PySpark

Forward Eye Technologies • Pune District

On-site
INR 1,500,000 - 2,100,000
Data Engineer - SQL/PySpark
Data Engineer - SQL/PySpark

Forward Eye Technologies • Dadri

Hybrid
INR 1,400,000 - 2,000,000
Data Engineer - ETL/PySpark
Data Engineer - ETL/PySpark

Forward Eye Technologies • Dadri

On-site
INR 1,800,000 - 2,400,000
Data Engineer - ETL/PySpark
Data Engineer - ETL/PySpark

Forward Eye Technologies • Pune District

On-site
INR 1,200,000 - 2,400,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Mumbai

On-site
INR 1,000,000 - 1,500,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Kolkata District

On-site
INR 1,000,000 - 1,500,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Maharashtra

On-site
INR 800,000 - 1,200,000
Data Engineer
Data Engineer

Advance Career Solutions • Pune District, Chennai District, Bengaluru

Hybrid
INR 1,200,000 - 2,800,000
Data Engineer
Data Engineer

Synergy Computer Solutions • Hyderabad

On-site
INR 1,200,000 - 2,000,000
Data Engineer
Data Engineer

Digitide • Bengaluru

On-site
INR 1,200,000 - 1,800,000