Data Engineer - ETL

Forward Eye Technologies

Pune District

On-site

INR 2,000,000 - 3,200,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Forward Eye Technologies in Pune is seeking a skilled Data Engineer to join our data team. You will design, build, and optimize large-scale data pipelines on AWS using PySpark and SQL, collaborating with data science and engineering peers to support analytics initiatives.

The role emphasizes cloud data infrastructure, real-time and batch processing, and data quality. You will work with S3, Glue, Redshift, and EMR, implementing efficient queries, governance, and CI/CD practices to ensure robust

Qualifications

  • 5+ years of data engineering experience with PySpark, AWS, and SQL.
  • Strong PySpark and ETL/ELT development experience.
  • Hands-on with AWS data services (S3, Glue, Redshift, EMR, Lambda).
  • Expertise in writing optimized SQL queries for large datasets.
  • Experience with data quality, governance, and CI/CD for data workflows.

Responsibilities

  • Design, build, and optimize ETL pipelines using PySpark on AWS.
  • Collaborate with data engineering and data science teams to translate requirements into scalable data solutions.
  • Develop and manage AWS data services (S3, Glue, Redshift, EMR, Lambda) for processing and storage.
  • Write and optimize complex SQL queries for data extraction and transformation on RDS/Redshift.
  • Implement data quality checks, governance, and security using AWS-native tools.
  • Automate repetitive tasks and maintain CI/CD pipelines for data workflows.
  • Troubleshoot data pipelines and support analytics initiatives with accessible data formats.

Skills

PySpark
AWS
SQL
Airflow
CI/CD
Git
Data pipelines
Data modelling

Tools

Airflow
Git

Job description

We are looking for a skilled Data Engineer with expertise in PySpark, AWS, and SQL to support data processing and analytical initiatives. This role involves working closely with data engineering and data science teams to build, maintain, and optimize large-scale data pipelines and integrations on AWS. The ideal candidate will be proficient in ETL processes using PySpark and SQL, with a deep understanding of cloud data infrastructure, specifically within the AWS ecosystem.

Key Responsibilities :
Data Pipeline Development :
  • - Design, build, and optimize ETL pipelines using PySpark for data ingestion, transformation, and storage on AWS.
  • - Collaborate with stakeholders to understand data requirements, translating them into scalable data solutions.
Cloud Infrastructure Management :
  • - Develop and manage AWS services such as S3, Glue, Lambda, EMR, Redshift, and RDS for data processing and storage.
  • - Implement data workflows that handle both batch and real-time processing needs, ensuring low-latency and efficient data access.
Database Management :
  • - Write and optimize complex SQL queries for data extraction and transformation from AWS RDS, Redshift, and other SQL-based databases.
  • - Leverage indexing, partitioning, and caching techniques to improve query performance on large datasets.
Data Quality and Governance :
  • - Ensure data accuracy, completeness, and consistency throughout the data lifecycle.
  • - Implement best practices for data quality, governance, and security using AWS-native tools and third-party solutions.
Automation and Optimization :
  • - Automate repetitive tasks and optimize workflows, ensuring efficient and resilient data processing.
  • - Use CI/CD pipelines and version control (e.g., Git) to deploy and manage data workflows.
Troubleshooting and Support :
  • - Identify and resolve issues within the data pipeline, providing solutions for data access and query performance.
  • - Support the team in data analytics and data science initiatives by preparing data in an accessible and structured format.
Required Skills and Experience :
  • - 5+ years of data engineering experience with a focus on PySpark, AWS, and SQL.
  • - Strong proficiency in PySpark for ETL and data transformation tasks.
  • - Hands-on experience with AWS data services (S3, Glue, Redshift, EMR, Lambda).
  • - Expertise in SQL and relational database management, with experience optimizing complex queries.
  • - Experience in data pipeline orchestration and monitoring (Airflow or similar).
  • - Knowledge of data partitioning, indexing, and caching techniques to enhance performance.
  • - Familiarity with CI/CD principles and tools, including version control systems like Git.
Nice to Have :
  • - Experience with AWS Redshift Spectrum and Athena for querying data on S3.
  • - Knowledge of streaming data processes and tools like Kinesis or Kafka.
  • - Background in big data technologies, such as Hadoop or Apache Spark.
  • - Experience with data lake architectures and building data marts for analytics.

Location : - Anywhere india ,Multiple Locations

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer - SQL/PySpark
Data Engineer - SQL/PySpark

Forward Eye Technologies • Dadri

Hybrid
INR 1,400,000 - 2,000,000
Data Engineer - ETL/PySpark
Data Engineer - ETL/PySpark

Forward Eye Technologies • Pune District

On-site
INR 1,200,000 - 2,400,000
Senior Data Engineer
Senior Data Engineer

AagatiServe Pvt Ltd • Delhi

On-site
INR 1,800,000 - 2,400,000
Data Engineer — PySpark + AWS Glue
Data Engineer — PySpark + AWS Glue

Yadimen Consulting Limited • Chennai District, Pune District, Bengaluru

On-site
INR 1,000,000 - 1,800,000
Data/ETL Engineer
Data/ETL Engineer

Ascendion Engineering • Bengaluru

On-site
INR 800,000 - 1,200,000
Data Engineer - ETL/PySpark
Data Engineer - ETL/PySpark

Forward Eye Technologies • Dadri

On-site
INR 1,800,000 - 2,400,000
AWS + Pyspark Data Engineer ( Gurugram)
AWS + Pyspark Data Engineer ( Gurugram)

PwC India • Gurugram District

Hybrid
INR 1,400,000 - 2,100,000
Data Engineer
Data Engineer

Intact Green Services (india) • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Industry-standard compensation
Data Engineer
Data Engineer

Bct Consulting • Bengaluru

On-site
INR 900,000 - 1,500,000
Data Engineer
Data Engineer

Accenture in India • Maharashtra

On-site
INR 1,200,000 - 2,200,000