AWS + Pyspark Data Engineer ( Gurugram)

PwC India

Gurugram District

On-site

INR 1,400,000 - 2,100,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

PwC India is seeking a Senior Data Engineer in Gurugram to design and optimize large-scale data pipelines and data warehouses. You will leverage Python, PySpark, and AWS services to manage high-volume data workloads and build modern cloud analytics platforms.

You will collaborate with Data Scientists, Analysts, and DevOps to ensure data quality, governance, and scalable architectures, while continuously tuning performance and enabling CI/CD for data projects.

Qualifications

  • 4–8 years of hands-on data engineering/big data experience.
  • Expert in Python and PySpark.
  • Strong SQL programming, tuning, and database design principles.
  • Extensive experience with AWS cloud services (S3, Glue, EMR, Athena, Redshift).
  • Familiarity with Apache Spark ecosystem and distributed computing concepts.
  • Experience with orchestration tools (Airflow), containerization, and Git.
  • Bachelor’s or Master’s degree in CS/IT/Engineering or related field.

Responsibilities

  • Design, build, and optimize large-scale data pipelines and data warehouses.
  • Develop cloud analytics platforms and ensure scalable data ingestion and processing on AWS.
  • Write complex SQL queries and tune performance for low latency.
  • Collaborate with data scientists, analysts, and DevOps to align data platforms with business needs.
  • Troubleshoot production performance bottlenecks in distributed data jobs.
  • Uphold data quality, security, governance, and CI/CD standards across deliverables.

Skills

Python
PySpark
SQL
AWS
Spark
Airflow
Git

Education

Bachelor's or Master's in Computer Science/IT/Engineering

Tools

Airflow
Docker

Job description

Job Overview
  • Role: Senior Data Engineer
  • Location: Gurugram, Haryana
  • Experience Required: 4 to 8 Years
  • Primary Tech Stack: AWS, PySpark, Advanced SQL
Role Overview

We are looking for a dynamic and results-driven Data Engineer with strong expertise in AWS, PySpark, and SQL to join our growing technology team in Gurugram. In this role, you will design, build, and optimize large-scale data pipelines, data warehouses, and modern cloud analytics platforms to handle high-volume data workloads.

Key Responsibilities
  • Pipeline Development: Design, develop, test, and maintain robust ETL/ELT data pipelines using PySpark and cloud-native services.
  • Cloud & Big Data Management: Build, monitor, and optimize scalable data ingestion, transformation, and processing workflows on AWS (e.g., S3, Glue, EMR, Athena, Redshift, Lambda).
  • SQL Optimization: Write complex SQL queries, perform query tuning, and manage database operations to ensure high performance and low latency.
  • Data Modeling & Architecture: Collaborate with cross-functional teams to build data models, schema designs, and data marts supporting analytical reporting.
  • Performance Tuning: Troubleshoot and resolve production performance bottlenecks in distributed data processing jobs.
  • Collaboration: Work closely with Data Scientists, Business Analysts, and DevOps teams to align data platform infrastructure with business requirements.
  • Best Practices: Ensure data quality, security, governance, and CI/CD automation standards are implemented across all deliverables.
Required Qualifications & Skills
  • Experience: 4 to 8 years of hands-on experience in Data Engineering, Big Data, or Business Intelligence roles.
  • Programming Languages: Expert-level proficiency in Python and PySpark.
  • Cloud Ecosystem: Strong production experience working with AWS cloud services (S3, Glue, EMR, Athena, Redshift, etc.).
  • Database & Querying: Strong command over SQL programming, performance tuning, and database design principles.
  • Big Data Frameworks: Familiarity with distributed computing principles and the Apache Spark ecosystem.
  • Tools & Version Control: Experience with orchestration tools (e.g., Apache Airflow), containerization, and version control systems (Git).
  • Education: Bachelors or Master’s degree in Computer Science, Information Technology, Engineering, or a related quantitative field.
Preferred / Good-to-Have Skills
  • Exposure to modern data lakehouse platforms like Databricks or Snowflake.
  • Experience working in fast-paced FinTech or Banking domains.
  • Familiarity with CI/CD deployment models and infrastructure-as-code concepts.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AWS & Pyspark- Data Engineer-Immediate joiners only
AWS & Pyspark- Data Engineer-Immediate joiners only

EY • Pune District, Bengaluru, Delhi

Hybrid
INR 1,200,000 - 1,800,000
Senior Data Engineer
Senior Data Engineer

AagatiServe Pvt Ltd • Delhi

On-site
INR 1,800,000 - 2,400,000
Data Engineer — PySpark + AWS Glue
Data Engineer — PySpark + AWS Glue

Yadimen Consulting Limited • Chennai District, Pune District, Bengaluru

On-site
INR 1,000,000 - 1,800,000
AWS Data Engineer
AWS Data Engineer

Weekday AI (YC W21) • Gurugram District

On-site
INR 1,800,000 - 2,500,000
Senior Data Engineer - Pyspark & AWS
Senior Data Engineer - Pyspark & AWS

RBM Software • Pune District

On-site
INR 4,000,000 - 7,000,000
Senior Data Engineer
Senior Data Engineer

Go Digital Technology Consulting • Mumbai

Hybrid
INR 1,200,000 - 1,800,000
Data Engineer -AWS, PySpark, SQL (8+ yrs)
Data Engineer -AWS, PySpark, SQL (8+ yrs)

Banking Tech MNC • Bengaluru

On-site
INR 800,000 - 1,200,000
AWS Data Engineer
AWS Data Engineer

Zorba AI • Bengaluru

On-site
INR 1,500,000 - 2,600,000
AWS Data Engineer
AWS Data Engineer

Zorba AI • Mumbai

On-site
INR 900,000 - 1,800,000
AWS Data Engineer
AWS Data Engineer

Zorba AI • Hyderabad

On-site
INR 900,000 - 1,500,000