PySpark Data Engineer

Code1 Tech Systems

India

On-site

INR 1,200,000 - 2,400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Code1 Tech in India is seeking a PySpark Data Engineer with 6+ years of hands-on experience to design, develop, and optimize large-scale data processing pipelines. You will work with Apache Spark, Python, Oracle SQL, and HDFS in production environments.

You will build scalable data ingestion frameworks, develop ETL/ELT workflows, and tune Spark performance. The role demands strong knowledge of Spark architecture, data cleansing, and scripting in Python and Shell.

Qualifications

  • 6+ years of hands-on experience designing, building, and maintaining PySpark data pipelines in production.

Responsibilities

  • Design and maintain scalable PySpark data pipelines for batch and large-scale processing.
  • Build data ingestion frameworks from multiple source systems.
  • Develop optimized ETL/ELT workflows for high-volume enterprise datasets.
  • Process, transform, and optimize datasets exceeding 500GB while ensuring performance.
  • Perform data cleansing, normalization, validation, and formatting to improve data quality.
  • Write efficient Oracle SQL queries for data extraction, transformation, and analysis.

Skills

PySpark
Spark Pipelines
Oracle SQL
HDFS
Python
Shell Scripting
Data Ingestion

Job description

At Code1 Tech, we drive innovations that shape the future of enterprise technology. Our expertise spans Data Engineering, AI/ML, Cloud Solutions, and Full-Stack Development. We empower businesses with cutting‑edge technology solutions, enabling digital transformation at scale. Join us to build impactful products with a passionate team of engineers and innovators.

About the Role

We are looking for a skilled PySpark Data Engineer with strong expertise in designing, developing, and optimizing large‑scale data processing pipelines. The ideal candidate should have hands‑on experience with Apache Spark, Python, Oracle SQL, and HDFS, along with a proven track record of building scalable data ingestion frameworks and processing high‑volume datasets in production environments.

Experience: 6+ Years

Key Responsibilities:
  • Design, develop, and maintain scalable Apache Spark (PySpark) data pipelines for batch and large‑scale data processing.
  • Build and enhance data ingestion frameworks integrating data from multiple source systems.
  • Develop optimized ETL/ELT workflows for high‑volume enterprise datasets.
  • Process, transform, and optimize datasets exceeding 500GB while ensuring performance and reliability.
  • Perform data cleansing, normalization, validation, and formatting to improve data quality.
  • Write efficient Oracle SQL queries for data extraction, transformation, and analysis.
  • Work with HDFS and distributed data processing environments.
  • Optimize Spark jobs by tuning partitioning, shuffling, caching, and resource utilization.
  • Automate manual data processing workflows using Python and Shell Scripting.
  • Monitor, troubleshoot, and optimize production data pipelines for performance and scalability.
  • Collaborate with data architects, analysts, and business stakeholders to deliver robust data solutions.
Mandatory Technical Skills:
  • Hands‑on expertise designing, building, and maintaining Apache Spark pipelines in production environments.
  • Experience building scalable data ingestion frameworks integrating multiple source systems.
  • Strong understanding of Spark Architecture including Driver/Executors, DAG, Partitioning, Shuffles, Caching, and Resource Management.
  • Experience processing and transforming datasets larger than 500GB.
  • Strong Oracle SQL and HDFS knowledge.
  • Experience handling data cleansing, normalization, and data formatting processes.
  • Strong Python, PySpark, and Shell Scripting skills.
  • Experience automating manual data processing workflows using Python.
Ready to make an impact?

Apply now and join our team of innovators.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Developer - PySpark
Developer - PySpark

Compunnel, Inc. • Pune District

On-site
INR 800,000 - 1,500,000
Pyspark Data Engineer
Pyspark Data Engineer

Synechron • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Data Engineer
Data Engineer

Intact Green Services (india) • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Industry-standard compensation
PySpark Big Data Developer
PySpark Big Data Developer

Citi • Maharashtra

On-site
INR 1,100,000 - 1,800,000
PySpark Developer / Senior Data Engineer
PySpark Developer / Senior Data Engineer

Alignity Solutions • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Senior PySpark ETL Lead Engineer
Senior PySpark ETL Lead Engineer

Relevantz Technology Services • Chennai District

On-site
INR 1,500,000 - 2,100,000
Data Engineer (Python & PySpark)
Data Engineer (Python & PySpark)

Techknomatic Services • Pune District

On-site
INR 700,000 - 1,200,000
Data Engineer
Data Engineer

EXL • Pune District

On-site
INR 1,200,000 - 2,400,000
Data Engineer (Spark/Scala)
Data Engineer (Spark/Scala)

Zorba AI • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Contractor - PySpark Engineer
Contractor - PySpark Engineer

Vivantify • Hyderabad

On-site
INR 1,800,000 - 2,600,000