Pyspark Data Engineer

Clover Infotech

Chennai District

On-site

INR 1,200,000 - 1,800,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Clover Infotech is seeking a data engineer with deep PySpark and Cloudera Data Platform expertise to design, build, and manage scalable data pipelines. You will ensure data quality and availability while implementing best practices for ingestion, transformation, and orchestration across the data stack.

The role emphasizes hands-on development, performance tuning, and collaboration with data teams to deliver business insights through reliable data solutions on CDP.

Qualifications

  • Strong PySpark experience and data engineering skills.
  • Hands-on with Cloudera Data Platform (CDP) and big data ecosystems.
  • Experience with data ingestion, transformation, and ETL best practices.
  • Exposure to orchestration tools like Airflow or Oozie.

Responsibilities

  • Data Pipeline Development: design, develop, and maintain scalable ETL pipelines using PySpark on CDP.
  • Data Ingestion: implement ingestion from databases, APIs, and files to data lake/warehouse on CDP.
  • Data Transformation: process and cleanse large datasets for analytics.
  • Performance Optimization: tune PySpark code and CDP components for efficiency.
  • Data Quality & Validation: implement quality checks and monitoring throughout pipelines.
  • Automation & Orchestration: automate workflows with Airflow/Oozie within CDP.
  • Monitoring & Maintenance: monitor, troubleshoot, and maintain data pipelines.

Skills

PySpark
Cloudera CDP
ETL design
Data quality
Airflow

Tools

Apache Airflow
Apache Oozie
Hadoop ecosystem

Job description

Job Descriptions:


We are seeking a highly skilled Data Engineer with deep expertise in PySpark and the Cloudera Data Platform (CDP) to join our data engineering team. As a Data Engineer, you will be responsible for designing, developing, and maintaining scalable data pipelines that ensure high data quality and availability across the organization. This role requires a strong background in big data ecosystems, cloud-native tools, and advanced data processing techniques.


The ideal candidate has hands‑on experience with data ingestion, transformation, and optimization on the Cloudera Data Platform, along with a proven track record of implementing data engineering best practices. You will work closely with other data engineers to build solutions that drive impactful business insights.


Responsibilities



  • Data Pipeline Development: Design, develop, and maintain highly scalable and optimized ETL pipelines using PySpark on the Cloudera Data Platform, ensuring data integrity and accuracy.

  • Data Ingestion: Implement and manage data ingestion processes from a variety of sources (e.g., relational databases, APIs, file systems) to the data lake or data warehouse on CDP.

  • Data Transformation and Processing: Use PySpark to process, cleanse, and transform large datasets into meaningful formats that support analytical needs and business requirements.

  • Performance Optimization: Conduct performance tuning of PySpark code and Cloudera components, optimizing resource utilization and reducing runtime of ETL processes.

  • Data Quality and Validation: Implement data quality checks, monitoring, and validation routines to ensure data accuracy and reliability throughout the pipeline.

  • Automation and Orchestration: Automate data workflows using tools like Apache Oozie, Airflow, or similar orchestration tools within the Cloudera ecosystem.

  • Monitoring and Maintenance: Monitor pipeline performance, troubleshoot issues

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer (PySpark + Cloudera)
Data Engineer (PySpark + Cloudera)

Zorba AI • Maharashtra

On-site
INR 1,000,000 - 1,500,000
Data Engineer (PySpark)
Data Engineer (PySpark)

Talent Basket • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Pyspark Data Engineer
Pyspark Data Engineer

Synechron • Bengaluru

On-site
INR 900,000 - 1,500,000
Pyspark Engineer
Pyspark Engineer

Synechron • Chennai District

Hybrid
INR 900,000 - 1,300,000
Data Engineer - SQL/PySpark
Data Engineer - SQL/PySpark

Forward Eye Technologies • Mumbai, Bengaluru, New Delhi

On-site
INR 1,800,000 - 3,200,000
PySpark Data Engineer
PySpark Data Engineer

Infosys • Hyderabad, Pune District, Bengaluru

On-site
INR 900,000 - 1,500,000
Data Engineer - SQL/PySpark
Data Engineer - SQL/PySpark

Forward Eye Technologies • Pune District

On-site
INR 1,500,000 - 2,100,000
Data Engineer-Pyspark
Data Engineer-Pyspark

Deloitte US-India Offices • Pune District

On-site
INR 2,000,000 - 4,000,000
Data Engineer
Data Engineer

Alike Thoughts • Bengaluru

On-site
INR 1,200,000 - 1,900,000
PySpark Data Engineer
PySpark Data Engineer

Code1 Tech Systems • India

On-site
INR 1,200,000 - 2,400,000