Data Engineer (PySpark)

Talent Basket

Bengaluru

On-site

INR 1,200,000 - 1,800,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading company in Human Resources Services seeks a skilled Data Engineer specializing in PySpark to design and maintain scalable data pipelines. The ideal candidate will have a solid background in the Cloudera Data Platform, and big data ecosystems, and play a crucial role in optimizing data processing and ensuring data quality across the organization.

Qualifications

  • 3+ years of experience as a Data Engineer focusing on PySpark and Cloudera.
  • Knowledge of data warehousing and ETL best practices.
  • Strong Linux scripting skills.

Responsibilities

  • Design, develop, and maintain scalable ETL pipelines using PySpark.
  • Implement data ingestion processes from various sources to the data lake.
  • Optimize PySpark code and Cloudera components for performance.

Skills

PySpark
Data Warehousing
Big Data Technologies
Orchestration and Scheduling
Scripting and Automation

Education

Bachelor’s or Master’s degree in Computer Science, Data Engineering, or related fields

Tools

Cloudera Data Platform
Apache Oozie
Airflow

Job description

Get AI-powered advice on this job and more exclusive features.

Job Title: Data Engineer (PySpark)

Location: ENBD Bangalore Office (5 Days)

Skill: PySpark, Cloudera, Hive, Apache Oozie

About The Role

We are seeking a highly skilled Data Engineer with deep expertise in PySpark and the Cloudera Data Platform (CDP) to join our data engineering team. As a Data Engineer, you will be responsible for designing, developing, and maintaining scalable data pipelines that ensure high data quality and availability across the organization. This role requires a strong background in big data ecosystems, cloud-native tools, and advanced data processing techniques.

The ideal candidate has hands-on experience with data ingestion, transformation, and optimization on the Cloudera Data Platform, along with a proven track record of implementing data engineering best practices. You will work closely with other data engineers to build solutions that drive impactful business insights.

Responsibilities
  • Data Pipeline Development: Design, develop, and maintain highly scalable and optimized ETL pipelines using PySpark on the Cloudera Data Platform, ensuring data integrity and accuracy.
  • Data Ingestion: Implement and manage data ingestion processes from various sources (relational databases, APIs, file systems) to the data lake or data warehouse on CDP.
  • Data Transformation and Processing: Use PySpark to process, cleanse, and transform large datasets into meaningful formats supporting analytical needs.
  • Performance Optimization: Conduct performance tuning of PySpark code and Cloudera components, optimizing resource utilization and reducing ETL runtime.
  • Data Quality and Validation: Implement data quality checks, monitoring, and validation routines to ensure data accuracy and reliability.
  • Automation and Orchestration: Automate data workflows using tools like Apache Oozie, Airflow, or similar within the Cloudera ecosystem.
  • Monitoring and Maintenance: Monitor pipeline performance, troubleshoot issues, and perform routine maintenance on the Cloudera Data Platform.
  • Collaboration: Work with data engineers, analysts, product managers, and stakeholders to understand data requirements and support data-driven initiatives.
  • Documentation: Maintain thorough documentation of data engineering processes, code, and pipeline configurations.
Qualifications

Education and Experience

  • Bachelor’s or Master’s degree in Computer Science, Data Engineering, or related fields.
  • 3+ years of experience as a Data Engineer, focusing on PySpark and Cloudera Data Platform.

Technical Skills

  • PySpark: Advanced proficiency, including RDDs, DataFrames, and optimization techniques.
  • Cloudera Data Platform: Experience with components like Cloudera Manager, Hive, Impala, HDFS, HBase.
  • Data Warehousing: Knowledge of data warehousing, ETL best practices, SQL tools like Hive and Impala.
  • Big Data Technologies: Familiarity with Hadoop, Kafka, and distributed computing tools.
  • Orchestration and Scheduling: Experience with Apache Oozie, Airflow, or similar frameworks.
  • Scripting and Automation: Strong Linux scripting skills.
Seniority Level

Mid-Senior level

Employment Type

Full-time

Job Function

Information Technology

Industries

Human Resources Services

Referrals increase your chances of interviewing at Talent Basket by 2x.

We’re unlocking community knowledge in a new way. Experts add insights directly into each article, started with the help of AI.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

Intact Green Services (india) • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Industry-standard compensation
Data Engineer
Data Engineer

EXL • Pune District

On-site
INR 1,200,000 - 2,400,000
Data Engineer
Data Engineer

InfoBeans • Pune City

Hybrid
INR 1,200,000 - 1,800,000
Data Engineer (PySpark + Cloudera)
Data Engineer (PySpark + Cloudera)

Zorba AI • Maharashtra

On-site
INR 1,000,000 - 1,500,000
Data Engineer
Data Engineer

Technogen • Gurugram District, Bengaluru, Pune District

On-site
INR 1,800,000 - 3,000,000
2702 - Data Engineer
2702 - Data Engineer

EXL • Pune City

On-site
INR 1,000,000 - 1,500,000
Spark + Scala+ Python + Github + Copilot
Spark + Scala+ Python + Github + Copilot

Hexaware Technologies • Bengaluru

Hybrid
INR 1,200,000 - 1,800,000
Data Engineer -AWS, PySpark, SQL (8+ yrs)
Data Engineer -AWS, PySpark, SQL (8+ yrs)

Banking Tech MNC • Bengaluru

On-site
INR 800,000 - 1,200,000
PySpark Data Engineer
PySpark Data Engineer

Code1 Tech Systems • India

On-site
INR 1,200,000 - 2,400,000
Data Engineer (Spark, Java, Cloud)
Data Engineer (Spark, Java, Cloud)

Redolent Infotech Pvt. Ltd. • Bengaluru

Hybrid
INR 900,000 - 1,400,000