Data Engineer

EXL

Pune District

On-site

INR 1,200,000 - 2,400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

EXL is seeking a Data Engineer in Pune to design, build, and optimize scalable data pipelines using PySpark, Python, Hive, and Oozie. You will collaborate with cross-functional teams to support analytics initiatives and handle enterprise data processing needs.

The role requires 5+ years of hands-on experience with big data technologies, SQL, ETL/ELT concepts, and cloud platforms, preferably AWS. Strong problem-solving and data governance mindset are essential.

Qualifications

  • 5+ years of hands-on experience in PySpark, Python, and SQL.
  • Strong experience with Hive and Oozie.
  • Solid ETL/ELT knowledge and data integration techniques.
  • Experience with AWS cloud platforms.
  • Experience handling large datasets and optimizing pipelines.

Responsibilities

  • Design, develop, and maintain scalable data pipelines using PySpark, Python, Hive, and Oozie.
  • Develop efficient SQL queries for data extraction, transformation, validation, and reporting.
  • Integrate data from multiple source systems ensuring accuracy and completeness.
  • Monitor, troubleshoot, and optimize data pipelines for performance and reliability.
  • Collaborate with stakeholders to understand data requirements and deliver effective solutions.
  • Implement best practices for data engineering, documentation, and coding standards.
  • Support data quality and governance initiatives.
  • Participate in code reviews, testing, deployment, and production support.

Skills

PySpark
Python
SQL
Big Data
Hive
Oozie
AWS
ETL/ELT
Data Warehousing
Databricks
Linux/Unix
Dagster
DBT

Tools

Databricks
Dagster
Linux/Unix

Job description

Position Summary

We are looking for a skilled Data Engineer with 5+ years of experience in building, maintaining, and optimizing scalable data pipelines using PySpark, Python, and SQL. The ideal candidate should possess strong hands‑on expertise in Big Data technologies such as Hive and Oozie, along with a solid understanding of ETL processes, data processing frameworks, and cloud environments (preferably AWS).

The role involves collaborating with cross‑functional teams to support enterprise data platforms, analytics initiatives, and business‑critical data processing requirements.

Key Responsibilities

  • Design, develop, and maintain scalable data pipelines using PySpark, Python, Hive, and Oozie.
  • Develop efficient SQL queries for data extraction, transformation, validation, and reporting.
  • Integrate data from multiple source systems while ensuring accuracy, consistency, and completeness.
  • Monitor, troubleshoot, and optimize data pipelines to improve performance and reliability.
  • Collaborate with business stakeholders, analysts, and engineering teams to understand data requirements and deliver effective solutions.
  • Implement best practices for data engineering, documentation, coding standards, and performance optimization.
  • Support data quality initiatives and ensure adherence to data governance standards.
  • Participate in code reviews, testing, deployment, and production support activities.

Required Skills

  • 5+ years of hands‑on experience in PySpark, Python, and SQL.
  • Strong experience working with Big Data technologies, including Hive and Oozie .
  • Good understanding of ETL/ELT concepts, data processing, and data integration techniques.
  • Working knowledge of cloud platforms, preferably AWS.
  • Experience handling large‑scale datasets and optimizing data processing workflows.
  • Strong analytical and problem‑solving skills.

Secondary Skills

  • Basic to intermediate experience with Databricks.
  • Familiarity with Linux/Unix commands and environment management on edge nodes.
  • Exposure to workflow scheduling and orchestration tools.
  • Knowledge of monitoring and debugging data pipelines.

Good to Have

  • Exposure to the Financial Services/Banking domain.
  • Understanding of Data Warehousing concepts and best practices.
  • Experience with modern data transformation frameworks such as DBT.
  • Exposure to orchestration platforms such as Dagster.
  • Knowledge of cloud-based data engineering architectures and modern data platforms.

Technical Skills

PySpark, Python, SQL, Hive, Oozie, Hadoop, AWS, Databricks, Linux/Unix, ETL/ELT, Data Warehousing, DBT, Dagster

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

Keka Technologies Private Limited • Indore District

On-site
INR 800,000 - 1,300,000
Data Engineer
Data Engineer

Qcentrio • Kanpur

On-site
INR 1,200,000 - 2,100,000
Data Engineer
Data Engineer

Qcentrio • New Delhi

On-site
INR 3,500,000 - 6,500,000
Data Engineer
Data Engineer

Intact Green Services (india) • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Industry-standard compensation
Data Engineer
Data Engineer

Advance Career Solutions • Pune District, Chennai District, Bengaluru

Hybrid
INR 1,200,000 - 2,800,000
Big Data Engineer - Python
Big Data Engineer - Python

Qcentrio • Ahmedabad District

On-site
INR 2,500,000 - 4,000,000
Big Data Engineer - Python
Big Data Engineer - Python

Qcentrio • Jaipur

On-site
INR 2,500,000 - 4,000,000
Big Data Engineer - Python
Big Data Engineer - Python

Qcentrio • Kolkata District

On-site
INR 1,800,000 - 3,200,000
Big Data Engineer - Python
Big Data Engineer - Python

Qcentrio • Surat

On-site
INR 350,000 - 600,000
Data Engineer
Data Engineer

Qcentrio • Hyderabad

On-site
INR 1,800,000 - 3,200,000