Data Engineer (Spark/Scala)

Zorba AI

Hyderabad

On-site

INR 1,800,000 - 3,000,000

Full time

30 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Zorba AI is seeking an experienced Data Engineer specializing in Spark/Scala to design and implement large-scale data pipelines across on-prem and cloud environments. You will work across HDFS, NAS, and S3, handling diverse data formats and ensuring reliable, scalable data processing.

Ideal candidates bring hands-on Spark/Databricks expertise, strong Python, PySpark, and SQL skills, and a track record of building hybrid on-prem/cloud data pipelines with robust testing and documentation.

Qualifications

  • Proven hands-on experience with Apache Spark and Databricks.
  • Strong Python and PySpark programming skills.
  • Solid SQL knowledge with complex joins and optimization.

Responsibilities

  • Design, develop, and maintain scalable data pipelines using Spark, Databricks, Scala Spark, and PySpark.
  • Build and support hybrid on-prem-to-cloud data integration solutions.
  • Integrate data across HDFS, NAS, on-prem file shares, and cloud storage like Amazon S3.

Skills

Apache Spark
Databricks
Scala
PySpark
Python
SQL
Amazon S3
HDFS
Airflow
Azure
On-Premise Data Engineering
ETL
Data Pipelines

Tools

Airflow
Prefect
React

Job description

We are looking for an experienced Data Engineer (Spark/Scala) with strong hands-on expertise in Apache Spark, Databricks, Scala, PySpark, Python, and SQL.

The role involves designing and developing large-scale data pipelines across on-premises and cloud environments, working with multiple file systems and data formats, modernizing legacy workflows, and supporting hybrid data architectures.

Key Responsibilities
  • Design, develop, and maintain scalable data pipelines using Apache Spark, Databricks, Scala Spark, and PySpark.
  • Build and support complex on-premises data workflows and hybrid on-prem-to-cloud data integration solutions.
  • Integrate data across HDFS, NAS, on-prem file shares, Amazon S3, and other storage platforms.
  • Work with multiple data formats including JSON, Parquet, CSV, Avro, Fixed-Length, and Excel.
  • Develop optimized SQL queries for data extraction, transformation, and loading.
  • Connect to multiple relational and non-relational databases and implement performance-efficient data extraction strategies.
  • Develop and maintain workflow orchestration using Apache Airflow or similar scheduling tools.
  • Write clean, production-grade Python code for data processing, automation, and engineering utilities.
  • Develop unit, integration, and data-quality tests for data pipelines.
  • Troubleshoot pipeline failures, performance bottlenecks, data quality issues, and complex multi-system integration problems.
  • Support migration and modernization of legacy on-premises data processes to hybrid/cloud environments.
  • Collaborate with Data Scientists, Analysts, Application Engineers, and other stakeholders.
  • Create technical documentation covering pipelines, data flows, architecture, and data lineage.
  • Support cloud integration initiatives, particularly across Azure environments.
  • Leverage coding assistants and AI agents to improve development productivity and automate engineering tasks.
Required Skills
  • Strong hands-on experience with Apache Spark and Databricks.
  • Strong experience with Scala/Spark Scala and PySpark.
  • Strong Python programming skills.
  • Strong SQL, including complex joins, query optimization, and performance tuning.
  • Hands-on experience with Amazon S3.
  • Experience working with HDFS, NAS, on-prem file systems, and cloud storage.
  • Strong experience handling JSON, Parquet, CSV, Avro, Fixed-Length, and Excel data formats.
  • Experience extracting data efficiently from multiple databases.
  • Strong understanding of complex on-premises data workflows and multi-system integrations.
  • Experience building hybrid on-prem/cloud data pipelines.
  • Strong troubleshooting and production support skills.
Secondary Skills
  • Azure cloud services.
  • Apache Airflow or similar workflow orchestration tools.
  • Automated unit and integration testing.
  • Data quality validation and monitoring.
  • Data lineage and technical documentation.
Good to Have
  • Working knowledge of Java.
  • Experience with Prefect.
  • Familiarity with React for internal tools or dashboards.
  • Experience using AI coding assistants and AI agents.
  • PBM / Pharmacy Benefit Management / Healthcare domain experience.
Preferred Candidate Profile

Candidates with strong experience in Scala + Spark + Databricks + PySpark, combined with on-premises data engineering and hybrid cloud integration, will be preferred.

Key Skills:

Apache Spark, Scala, Spark Scala, Databricks, PySpark, Python, SQL, Amazon S3, HDFS, Airflow, Azure, On-Premise Data Engineering, ETL, Data Pipelines.

Skills: cloud,apache spark,scala

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Kolkata District

On-site
INR 1,000,000 - 1,500,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Maharashtra

On-site
INR 800,000 - 1,200,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Mumbai

On-site
INR 1,000,000 - 1,500,000
Data Engineer
Data Engineer

HGS • Hyderabad

On-site
INR 1,200,000 - 2,100,000
Data Engineer – Databricks, Python, SQL & Cloud Migration
Data Engineer – Databricks, Python, SQL & Cloud Migration

Qloron Pvt Ltd • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Senior Data Specialist
Senior Data Specialist

Jobtailor • Bengaluru

On-site
INR 1,500,000 - 2,600,000
Data Engineer
Data Engineer

Altysys • Gurugram District

On-site
INR 800,000 - 1,200,000
Data Engineer (Azure & Databricks)
Data Engineer (Azure & Databricks)

Lufthansa Technik Services India Pvt Ltd • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Data Engineer (Azure & Databricks)
Data Engineer (Azure & Databricks)

Lufthansa Technik Services India • Bengaluru

On-site
INR 1,200,000 - 1,500,000
Data Engineer
Data Engineer

Tekskills • Chennai District

On-site
INR 2,000,000 - 4,000,000