Data Engineer-Spark,Scala

Zorba AI

Chennai District

On-site

INR 900,000 - 1,500,000

Full time

22 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Zorba AI in Chennai is seeking an experienced Data Engineer (Spark/Scala) to design, build, and optimize data workflows across on-premises and cloud environments. You will work with Spark, Databricks, Scala/PySpark, SQL, and Python to create robust, high-performance pipelines across diverse file systems and formats.

You will migrate legacy processes, ensure reliable data movement, and contribute to hybrid on-prem-to-cloud integrations while collaborating with data science and analytics teams.

Qualifications

  • Strong hands-on experience with Apache Spark and Databricks for large-scale data processing.
  • Proficiency with Amazon S3 for data storage and pipeline integration.
  • Strong SQL skills and query optimization.
  • Experience integrating data across various file systems and formats.
  • Knowledge of Scala Spark and PySpark for distributed processing.
  • Strong Python programming skills for scripting and data engineering tasks.
  • Experience connecting to and extracting data from multiple databases with performance awareness.
  • Experience with on-premises data workflows and hybrid on-prem/cloud pipelines.
  • Experience leveraging coding assistant tools and AI agents to boost productivity.

Responsibilities

  • Design, develop, and maintain large-scale data pipelines using Spark and Databricks.
  • Support on-premises data workflows with hybrid on-prem-to-cloud integration.
  • Integrate data across file systems and formats (JSON, Parquet, CSV, etc.).
  • Write optimized SQL for data extraction and loading.
  • Connect to various databases and optimize extraction.
  • Develop workflow orchestration with Airflow for reliable pipelines.
  • Write production-grade Python code for data processing.
  • Maintain tests and documentation for pipelines and architecture.
  • Troubleshoot pipeline failures and performance bottlenecks.
  • Collaborate with data science and analytics teams.
  • Support cloud integration efforts, especially Azure, as workloads evolve.

Skills

Spark & Databricks
SQL
Python
Scala
Data pipelines
Airflow

Tools

Amazon S3
Parquet/JSON
Azure

Job description

About The Role

We are seeking an experienced Data Engineer to design, build, and optimize complex data workflows across on‑premises and cloud environments. This role requires deep hands‑on expertise in Apache Spark, Databricks, and Scala/PySpark, along with strong SQL and Python skills, to build robust, high‑performance data pipelines. You will work extensively on complex on‑prem workflows, integrating data across multiple file systems and formats, migrating and modernizing legacy processes, and ensuring efficient, reliable data movement across heterogeneous environments.

Data Engineer (Spark/Scala)

We are seeking an experienced Data Engineer to design, build, and optimize complex data workflows across on‑premises and cloud environments. This role requires deep hands‑on expertise in Apache Spark, Databricks, and Scala/PySpark, along with strong SQL and Python skills, to build robust, high‑performance data pipelines. You will work extensively on complex on‑prem workflows, integrating data across multiple file systems and formats, migrating and modernizing legacy processes, and ensuring efficient, reliable data movement across heterogeneous environments.

Key Responsibilities
  • Design, develop, and maintain large‑scale data pipelines using Apache Spark, Databricks, Scala Spark, and PySpark
  • Build and support complex on‑premises data workflows, including migration/hybrid on‑prem-to-cloud integration patterns
  • Integrate data across diverse file systems (on‑prem file shares, NAS, HDFS, S3) and formats JSON, Parquet, Fixed‑Length, CSV, Excel, Avro
  • Write efficient, optimized SQL for data extraction, transformation, and loading across relational databases
  • Connect to and extract data efficiently from various source databases, tuning queries and pipelines for performance at scale
  • Develop and maintain workflow orchestration using Airflow (or similar schedulers) for reliable, monitored pipeline execution
  • Write clean, production‑grade Python code for data processing, automation, and tooling
  • Build and maintain unit/integration tests for data pipelines to ensure data quality and reliability
  • Create and maintain clear technical documentation for pipelines, data flows, and system architecture
  • Troubleshoot and resolve data pipeline failures, performance bottlenecks, and data quality issues in complex, multi‑system workflows
  • Collaborate with cross‑functional teams (data science, analytics, application engineering) to support downstream data consumption
  • Support cloud integration efforts, particularly with Azure, as workloads evolve from on‑prem to hybrid/cloud architectures
Required Qualifications
Primary Skills
  • Strong hands‑on experience with Apache Spark and Databricks for large‑scale data processing
  • Proficiency with Amazon S3 for data storage and pipeline integration
  • Strong SQL skills - query optimization, complex joins, performance tuning
  • Proven experience integrating data across various file systems and formats: JSON, Parquet, Fixed‑Length, CSV, Excel, Avro, etc.
  • Strong knowledge of Scala Spark and PySpark for distributed data processing
  • Strong Python programming skills for scripting, automation, and data engineering tasks
  • Strong experience connecting to and efficiently extracting data from databases (relational/other), including performance‑conscious extraction strategies
  • Demonstrated experience working on complex on‑prem data workflows (multi‑system integration, legacy system data extraction, hybrid on‑prem/cloud pipelines)
  • Experience leveraging coding assistant tools and implementing AI agents to enhance development productivity and task execution.
Secondary Skills
  • Experience with Azure cloud services (storage, compute, data services)
  • Experience with Apache Airflow for workflow orchestration and scheduling
  • Experience writing automated tests for data pipelines (unit, integration, data quality checks)
  • Strong documentation skills able to clearly document pipelines, data lineage, and technical designs
Good To Have
  • Working knowledge of Java
  • Familiarity with React for building internal tooling/dashboards
  • Experience with Prefect for workflow orchestration
  • PBM (Pharmacy Benefit Management) / Healthcare domain knowledge
Data Engineer (Spark/Scala)
Good To Have
  • Working knowledge of Java
  • Familiarity with React for building internal tooling/dashboards
  • Experience with Prefect for workflow orchestration
  • PBM (Pharmacy Benefit Management) / Healthcare domain knowledge.
Data Engineer (Spark/Scala)
Good To Have
  • Working knowledge of Java
  • Familiarity with React for building internal tooling/dashboards
  • Experience with Prefect for workflow orchestration
  • PBM (Pharmacy Benefit Management) / Healthcare domain knowledge

Skills: scala,pipelines,data,spark

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer-Spark,Scala
Data Engineer-Spark,Scala

Zorba AI • Bengaluru

On-site
INR 1,200,000 - 2,400,000
Data Engineer (Spark/Scala)
Data Engineer (Spark/Scala)

Zorba AI • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Kolkata District

On-site
INR 1,000,000 - 1,500,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Mumbai

On-site
INR 1,000,000 - 1,500,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Maharashtra

On-site
INR 800,000 - 1,200,000
Senior Data Specialist
Senior Data Specialist

Jobtailor • Bengaluru

On-site
INR 1,500,000 - 2,600,000
Data engineer with Scala + spark + SQL
Data engineer with Scala + spark + SQL

HMG AMERICA LLC • Bangalore Rural

On-site
INR 1,800,000 - 3,200,000
Senior Data Engineer
Senior Data Engineer

SourcingXPress • Hyderabad

On-site
INR 1,500,000 - 2,500,000
Data Engineer
Data Engineer

HGS • Hyderabad

On-site
INR 1,200,000 - 2,100,000
Data Engineer (Azure & Databricks)
Data Engineer (Azure & Databricks)

Lufthansa Technik Services India • Bengaluru

On-site
INR 1,200,000 - 1,500,000