Senior Software Engineer

Infinite Computer Solutions

Chennai District

On-site

INR 1,200,000 - 1,800,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Infinite Computer Solutions is seeking an experienced Big Data Engineer to design and optimize large-scale batch and streaming data pipelines on the Hadoop ecosystem using Apache Spark and Scala. The role involves high-volume ingestion, transformation, and enrichment of clickstream, network, and location datasets, with ownership of performance, reliability, and data quality in production.

The candidate will work with data architects, platform engineering, and analytics teams in a hands-on role,

Qualifications

  • 4+ years of data engineering experience with Spark/Scala on production workloads.
  • Deep knowledge of the Hadoop ecosystem: HDFS, Hive, YARN, HBase.
  • Advanced SQL and data modeling for large datasets.
  • Experience tuning Spark performance using UI, logs and execution plans.

Responsibilities

  • Design, develop, and maintain Spark-based pipelines in Scala.
  • Ingest, transform, and enrich data at scale from clickstream, network, and geospatial sources.
  • Tune job performance, memory sizing, and partitioning strategies.
  • Collaborate with data architects, platform teams and downstream analytics.
  • Automate orchestration with Airflow, Oozie, or similar tools.
  • Participate in code reviews, CI/CD, and production support.

Skills

Spark development (Scala)
Scala fundamentals
Hadoop ecosystem
SQL & data modeling
Spark performance tuning
CI/CD tooling

Education

Bachelor's degree

Tools

Spark
NiFi
Kafka
Airflow/Oozie
Hive/Hadoop tools
Jenkins

Job description

Job Description

Job title: Hadoop, Spark Engineer — Apache Spark / Scala / Hadoop

Role Summary

We are seeking an experienced Big Data Engineer to design, build, and optimize large-scale batch and streaming data pipelines on the Hadoop ecosystem using Apache Spark and Scala. The role supports high-volume ingestion, transformation, and enrichment of clickstream, network, and location datasets, working closely with data architects, platform engineering, and downstream analytics teams. This is a hands-on engineering role with ownership of pipeline performance, reliability, and data quality in production.

Key Responsibilities
  • Design, develop, and maintain distributed data pipelines using Apache Spark (Core, SQL, Streaming) written in Scala.
  • Build ingestion and transformation workflows across the Hadoop ecosystem — HDFS, Hive, YARN, MapReduce — for structured and semi-structured data at TB-PB scale.
  • Tune and optimize Spark jobs: partitioning strategy, caching, broadcast joins, shuffle reduction, data skew handling, and executor/memory sizing.
  • Implement real-time and near-real-time ingestion using Apache NiFi and/or Kafka.
  • Embed data quality, reconciliation, and validation controls directly into pipelines.
  • Author and optimize HiveQL and Spark SQL for curated and consumption layers.
  • Automate orchestration and scheduling using Airflow, Oozie, or Control-M.
  • Participate in code reviews, CI/CD automation, unit and integration testing, and production support.
  • Troubleshoot job failures, SLA breaches, and performance regressions; drive root-cause analysis to permanent fixes.
  • Document data flows, lineage, transformation logic, and operational runbooks.
Required Qualifications
  • 4+ years of data engineering experience, with 2+ years hands-on Apache Spark development in Scala on production workloads.
  • Strong Scala fundamentals — functional programming constructs, collections API, case classes, pattern matching, implicit, and error handling.
  • Deep working knowledge of the Hadoop ecosystem: HDFS, Hive, YARN, HBase.
  • Advanced SQL and data modeling skills across dimensional and big-data denormalized patterns.
  • Demonstrated Spark performance tuning and debugging using the Spark UI, event logs, and physical execution plans.
  • Proficiency with columnar and serialization formats — Parquet, ORC, Avro — including compression and partitioning trade-offs.
  • Linux and shell scripting, Git, Maven or SBT, and Jenkins or equivalent CI/CD tooling.
  • Ability to work independently in an onshore-offshore delivery model.
Preferred Qualifications
  • Kafka and Spark Structured Streaming for event-driven pipelines.
  • Cloud data platform exposure — GCP (BigQuery, Dataproc), AWS EMR, or Azure Databricks.
  • Telecom domain experience with clickstream, network, or geospatial/location data.
  • Python or PySpark as a secondary development language.
  • Data governance and security frameworks — Apache Ranger, Kerberos, PII masking and tokenization.
Nice to Have
  • Apache NiFi flow design, configuration, and administration.
Education

Bachelor’s degree in computer science, Information Technology, Engineering, or a related discipline — or equivalent demonstrable practical experience.

Qualifications

Bachelor's

Range Of Year Experience

Min Year: 4

Max Year: 6

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer
Senior Data Engineer

SourcingXPress • Hyderabad

On-site
INR 1,500,000 - 2,500,000
Big Data Engineer
Big Data Engineer

Tata Consultancy Services • Pune District

On-site
INR 1,400,000 - 2,100,000
Big Data Developer
Big Data Developer

Metaplore Solutions Pvt Ltd • Bengaluru

On-site
INR 4,500,000 - 7,500,000
Hadoop Data Engineer
Hadoop Data Engineer

Incedo Inc. • Hyderabad

On-site
INR 1,200,000 - 1,900,000
Software Engineer – Data Platform
Software Engineer – Data Platform

Alegeus • Bengaluru

On-site
INR 4,200,000 - 6,600,000
Software Engineer
Software Engineer

Alegeus • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Data Engineering - Senior / Lead Engineer
Data Engineering - Senior / Lead Engineer

Paytm • Dadri

On-site
INR 1,200,000 - 2,000,000
Data Engineer (Spark/Scala)
Data Engineer (Spark/Scala)

Zorba AI • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Senior Data Engineer
Senior Data Engineer

Luxoft • Bengaluru

On-site
INR 1,500,000 - 2,000,000
Big Data Developer
Big Data Developer

Unify Technologies • Hyderabad

On-site
INR 1,200,000 - 1,800,000