Data Engineer

Infinite Computer Solutions

Town of Texas (WI)

On-site

USD 120,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Infinite Computer Solutions is seeking an experienced Big Data Engineer to design, build, and optimize large-scale batch and streaming data pipelines on the Hadoop ecosystem using Apache Spark and Scala.

You will work with data architects and analytics teams to ingest, transform, and enrich clickstream, network, and location datasets, ensuring data quality and reliability in production environments.

Qualifications

  • 6+ years of data engineering experience with 4+ years Spark in Scala on production workloads.
  • Strong Scala fundamentals and functional programming concepts.
  • Deep Hadoop ecosystem knowledge: HDFS, Hive, YARN, HBase.
  • Advanced SQL and data modeling skills for big data denormalized patterns.
  • Experience with Spark performance tuning and debugging via Spark UI and logs.
  • Familiarity with Parquet/ORC and data serialization formats and their trade-offs.

Responsibilities

  • Design, develop, and maintain distributed data pipelines using Spark (Core, SQL, Streaming) in Scala.
  • Build ingestion and transformation workflows across Hadoop ecosystem for TB-PB scale data.
  • Tune Spark jobs with partitioning, caching, joins, and memory sizing to optimize performance.
  • Implement real-time/near-real-time ingestion with NiFi and/or Kafka.
  • Embed data quality controls within pipelines and data lineage documentation.
  • Write and optimize HiveQL and Spark SQL for curated and consumption layers.
  • Automate orchestration with Airflow, Oozie, or Control-M.
  • Participate in code reviews, CI/CD, testing, and production support.
  • Troubleshoot failures, SLA breaches, and performance regressions; drive root-cause fixes.
  • Document data flows, lineage, and operational runbooks.

Skills

Spark (Scala)
Hadoop
SQL
Data pipelines
Python/PySpark
CI/CD tooling
Linux scripting

Education

Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline

Tools

HDFS
Hive
YARN
MapReduce
Parquet/ORC
Avro
Git
Maven
SBT
Jenkins

Job description

Job Description

  • Design, develop, and maintain distributed data pipelines using Apache Spark (Core, SQL, Streaming) written in Scala.
  • Build ingestion and transformation workflows across the Hadoop ecosystem - HDFS, Hive, YARN, MapReduce - for structured and semi-structured data at TB-PB scale.
  • Tune and optimize Spark jobs: partitioning strategy, caching, broadcast joins, shuffle reduction, data skew handling, and executor/memory sizing.
  • Implement real-time and near-real-time ingestion using Apache NiFi and/or Kafka.
  • Embed data quality, reconciliation, and validation controls directly into pipelines.
  • Author and optimize HiveQL and Spark SQL for curated and consumption layers.
  • Automate orchestration and scheduling using Airflow, Oozie, or Control-M.
  • Participate in code reviews, CI/CD automation, unit and integration testing, and production support.
  • Troubleshoot job failures, SLA breaches, and performance regressions; drive root-cause analysis to permanent fixes.
  • Document data flows, lineage, transformation logic, and operational runbooks.

Job Description

Role Summary We are seeking an experienced Big Data Engineer to design, build, and optimize large-scale batch and streaming data pipelines on the Hadoop ecosystem using Apache Spark and Scala. The role supports high-volume ingestion, transformation, and enrichment of clickstream, network, and location datasets, working closely with data architects, platform engineering, and downstream analytics teams. This is a hands-on engineering role with ownership of pipeline performance, reliability, and data quality in production.

Key Responsibilities

  • Design, develop, and maintain distributed data pipelines using Apache Spark (Core, SQL, Streaming) written in Scala.
  • Build ingestion and transformation workflows across the Hadoop ecosystem - HDFS, Hive, YARN, MapReduce - for structured and semi-structured data at TB-PB scale.
  • Tune and optimize Spark jobs: partitioning strategy, caching, broadcast joins, shuffle reduction, data skew handling, and executor/memory sizing.
  • Implement real-time and near-real-time ingestion using Apache NiFi and/or Kafka.
  • Embed data quality, reconciliation, and validation controls directly into pipelines.
  • Author and optimize HiveQL and Spark SQL for curated and consumption layers.
  • Automate orchestration and scheduling using Airflow, Oozie, or Control-M.
  • Participate in code reviews, CI/CD automation, unit and integration testing, and production support.
  • Troubleshoot job failures, SLA breaches, and performance regressions; drive root-cause analysis to permanent fixes.
  • Document data flows, lineage, transformation logic, and operational runbooks.

Required Qualifications

  • 6+ years of data engineering experience, with 4+ years hands-on Apache Spark development in Scala on production workloads.
  • Strong Scala fundamentals - functional programming constructs, collections API, case classes, pattern matching, implicits, and error handling.
  • Deep working knowledge of the Hadoop ecosystem: HDFS, Hive, YARN, HBase.
  • Advanced SQL and data modeling skills across dimensional and big-data denormalized patterns.
  • Demonstrated Spark performance tuning and debugging using the Spark UI, event logs, and physical execution plans.
  • Proficiency with columnar and serialization formats - Parquet, ORC, Avro - including compression and partitioning trade-offs.
  • Linux and shell scripting, Git, Maven or SBT, and Jenkins or equivalent CI/CD tooling.
  • Ability to work independently in a distributed onshore-offshore delivery model.

Preferred Qualifications

  • Kafka and Spark Structured Streaming for event-driven pipelines.
  • Cloud data platform exposure - GCP (BigQuery, Dataproc), AWS EMR, or Azure Databricks.
  • Telecom domain experience with clickstream, network, or geospatial/location data.
  • Python or PySpark as a secondary development language.
  • Data governance and security frameworks - Apache Ranger, Kerberos, PII masking and tokenization

Nice to Have

  • Apache NiFi flow design, configuration, and administration.

Education Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline - or equivalent demonstrable practical experience.

Qualifications Bachelor

Range Of Year Experience-Min Year 15

Range Of Year Experience-Max Year 18

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

Tata Consultancy Services • Irving (TX)

On-site
USD 125,000 - 140,000
Big Data Lead
Big Data Lead

Veriipro • United States

On-site
USD 180,000 - 240,000
Senior Data Engineer
Senior Data Engineer

Compunnel, Inc. • Bentonville (AR)

On-site
USD 100,000 - 130,000
Java Spark Engineer
Java Spark Engineer

Veriipro • Berkeley Heights (NJ)

On-site
USD 140,000 - 190,000
Data Engineer ETL
Data Engineer ETL

Compunnel, Inc. • Durham (NC)

On-site
USD 100,000 - 130,000
Senior Data Engineer
Senior Data Engineer

Compunnel, Inc. • Sunnyvale (CA)

On-site
USD 100,000 - 140,000
Bigdata Engineer
Bigdata Engineer

Disys - Oak Brook • Tampa (FL)

On-site
USD 90,000 - 120,000
Data Engineer
Data Engineer

The Value Maximizer • United States

On-site
USD 90,000 - 120,000
Engineer
Engineer

Tata Consultancy Services Limited • Irving (TX)

On-site
USD 70,000 - 80,000
Big Data Engineer
Big Data Engineer

Ex • Pittsburgh

On-site
USD 120,000 - 180,000