Java Spark Engineer

Delta System & Software, Inc.

Berkeley Heights (NJ)

On-site

USD 150,000 - 210,000

Full time

5 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Delta System & Software, Inc. is seeking a senior data engineer to architect and build scalable, fault-tolerant data pipelines using Java and Apache Spark. You will lead design of batch and streaming ETL/ELT systems, optimize performance, and set coding standards.

You will mentor engineers, own production reliability, partner with product and analytics teams, and drive decisions on data architecture, storage formats, and cluster orchestration.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.
  • 7+ years of professional Java development experience.
  • 5+ years hands-on experience with Apache Spark in production environments.
  • Expert-level understanding of distributed systems: fault tolerance, data locality, shuffle mechanics, resource management.
  • Proven track record designing systems processing terabyte+ scale data.
  • Strong SQL skills and familiarity with columnar storage formats.

Responsibilities

  • Architect and build scalable, fault-tolerant data pipelines using Java and Apache Spark.
  • Lead design of batch and streaming ETL/ELT systems handling large data volumes.
  • Deep-dive performance tuning: partitioning strategy, memory management, shuffle/skew optimization, job cost reduction.
  • Set coding standards and lead code/design reviews across the team.
  • Drive technical decisions on data architecture, storage formats, and pipeline orchestration.
  • Mentor mid-level and junior engineers; act as a technical escalation point.
  • Partner with product, analytics, and platform teams to translate requirements into scalable systems.
  • Own production reliability — on-call ownership, incident response, root-cause analysis for pipeline failures.
  • Evaluate and introduce new tools/frameworks where they improve the system.
  • Contribute to capacity planning and cost optimization for cluster infrastructure.

Skills

Java
Apache Spark
Distributed systems
SQL
CI/CD
Kubernetes
Performance tuning
Kafka

Education

Bachelor’s or Master’s degree in Computer Science, Engineering, or related field

Tools

YARN
Kubernetes
Delta Lake
Parquet/ORC/Avro
Kafka

Job description

  • Architect and build scalable, fault-tolerant data pipelines using Apache Spark (Java)
  • Lead design of batch and streaming ETL/ELT systems handling large data volumes
  • Deep-dive performance tuning: partitioning strategy, memory management, shuffle/skew optimization, job cost reduction
  • Set coding standards and lead code/design reviews across the team
  • Drive technical decisions on data architecture, storage formats, and pipeline orchestration
  • Mentor mid-level and junior engineers; act as a technical escalation point
  • Partner with product, analytics, and platform teams to translate requirements into scalable systems
  • Own production reliability — on-call ownership, incident response, root-cause analysis for pipeline failures
  • Evaluate and introduce new tools/frameworks where they improve the system
  • Contribute to capacity planning and cost optimization for cluster infrastructure
Required Qualifications
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or related field
  • 7+ years of professional Java development experience
  • 5+ years hands-on experience with Apache Spark in production environments
  • Expert-level understanding of distributed systems: fault tolerance, data locality, shuffle mechanics, resource management
  • Proven track record designing systems processing terabyte+ scale data
  • Strong SQL skills and deep familiarity with columnar storage formats (Parquet, ORC, Avro, Delta Lake/Iceberg)
  • Experience with cluster managers (YARN, Kubernetes) and cloud-managed Spark
  • Proficiency with Kafka
  • Strong grasp of CI/CD, containerization, and infrastructure-as-code practices
Preferred Qualifications
  • Experience with Flink or other stream-processing frameworks
  • Familiarity with data governance, lineage, and quality frameworks
  • Experience with workflow orchestration at scale
  • Background in system design for multi-tenant or multi-region data platforms
  • Prior experience leading a team or acting as a technical lead
  • Excellent communication — able to explain technical tradeoffs to non-technical stakeholders
  • Strong mentorship and coaching ability
  • Comfortable driving ambiguous, cross-team technical initiatives
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Java Spark Engineer
Java Spark Engineer

Saransh Inc • Berkley (AL)

On-site
USD 150,000 - 210,000
Data Engineer ETL
Data Engineer ETL

Compunnel, Inc. • Durham (NC)

On-site
USD 100,000 - 130,000
Pyspark Developer
Pyspark Developer

Tieto • Irving (TX)

On-site
USD 140,000 - 190,000
Pyspark Architect
Pyspark Architect

Avance Consulting • Charlotte (NC)

On-site
USD 100,000 - 130,000
Senior Data Engineer – Enterprise Data Frameworks
Senior Data Engineer – Enterprise Data Frameworks

Citizens • Rhode Island

On-site
USD 120,000 - 160,000
PySpark Data Engineer
PySpark Data Engineer

Tata Consultancy Services • Irving (TX)

On-site
USD 90,000 - 120,000
Discretionary Annual Incentive
Medical Coverage (Health, Dental & Vis
401K Plan
+1
Data Engineer
Data Engineer

The Value Maximizer • United States

On-site
USD 90,000 - 120,000
Lead Streaming Data Engineer / Technical Lead
Lead Streaming Data Engineer / Technical Lead

Compunnel, Inc. • Tampa (FL)

On-site
USD 150,000 - 190,000
Senior Data Engineer
Senior Data Engineer

Compunnel, Inc. • Bentonville (AR)

On-site
USD 100,000 - 130,000
Scala Developer
Scala Developer

Compunnel, Inc. • Greenwood Village (CO)

On-site
USD 90,000 - 130,000