Java Spark Engineer

Veriipro

Berkeley Heights (NJ)

On-site

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Veriipro in Berkeley Heights, NJ, seeks a senior data engineer to architect and implement scalable data pipelines using Spark (Java). You will lead batch and streaming ETL/ELT, optimize performance, and set coding standards across the team.

You will mentor engineers, drive data architecture decisions, partner with product and analytics, own production reliability, and help control cost and capacity for cluster infrastructure.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.
  • 7+ years of professional Java development experience.
  • 5+ years hands-on experience with Apache Spark in production environments.
  • Expert-level understanding of distributed systems: fault tolerance, data locality, shuffle mechanics, resource management.
  • Proven track record designing systems processing terabyte+ scale data.
  • Strong SQL skills and deep familiarity with columnar storage formats (Parquet, ORC, Avro, Delta Lake/Iceberg).
  • Experience with cluster managers (YARN, Kubernetes) and cloud-managed Spark.
  • Proficiency with Kafka
  • Strong grasp of CI/CD, containerization, and infrastructure-as-code practices.

Responsibilities

  • Architect and build scalable, fault-tolerant data pipelines using Apache Spark (Java)
  • Lead design of batch and streaming ETL/ELT systems handling large data volumes
  • Deep-dive performance tuning: partitioning strategy, memory management, shuffle/skew optimization, job cost reduction
  • Set coding standards and lead code/design reviews across the team
  • Drive technical decisions on data architecture, storage formats, and pipeline orchestration
  • Mentor mid-level and junior engineers; act as a technical escalation point
  • Partner with product, analytics, and platform teams to translate requirements into scalable systems
  • Own production reliability — on-call ownership, incident response, root‑cause analysis for pipeline failures
  • Evaluate and introduce new tools/frameworks where they improve the system
  • Contribute to capacity planning and cost optimization for cluster infrastructure

Skills

Java
Spark
Distributed systems
SQL
Kafka
CI/CD
Containers
Kubernetes
YARN
Delta Lake

Education

Bachelor’s or Master’s degree in CS/Engineering

Tools

YARN
Kubernetes
Docker

Job description

Primary Responsibilities
  • Architect and build scalable, fault-tolerant data pipelines using Apache Spark (Java)
  • Lead design of batch and streaming ETL/ELT systems handling large data volumes
  • Deep-dive performance tuning: partitioning strategy, memory management, shuffle/skew optimization, job cost reduction
  • Set coding standards and lead code/design reviews across the team
  • Drive technical decisions on data architecture, storage formats, and pipeline orchestration
  • Mentor mid-level and junior engineers; act as a technical escalation point
  • Partner with product, analytics, and platform teams to translate requirements into scalable systems
  • Own production reliability — on-call ownership, incident response, root‑cause analysis for pipeline failures
  • Evaluate and introduce new tools/frameworks where they improve the system
  • Contribute to capacity planning and cost optimization for cluster infrastructure
Required Qualifications
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or related field
  • 7+ years of professional Java development experience
  • 5+ years hands‑on experience with Apache Spark in production environments
  • Expert‑level understanding of distributed systems: fault tolerance, data locality, shuffle mechanics, resource management
  • Proven track record designing systems processing terabyte+ scale data
  • Strong SQL skills and deep familiarity with columnar storage formats (Parquet, ORC, Avro, Delta Lake/Iceberg)
  • Experience with cluster managers (YARN, Kubernetes) and cloud‑managed Spark
  • Proficiency with Kafka
  • Strong grasp of CI/CD, containerization, and infrastructure‑as‑code practices
Preferred Qualifications
  • Experience with Flink or other stream‑processing frameworks
  • Familiarity with data governance, lineage, and quality frameworks
  • Experience with workflow orchestration at scale
  • Background in system design for multi‑tenant or multi‑region data platforms
  • Prior experience leading a team or acting as a technical lead
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer
Senior Software Engineer

Infinite Computer Solutions • Town of Texas (WI)

On-site
USD 110,000 - 140,000
Senior Technical Lead
Senior Technical Lead

Infinite Computer Solutions • Town of Texas (WI)

On-site
USD 120,000 - 180,000
Data Engineer
Data Engineer

Infinite Computer Solutions • Town of Texas (WI)

On-site
USD 120,000 - 170,000
Spark Engineer
Spark Engineer

Veriipro • Jacksonville (FL)

On-site
USD 90,000 - 130,000
Senior Data Engineer
Senior Data Engineer

Compunnel, Inc. • Bentonville (AR)

On-site
USD 100,000 - 130,000
Senior Data Engineer
Senior Data Engineer

Compunnel, Inc. • Sunnyvale (CA)

On-site
USD 100,000 - 140,000
Data Engineer ETL
Data Engineer ETL

Compunnel, Inc. • Durham (NC)

On-site
USD 100,000 - 130,000
Senior Data Engineer – Enterprise Data Frameworks
Senior Data Engineer – Enterprise Data Frameworks

Citizens • Rhode Island

On-site
USD 120,000 - 160,000
PySpark / Java Developer
PySpark / Java Developer

Veriipro • Whitpain Township (PA)

On-site
USD 90,000 - 120,000
Pyspark Architect
Pyspark Architect

Avance Consulting • Charlotte (NC)

On-site
USD 100,000 - 130,000