Java Spark Engineer

ApTask

New York (NY)

On-site

USD 150,000 - 210,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

ApTask is seeking a Java Spark Engineer to architect and build scalable data pipelines using Spark (Java) in a 5-days-in-office setting in Berkley Heights, NJ. You will lead design reviews, mentor engineers, and own production reliability across batch and streaming workloads.

The role demands 7+ years of Java development and 5+ years Spark in production, deep distributed-systems expertise, and strong SQL with Parquet/ORC/Avro/Delta Lake.

Qualifications

  • 7+ years of professional Java development experience.
  • 5+ years with Apache Spark in production.
  • Strong understanding of distributed systems and data locality.

Responsibilities

  • Architect and build scalable data pipelines with Spark (Java).
  • Lead batch and streaming ETL/ELT design.
  • Tune performance: partitioning, memory, shuffle, cost.
  • Set coding standards and review code/design.
  • Drive data-architecture decisions and tooling.
  • Mentor engineers and serve as escalation point.
  • Collaborate with product/analytics/platform teams.
  • Ensure production reliability and on-call incident response.
  • Evaluate new tools to improve systems.
  • Contribute to capacity planning and cost optimization.

Skills

Java
Apache Spark
Distributed systems
SQL
Kafka
Kubernetes
CI/CD
Parquet
Delta Lake

Education

Bachelor's or Master’s in CS/Engineering

Tools

YARN
Kubernetes
EMR
Docker

Job description

Role: Java Spark Engineer

Location: Berkley Heights, NJ

Work Mode: 5-days in office (flexible to support weekend)

Responsibilities
  • Architect and build scalable, fault-tolerant data pipelines using Apache Spark (Java)
  • Lead design of batch and streaming ETL/ELT systems handling large data volumes
  • Deep-dive performance tuning: partitioning strategy, memory management, shuffle/skew optimization, job cost reduction
  • Set coding standards and lead code/design reviews across the team
  • Drive technical decisions on data architecture, storage formats, and pipeline orchestration
  • Mentor mid-level and junior engineers; act as a technical escalation point
  • Partner with product, analytics, and platform teams to translate requirements into scalable systems
  • Own production reliability — on-call ownership, incident response, root-cause analysis for pipeline failures
  • Evaluate and introduce new tools/frameworks where they improve the system
  • Contribute to capacity planning and cost optimization for cluster infrastructure
Required Qualifications
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or related field
  • 7 years of professional Java development experience
  • 5 years hands-on experience with Apache Spark in production environments
  • Expert-level understanding of distributed systems: fault tolerance, data locality, shuffle mechanics, resource management
  • Proven track record designing systems processing terabyte scale data
  • Strong SQL skills and deep familiarity with columnar storage formats (Parquet, ORC, Avro, Delta Lake/Iceberg)
  • Experience with cluster managers (YARN, Kubernetes) and cloud-managed Spark
  • Proficiency with Kafka
  • Strong grasp of CI/CD, containerization, and infrastructure-as-code practices
Preferred Qualifications
  • Experience with Flink or other stream-processing frameworks
  • Familiarity with data governance, lineage, and quality frameworks
  • Experience with workflow orchestration at scale
  • Background in system design for multi-tenant or multi-region data platforms
  • Prior experience leading a team or acting as a technical lead
Soft Skills / Leadership
  • Excellent communication — able to explain technical tradeoffs to non-technical stakeholders
  • Strong mentorship and coaching ability
  • Comfortable driving ambiguous, cross-team technical initiatives
Behavioral Skills
  • Good Communication skills
  • 5 days Work from Office at Berkley Heights, NJ
  • Team Player
  • Ability to work in a changing environment
  • Strong problem solving and analytical skills
  • Ability to work independently or within a team
  • Manage day-to-day challenges and communicate developmental risks with the technical team
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Java Spark Engineer
Java Spark Engineer

Saransh Inc • Berkley (AL)

On-site
USD 150,000 - 210,000
Java Spark Engineer
Java Spark Engineer

Delta System & Software, Inc. • Berkeley Heights (NJ)

On-site
USD 150,000 - 210,000
Java Spark Developer
Java Spark Developer

SmartRecruiters, Inc. • New York (NY)

On-site
USD 140,000 - 210,000
Senior Java Spark Architect – Scalable Data Pipelines
Senior Java Spark Architect – Scalable Data Pipelines

ApTask • New York (NY)

On-site
USD 150,000 - 210,000
Senior Data Engineer – Enterprise Data Frameworks
Senior Data Engineer – Enterprise Data Frameworks

Citizens • Rhode Island

On-site
USD 120,000 - 160,000
Senior Java Spark Engineer Onsite: Data Pipeline Architect
Senior Java Spark Engineer Onsite: Data Pipeline Architect

Saransh Inc • Berkley (AL)

On-site
USD 150,000 - 210,000
Apache Spark Java Developer
Apache Spark Java Developer

Elite Workforce Inc. • Brooklyn Park (MN)

On-site
USD 80,000 - 110,000
Data Engineer ETL
Data Engineer ETL

Compunnel, Inc. • Durham (NC)

On-site
USD 100,000 - 130,000
Senior Software Engineer - Core Java & Apache Spark
Senior Software Engineer - Core Java & Apache Spark

Citi • New York (NY)

On-site
USD 120,000 - 160,000
Senior Java Spark Architect & Data Pipeline Tech Lead
Senior Java Spark Architect & Data Pipeline Tech Lead

Mphasis • New York (NY)

On-site
USD 63,000 - 135,000