Big Data Consultant

HMG AMERICA LLC

San Francisco (CA)

On-site

USD 120,000 - 150,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

A tech data solutions provider is seeking a Spark job migration specialist in San Francisco, California. This role focuses on migrating data pipelines, JAR tasks, and analytics workloads to modern platforms, requiring over 5 years of experience with Apache Spark and cloud services like Azure or AWS. Responsibilities include refactoring code, performance optimization, and regression testing. Ideal candidates have strong skills in HDFS and the Hadoop ecosystem, as well as expertise in SQL and scripting.

Qualifications

  • 5+ years experience with Apache Spark (PySpark/Scala) and Cloud platforms.
  • Strong experience with HDFS and Hadoop ecosystem.

Responsibilities

  • Migrate JVM workloads and Spark-Submit tasks to Databricks.
  • Convert HiveQL scripts and Oozie workflows into optimized Spark applications.
  • Implement Adaptive Query Execution in Spark 3 to improve performance.
  • Perform regression testing to validate output consistency.

Skills

Apache Spark (PySpark/Scala)
HDFS
Cloud platforms (Azure/AWS)
Data migration
SQL and performance tuning
Scripting (Python, Shell, Scala)

Job description

Overview

A Spark job migration specialist migrates data pipelines, JAR tasks, and analytics workloads from legacy systems (like Hadoop/CDH or AWS EMR) to ACOS modern platforms. This involves refactoring code (e.g., Hive to PySpark), performance testing, and updating Spark 2.x to 3.x.

About the Role

A Spark job migration specialist migrates data pipelines, JAR tasks, and analytics workloads from legacy systems (like Hadoop/CDH or AWS EMR) to ACOS modern platforms. This involves refactoring code (e.g., Hive to PySpark), performance testing, and updating Spark 2.x to 3.x.

Responsibilities
  • Workload Migration: Migrate JVM workloads and Spark-Submit tasks to Databricks JAR tasks or Notebook tasks.
  • Pipeline Re-engineering: Convert existing HiveQL scripts and Oozie workflows into optimized Spark SQL or PySpark applications.
  • Refactoring: Adapt data pipelines from Azure Synapse to any cloud platform, including updating library dependencies and notebook references.
  • Performance Optimization: Implement Adaptive Query Execution (AQE) in Spark 3 to improve shuffle performance and fix skew joins.
  • Testing & Validation: Perform regression testing to ensure output consistency between old and new systems using validation scripts.
  • Job Customization: Use spark.sparkContext.setJobDescription() to label, monitor, and troubleshoot specific Spark tasks in the UI.
Qualifications

Experience: 5+ years experience with Apache Spark (PySpark/Scala) and Cloud platforms (Azure/AWS).

Required Skills
  • Strong experience with HDFS, Hadoop ecosystem (Hive, Spark, HBase, MapReduce).
  • Experience in data migration to cloud / enterprise data platforms.
  • Knowledge of:
  • Cloud storage (ADLS, S3, Blob Storage)
  • Distributed processing frameworks
  • SQL and performance tuning expertise.
  • Experience in scripting (Python, Shell, Scala).
Preferred Skills
  • Data Pipelines: Ensuring schema evolution, data correctness, and testing with golden datasets.
  • Job Definitions: Reconfiguring job properties, cluster settings, and Spark configurations.
Pay range and compensation package

Pay range or salary or compensation details not provided.

Equal Opportunity Statement

We are committed to diversity and inclusivity.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

PySpark Data Engineer
PySpark Data Engineer

Tata Consultancy Services • Irving (TX)

On-site
USD 90,000 - 120,000
Discretionary Annual Incentive
Medical Coverage (Health, Dental & Vis
401K Plan
+1
Data Engineer
Data Engineer

Infinite Computer Solutions • Town of Texas (WI)

On-site
USD 120,000 - 170,000
Data Engineer
Data Engineer

VDart Inc • Seattle (WA)

On-site
USD 96,000 - 152,000
Senior PySpark Data Developer
Senior PySpark Data Developer

Tata Consultancy Services • Irving (TX)

On-site
USD 90,000 - 150,000
Discretionary Annual Incentive
Comprehensive Medical Coverage
Parental Leaves
+4
Data Engineer
Data Engineer

Tata Consultancy Services • Irving (TX)

On-site
USD 125,000 - 140,000
Data Engineer
Data Engineer

JPS Tech Solutions • Seattle (WA)

On-site
USD 170,000 - 210,000
Spark Migration Architect for Big Data Pipelines
Spark Migration Architect for Big Data Pipelines

HMG AMERICA LLC • San Francisco (CA)

On-site
USD 120,000 - 150,000
Lead Big Data Engineer
Lead Big Data Engineer

SoftServe • Town of Poland (NY)

On-site
USD 120,000 - 190,000
Big Data Platform Engineer
Big Data Platform Engineer

Compunnel, Inc. • Rockville (MD)

On-site
USD 140,000 - 190,000
Senior Data Engineer (On Site, Washington, DC)
Senior Data Engineer (On Site, Washington, DC)

Agile5 Technologies, Inc. • Northern (KY)

Hybrid
USD 90,000 - 165,000