Senior Serverless Spark Migration Engineer

Virtasant

Austin (TX)

Hybrid

USD 140,000 - 190,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Virtasant is seeking a Senior Platform Engineer to lead migrations of Apache Spark workloads from on-prem Hadoop/Spark environments to cloud-native and serverless architectures across AWS and GCP. This hands-on role covers assessment, architecture, refactoring, migration execution, performance optimization, automation, and production readiness.

You will determine target architectures per workload and establish repeatable patterns for enterprise-scale migrations, balancing legacy systems with

Qualifications

  • 8+ years in data engineering, distributed systems, cloud engineering, or platform engineering.
  • 5+ years of hands-on Apache Spark experience in enterprise environments.
  • Strong PySpark and/or Scala development experience.
  • Proven experience migrating large-scale Spark workloads between infrastructure platforms.
  • Hands-on experience with both AWS and GCP.
  • Experience with on-premise Hadoop/Spark ecosystems, including HDFS, YARN and Hive.

Responsibilities

  • Lead migrations of enterprise Spark workloads from on-premise environments to AWS and GCP.
  • Assess Spark applications, clusters, configurations, dependencies, data flows, and resource utilization.
  • Determine migration approaches: rehost, replatform, refactor, modernize, or retire.
  • Modernize traditional cluster workloads for serverless Spark where appropriate.
  • Design architectures using AWS EMR Serverless, S3, Glue, Lake Formation, GCP Dataproc Serverless, GCS, and BigQuery.
  • Refactor PySpark/Scala/Spark SQL applications for cloud portability and reliability.
  • Migrate workloads from Hadoop, HDFS, YARN, Hive, and on-prem Spark clusters.
  • Troubleshoot and optimize Spark workloads: partitioning, joins, data skew, execution plans, and SQL execution.
  • Benchmark performance and optimize serverless workloads for performance and cost.
  • Build reusable migration tooling, automation, and frameworks.
  • Implement CI/CD and Infrastructure as Code using Terraform and related tools.
  • Define testing, validation, cutover, rollback, observability, and production-readiness patterns.
  • Partner with Data Engineering, ML, Cloud Architecture, Platform Engineering, DevOps/SRE, Security, Governance, and FinOps teams.

Skills

PySpark
Scala
SQL
Cloud architecture

Tools

AWS
GCP
Hadoop
YARN
Hive
Terraform
Spark

Job description

Senior Platform Engineer, Cloud Infrastructure

Type: Remote (Brazil, Mexico)
Coverage: Pacific Hours (8:00 AM - 5:00 PM PST)

About Virtasant

Virtasant is a global cloud and technology services company helping organizations modernize, optimize, and build at scale. We work with enterprise customers on complex cloud, data, infrastructure, and AI initiatives, bringing together deep technical expertise and hands-on delivery.

About the Role

We’re looking for a Senior Serverless Spark Migration Engineer to help modernize a large-scale enterprise data platform.

You’ll lead the migration of production Apache Spark workloads from on-premise Hadoop/Spark environments to cloud-native and serverless architectures across AWS and GCP. This is a hands-on engineering role spanning workload assessment, architecture, application refactoring, migration execution, performance optimization, automation, and production readiness.

The goal is not simply to lift and shift existing workloads. You’ll determine the right target architecture for each workload and establish repeatable patterns that can eventually support migration at significant enterprise scale.

What You’ll Do
  • Lead migrations of enterprise Spark workloads from on-premise environments to AWS and GCP.

  • Assess Spark applications, clusters, configurations, dependencies, data flows, and resource utilization.

  • Determine the right migration approach across rehost, replatform, refactor, modernize, or retire.

  • Modernize traditional cluster-based workloads for serverless Spark where appropriate.

  • Design and implement architectures using technologies such as AWS EMR Serverless, S3, Glue, Lake Formation, GCP Dataproc Serverless, GCS, and BigQuery.

  • Refactor legacy PySpark/Scala/Spark SQL applications for cloud portability, scalability, and reliability.

  • Migrate workloads from environments using Hadoop, HDFS, YARN, Hive, and on-prem Spark clusters.

  • Troubleshoot and optimize Spark workloads across partitioning, shuffle behavior, joins, data skew, execution plans, executor configuration, serialization, and SQL execution.

  • Benchmark performance and optimize serverless workloads for performance, reliability, and cloud cost.

  • Build reusable migration tooling, automation, templates, and frameworks.

  • Implement CI/CD and Infrastructure as Code using tools such as Terraform.

  • Define testing, validation, cutover, rollback, observability, and production-readiness patterns.

  • Partner with Data Engineering, ML, Cloud Architecture, Platform Engineering, DevOps/SRE, Security, Governance, and FinOps teams.

What We’re Looking For
  • 8+ years of experience across data engineering, distributed systems, cloud engineering, or platform engineering.

  • 5+ years of hands-on Apache Spark experience in enterprise environments.

  • Strong PySpark and/or Scala development experience.

  • Proven experience migrating large-scale Spark workloads between infrastructure platforms.

  • Hands-on experience with both AWS and GCP.

  • Experience with on-premise Hadoop/Spark ecosystems, including technologies such as HDFS, YARN, and Hive.

  • Deep understanding of Spark internals and distributed processing.

  • Strong SQL and data engineering fundamentals.

  • Experience with cloud data lakes and object storage.

  • Strong production troubleshooting and performance-tuning experience.

  • Experience with CI/CD, Git, and Infrastructure as Code.

  • Ability to own migration work end-to-end, from discovery and architecture through production cutover and optimization.

Nice to Have

Experience with EMR/EMR Serverless, Dataproc/Dataproc Serverless, Glue, Lake Formation, BigQuery, Delta Lake, Iceberg, Kafka, Airflow, Terraform, Docker, or Kubernetes is valuable.

What Success Looks Like

You can take ownership of the complete migration lifecycle:

Discover -> Assess -> Design -> Refactor -> Migrate -> Validate -> Optimize -> Operate

You understand both the legacy Hadoop/Spark world and modern cloud-native data platforms, and can make pragmatic architecture decisions based on workload characteristics rather than simply reproducing an existing environment in the cloud.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Cloud-Native Spark Migration Architect (Serverless)
Cloud-Native Spark Migration Architect (Serverless)

Virtasant • Austin (TX)

Hybrid
USD 140,000 - 190,000
Remote - Senior PySpark Developer
Remote - Senior PySpark Developer

Resource Informatics Group, Inc • United States

On-site
USD 140,000 - 200,000
Big Data Consultant
Big Data Consultant

HMG AMERICA LLC • San Francisco (CA)

On-site
USD 120,000 - 150,000
Java Spark Engineer
Java Spark Engineer

Delta System & Software, Inc. • Berkeley Heights (NJ)

On-site
USD 150,000 - 210,000
Data Engineer
Data Engineer

Compunnel, Inc. • Dallas (TX)

On-site
USD 90,000 - 120,000
Sr AWS Migration Engineer
Sr AWS Migration Engineer

V R Della Infotech Inc • Minneapolis (MN)

On-site
USD 120,000 - 180,000
Data Engineer ETL
Data Engineer ETL

Compunnel, Inc. • Durham (NC)

On-site
USD 100,000 - 130,000
Pyspark Architect
Pyspark Architect

Avance Consulting • Charlotte (NC)

On-site
USD 100,000 - 130,000
Senior Data Engineer – Spark / Scala
Senior Data Engineer – Spark / Scala

Astyra Corporation • Richmond (VA)

On-site
USD 150,000 - 190,000
Senior Solutions Architect – Healthcare Data Migration
Senior Solutions Architect – Healthcare Data Migration

Jobtailor • Mesa (AZ)

On-site
USD 180,000 - 240,000