Apache Spark Developer

Bright Vision Technologies

Austin (TX)

Remote

USD 125,000 - 185,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Bright Vision Technologies seeks an experienced Apache Spark Developer to design, develop, and optimize large-scale distributed data processing applications supporting analytics, ML, real-time reporting, and cloud platforms. The role emphasizes building high-performance Spark pipelines across structured and semi-structured data, delivering scalable, cost-efficient solutions.

You will collaborate with data architects, cloud engineers, ML engineers, and BI teams to leverage Spark, Databricks, and

Qualifications

  • 6+ years of professional software or data engineering experience.
  • 4+ years of hands-on Apache Spark development in production.
  • Strong proficiency in PySpark, Scala, or Spark SQL for distributed processing.
  • Deep understanding of Spark architecture (RDDs, DataFrames, Datasets, Catalyst, DAG, Tungsten).
  • Experience with distributed computing concepts (partitioning, shuffling, caching, broadcasts, fault tolerance).
  • Advanced SQL skills with SQL Server, Oracle, PostgreSQL, Snowflake, or Teradata.

Responsibilities

  • Design, develop, and maintain high-performance distributed data-processing applications using Spark.
  • Build scalable batch and real-time ETL/ELT pipelines for enterprise data.
  • Develop Spark apps using PySpark, Scala, or Spark SQL for data transformation and analytics.
  • Optimize Spark jobs for memory, partitioning, and shuffle performance.
  • Ingest data from enterprise sources, APIs, Kafka, and data lakes.
  • Collaborate with cloud teams to deploy Spark on Databricks, EMR, Azure Synapse, or Kubernetes.
  • Implement data quality validation, monitoring, and error handling across pipelines.
  • Integrate Spark with data warehouses, lakehouses, and BI tools.
  • Participate in architecture and code reviews and Agile ceremonies.
  • Troubleshoot production issues and optimize performance in distributed setups.
  • Support cloud migration by modernizing legacy ETL into Spark-based architectures.

Skills

PySpark
Scala
Spark SQL
Spark architecture
Distributed systems
SQL proficiency

Tools

Databricks
Kafka
Hadoop ecosystem
Delta Lake
Kubernetes

Job description

Apache Spark Developer – Remote

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.

Job Title: Apache Spark Developer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $125,000–$185,000 Annually
Experience Required: 6+ years

Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.

Job Summary

We are seeking an experienced Apache Spark Developer to design, develop, and optimize large-scale distributed data processing applications supporting enterprise analytics, machine learning, real-time reporting, and cloud-based data platforms. This role focuses on building high-performance Spark applications capable of processing billions of records across structured and semi-structured data sources while delivering scalable, reliable, and cost-efficient data pipelines.

You will work closely with data architects, data engineers, cloud platform teams, machine learning engineers, and business intelligence developers to build modern data processing solutions leveraging Apache Spark, cloud-native technologies, and distributed computing frameworks. The ideal candidate possesses deep expertise in Spark architecture, distributed systems, performance optimization, and cloud-based big data ecosystems.

Key Responsibilities
  • Design, develop, and maintain high-performance distributed data processing applications using Apache Spark.
  • Build scalable batch and real-time ETL/ELT pipelines processing large volumes of enterprise data.
  • Develop Spark applications using PySpark, Scala, or Spark SQL for data transformation, aggregation, and analytics.
  • Optimize Spark jobs for memory utilization, partitioning strategies, shuffle performance, and execution efficiency.
  • Process structured, semi-structured, and streaming data from enterprise databases, APIs, Kafka, cloud storage, and data lakes.
  • Develop reusable Spark libraries, data processing frameworks, and metadata-driven ingestion pipelines.
  • Collaborate with cloud engineering teams to deploy Spark workloads on Databricks, EMR, Azure Synapse, or Kubernetes.
  • Implement data quality validation, reconciliation, monitoring, and automated error handling across distributed pipelines.
  • Integrate Spark applications with enterprise data warehouses, lakehouses, and reporting platforms.
  • Participate in architecture reviews, code reviews, technical design discussions, and Agile development activities.
  • Troubleshoot production issues involving distributed processing, cluster performance, resource utilization, and data quality.
  • Support cloud migration initiatives by modernizing legacy ETL workloads into Spark-based architectures.
Required Skills
  • Six or more years of professional software or data engineering experience.
  • Four or more years of hands-on Apache Spark development experience in enterprise production environments.
  • Strong proficiency in PySpark, Scala, or Spark SQL for distributed data processing.
  • Deep understanding of Apache Spark architecture including RDDs, DataFrames, Datasets, Catalyst Optimizer, DAG execution, and Tungsten engine.
  • Strong experience with distributed computing concepts including partitioning, shuffling, caching, broadcast joins, and fault tolerance.
  • Advanced SQL skills with databases such as SQL Server, Oracle, PostgreSQL, Snowflake, or Teradata.
  • Experience working with Hadoop ecosystem technologies including Hive, HDFS, YARN, and Parquet.
  • Experience processing streaming data using Spark Structured Streaming, Apache Kafka, or Event Hubs.
  • Hands-on experience with cloud platforms including Azure Databricks, AWS EMR, AWS Glue, Azure Synapse Analytics, or Google Dataproc.
  • Experience integrating Spark applications with Delta Lake, Apache Iceberg, or Apache Hudi.
  • Strong understanding of data warehousing concepts, dimensional modeling, and data lake architecture.
  • Experience using Git, CI/CD pipelines, Azure DevOps, GitHub Actions, or Jenkins.
  • Strong debugging, troubleshooting, and Spark performance tuning skills.
  • Experience working in Agile Scrum development environments.
Preferred Qualifications
  • Experience building enterprise Lakehouse architectures using Databricks or Delta Lake.
  • Familiarity with Apache Airflow, Azure Data Factory, AWS Step Functions, or Control-M for workflow orchestration.
  • Experience with machine learning workflows using Spark MLlib, MLflow, or feature engineering pipelines.
  • Knowledge of Kubernetes, Docker, and containerized Spark deployments.
  • Experience implementing Data Quality frameworks using Great Expectations or Deequ.
  • Familiarity with Apache NiFi, Apache Flink, Trino, or Presto.
  • Experience working with cloud object storage including Amazon S3, Azure Data Lake Storage (ADLS Gen2), or Google Cloud Storage.
  • Knowledge of Infrastructure as Code using Terraform or ARM templates.
  • Experience with enterprise monitoring tools including Prometheus, Grafana, Datadog, or OpenTelemetry.
  • Cloud certifications in Azure, AWS, Databricks, or Apache Spark-related technologies are highly desirable.
Project Environment

You will be joining a modern data engineering team responsible for building cloud-native big data platforms supporting enterprise analytics, AI, and business intelligence initiatives. Current projects include:

  • Enterprise data lakehouse implementation using Databricks and Delta Lake
  • Real-time streaming analytics processing billions of daily events
  • Large-scale customer analytics and behavioral data platforms
  • Financial risk modeling and fraud detection pipelines
  • Healthcare clinical and operational analytics solutions
  • Cloud migration of legacy Hadoop and ETL workloads
  • Machine learning feature engineering and model training pipelines
  • Enterprise reporting platforms supporting executive dashboards and self-service analytics
  • Distributed data processing infrastructure deployed on Azure and AWS

This is a hands-on engineering role where you will contribute to distributed system architecture, Spark application development, cloud migration, performance optimization, production support, and continuous improvement of enterprise-scale data processing platforms.

Equal Employment Opportunity (EEO) Statement

Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall.

BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Technical Lead AWS PySpark
Technical Lead AWS PySpark

Argyllinfotech • South San Francisco (CA)

Hybrid
USD 150,000 - 190,000
Health insurance
401(k) plan
PySpark Data Engineer
PySpark Data Engineer

Tata Consultancy Services • Irving (TX)

On-site
USD 90,000 - 120,000
Discretionary Annual Incentive
Medical Coverage (Health, Dental & Vis
401K Plan
+1
Big Data Architect
Big Data Architect

Bright Vision Technologies • Columbus (OH), Dublin (OH)

Remote
USD 100,000 - 150,000
ETL Developer
ETL Developer

Bright Vision Technologies • Austin (TX)

Remote
USD 105,000 - 175,000
Senior Data Engineer - Apache Spark and SQL - Vice President
Senior Data Engineer - Apache Spark and SQL - Vice President

Citi • New York (NY)

Hybrid
USD 26,000 - 48,000
AI Data Platform Engineer
AI Data Platform Engineer

Bright Vision Technologies • Cedar Park (TX)

Remote
USD 135,000 - 170,000
Sr. DATA Engineer
Sr. DATA Engineer

LTM • Irving (TX)

On-site
USD 100,000 - 130,000
Comprehensive Medical Plan Covering Medical, Dental, Vision
Short Term and Long-Term Disability Coverage
401(k) Plan with Company match
+2
Senior Apache Spark Engineer - Remote Work | REF#303632
Senior Apache Spark Engineer - Remote Work | REF#303632

BairesDev • United States

Remote
CLP 112,888,000 - 150,517,000
Remote work
USD salary option
Hardware provided
+3
Data Engineering Specialist – AI
Data Engineering Specialist – AI

Bright Vision Technologies • Columbus (OH), Dublin (OH)

Remote
USD 100,000 - 150,000
Senior Data Scientist
Senior Data Scientist

Bright Vision Technologies • Austin (TX)

Remote
USD 160,000 - 180,000