Senior Data Software Engineer/ Databricks, Apache Spark, PySpark

EPAM Systems Inc

United States

Remote

USD 140,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

EPAM Systems is seeking a Senior Data Software Engineer to design and implement scalable data pipelines using PySpark and SQL within a Databricks-centric environment. You will write clean, well-tested code and collaborate with data scientists and engineers to deliver reliable data solutions.

Ideal candidates have 3+ years in data software engineering, strong PySpark/SQL capabilities, experience with Databricks, and a solid grounding in unit testing with pytest; English at B2+ and familiarity

Qualifications

  • 3+ years of experience in Data Software Engineering.
  • Proven experience with Apache Spark, ideally in a Databricks environment.
  • Proficiency in PySpark and SQL.
  • Background in unit testing frameworks, especially pytest.
  • Understanding of data engineering principles and ETL processes.
  • Familiarity with version control systems (e.g., Git).
  • Ability to work independently and in a collaborative team setting.
  • Excellent problem-solving and communication skills.
  • Proficiency in English at an Upper-Intermediate level (B2) or higher.

Responsibilities

  • Design, develop, and maintain scalable data pipelines using Apache Spark (Databricks preferred).
  • Write efficient and optimized PySpark code for data transformation and processing.
  • Develop and execute complex SQL queries for data extraction, validation, and reporting.
  • Implement unit tests using pytest to ensure code reliability and maintainability.
  • Collaborate with data scientists, analysts, and other engineers to deliver high-quality data solutions.
  • Monitor and troubleshoot data workflows and performance issues.
  • Document technical designs, processes, and best practices.

Skills

PySpark
SQL
Unit testing
Apache Spark
Databricks

Tools

Databricks
Git
Delta Lake
MLflow

Job description

We are seeking a skilled Senior Data Software Engineer with strong expertise in PySpark , SQL , and unit testing to join our data engineering team. The ideal candidate will have hands-on experience with Apache Spark , preferably within the Databricks environment, and will be responsible for building scalable data pipelines, optimizing data workflows, and ensuring code quality through rigorous testing practices.

Responsibilities
  • Design, develop, and maintain scalable data pipelines using Apache Spark (Databricks preferred)
  • Write efficient and optimized PySpark code for data transformation and processing
  • Develop and execute complex SQL queries for data extraction, validation, and reporting
  • Implement unit tests using pytest to ensure code reliability and maintainability
  • Collaborate with data scientists, analysts, and other engineers to deliver high-quality data solutions
  • Monitor and troubleshoot data workflows and performance issues
  • Document technical designs, processes, and best practices
Requirements
  • 3+ years of experience in Data Software Engineering
  • Proven experience with Apache Spark, ideally in a Databricks environment
  • Proficiency in PySpark and SQL
  • Background in unit testing frameworks, especially pytest
  • Understanding of data engineering principles and ETL processes
  • Familiarity with version control systems (e.g., Git)
  • Ability to work independently and in a collaborative team setting
  • Excellent problem-solving and communication skills
  • Proficiency in English at an Upper-Intermediate level (B2) or higher
  • Nice to have Experience with cloud platforms (e.g., Azure, AWS, GCP)
  • Knowledge of CI/CD pipelines and DevOps practices
  • Familiarity with Delta Lake, MLflow, or other Databricks-native tools
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

The Value Maximizer • United States

On-site
USD 90,000 - 120,000
Sr. Data Engineer
Sr. Data Engineer

Techgene Solutions • Pasadena (CA)

Hybrid
USD 120,000 - 180,000
Databricks Data Engineer
Databricks Data Engineer

Henderson Scott • Irving (TX)

On-site
USD 100,000 - 130,000
Data Tester
Data Tester

Compunnel, Inc. • Louisville (KY)

On-site
USD 90,000 - 110,000
Senior Data Software Engineer/ AI Agents, Azure, Spark
Senior Data Software Engineer/ AI Agents, Azure, Spark

EPAM Systems Inc • United States

Remote
USD 140,000 - 180,000
Senior Data Software Engineer with Databricks
Senior Data Software Engineer with Databricks

EPAM Systems Inc • United States

Remote
USD 120,000 - 190,000
Senior Data Engineer - Databricks
Senior Data Engineer - Databricks

DATAECONOMY Inc • New Jersey

On-site
USD 130,000 - 180,000
Lead Data Engineer (Databricks & Spark)
Lead Data Engineer (Databricks & Spark)

EPAM Systems Inc • United States

On-site
USD 140,000 - 190,000
Databricks developer
Databricks developer

DS Technologies Inc • Atlanta (GA)

On-site
USD 120,000 - 150,000
Lead Data Engineer with Databricks
Lead Data Engineer with Databricks

Univedge Consulting LLC • St. Louis (MO)

On-site
USD 120,000 - 180,000