Databricks Engineer

Jobtailor

Gaithersburg (MD)

On-site

USD 120,000 - 180,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Jobtailor in Gaithersburg, MD seeks a data engineer to design, build, and optimize scalable data pipelines using Databricks, PySpark, SQL, and Delta Lake.

You will develop and maintain Databricks notebooks, jobs, and workflows for high-volume analytics, and collaborate with data scientists, analysts, engineers, and platform teams.

This role emphasizes data governance, cloud security, and CI/CD practices in a regulated environment.

Qualifications

  • Bachelor's degree in computer science, software engineering, data engineering, or a related technical field.
  • 5+ years of data engineering experience with significant Databricks exposure.
  • Strong experience with PySpark, SQL, Spark, Delta Lake, and Databricks.
  • Experience building production-grade data pipelines and workflows.
  • Experience with cloud data platforms and storage (AWS S3, ADLS, GCS).
  • Experience with Git and CI/CD practices.
  • Understanding of distributed data processing, data modeling, data quality, and pipeline performance optimization.
  • Experience troubleshooting production data workloads.
  • Understanding of cloud security concepts like RBAC and data protection.
  • Strong communication and collaboration skills.

Responsibilities

  • Design, build, and optimize scalable ETL/ELT pipelines with Databricks, PySpark, SQL, and Delta Lake
  • Maintain Databricks notebooks, jobs, and workflows for high-volume analytics
  • Build data ingestion, transformation, validation, and integration processes
  • Migrate and modernize existing data workloads
  • Optimize Spark workloads with partitioning, caching, and performance tuning
  • Implement automated testing, data quality checks, monitoring, and logging
  • Support CI/CD and infrastructure automation with Git, Azure DevOps, or GitHub Actions
  • Configure Databricks compute, clusters, runtimes, and job execution
  • Enforce RBAC, data governance, and access controls in cloud environments
  • Troubleshoot production data issues and improve reliability
  • Collaborate across data scientists, analysts, engineers, architects, and platform teams
  • Deliver governed datasets for sensitive federal and non-federal data

Skills

Databricks
PySpark
SQL
Apache Spark
Delta Lake
Data Pipeline Development
Data Quality Checks
Performance Tuning
Data Modeling
Infrastructure Automation

Education

Bachelor's degree in computer science, software engineering, data engineering, or a related technical field

Tools

Git
Azure DevOps
GitHub Actions
AWS S3
Azure Data Lake Storage
Google Cloud Storage
Terraform
Databricks Unity Catalog
Databricks APIs
Databricks Lakeflow

Job description


  • Design, build, and optimize scalable ETL/ELT pipelines using Databricks, PySpark, SQL, and Delta Lake

  • Develop and maintain Databricks notebooks, jobs, and workflows for high-volume analytical workloads

  • Build data ingestion, transformation, validation, and integration processes

  • Migrate and modernize existing data workloads

  • Optimize Spark workloads through partitioning, caching, joins, file management, and performance tuning

  • Implement automated testing, data quality checks, monitoring, logging, and operational processes

  • Support CI/CD and infrastructure automation using Git, Azure DevOps, or GitHub Actions

  • Configure and optimize Databricks compute, clusters, runtimes, and job execution

  • Work in secure, role-based cloud environments and implement data governance and access controls

  • Troubleshoot production issues, perform root-cause analysis, and improve data-service reliability

  • Collaborate with data scientists, researchers, analysts, engineers, architects, and platform teams

  • Deliver reliable, governed, accessible analytical datasets supporting the NIA Data Enclave and sensitive federal and non-federal datasets


Requirements


  • Bachelor's degree in computer science, software engineering, data engineering, or a related technical field

  • 5+ years of data engineering experience, including significant hands-on experience with Databricks

  • Strong experience with PySpark, SQL, Apache Spark, Delta Lake, and Databricks

  • Experience developing production-grade data pipelines and workflows

  • Experience with cloud-based data platforms and storage such as AWS S3, Azure Data Lake Storage, or Google Cloud Storage

  • Experience with Git and CI/CD practices

  • Understanding of distributed data processing, data modeling, data quality, and pipeline performance optimization

  • Experience troubleshooting and supporting production data workloads

  • Understanding of cloud security concepts including role-based access control, identity management, least-privilege access, and data protection

  • Strong communication and collaboration skills

  • 7+ years of related experience listed in job qualifications

  • Preferred: Experience with Databricks Unity Catalog and enterprise data governance

  • Preferred: Experience in FISMA Moderate/High or other regulated environments

  • Preferred: Experience with AWS, Azure, and/or GCP in a multi-cloud environment

  • Preferred: Familiarity with CMS, federal, healthcare, biomedical, or other sensitive datasets

  • Preferred: Experience with secure data enclaves, restricted-access environments, or federated data platforms

  • Preferred: Experience with Terraform or other infrastructure-as-code technologies

  • Preferred: Experience with Databricks Lakeflow Declarative Pipelines / Delta Live Tables

  • Preferred: Experience implementing data lineage, metadata management, monitoring, and audit-ready logging

  • Preferred: Familiarity with Databricks APIs, SDKs, or automation frameworks

  • US citizenship is not required

  • No clearance is required


Core Competencies

Demonstrates expertise in designing and optimizing ETL/ELT pipelines using Databricks, PySpark, and SQL, while ensuring data quality and governance in cloud environments. Proficient in implementing CI/CD practices and troubleshooting production data workloads to enhance reliability and performance.


Highest-signal resume keywords


  • Databricks ETL/ELT Pipeline Development

  • PySpark Programming

  • SQL Query Optimization

  • CI/CD Practices with Git

  • Cloud Data Platform Experience


ATS Optimization Keywords

Hard Skills


  • Databricks

  • PySpark

  • SQL

  • Apache Spark

  • Delta Lake

  • Data Pipeline Development

  • Data Quality Checks

  • Performance Tuning

  • Data Modeling

  • Infrastructure Automation


Soft Skills


  • Strong Communication

  • Collaboration Skills


Industry Keywords


  • Data Engineering

  • Cloud Security

  • Data Governance

  • FISMA Compliance

  • Healthcare Data

  • Biomedical Data

  • Sensitive Datasets

  • Role-Based Access Control

  • Data Lineage

  • Metadata Management


Tools & Technologies


  • Git

  • Azure DevOps

  • GitHub Actions

  • AWS S3

  • Azure Data Lake Storage

  • Google Cloud Storage

  • Terraform

  • Databricks Unity Catalog

  • Databricks APIs

  • Databricks Lakeflow

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

Jobtailor • Town of Montana (WI)

On-site
USD 120,000 - 180,000
Staff Data Engineer
Staff Data Engineer

Jobtailor • Missouri

On-site
USD 150,000 - 190,000
Databricks Administrator
Databricks Administrator

CloudIngest • Dallas (TX)

On-site
USD 110,000 - 170,000
Principal Data Engineer
Principal Data Engineer

Jobtailor • Vienna (VA)

On-site
USD 150,000 - 190,000
Cloud Engineer
Cloud Engineer

Jobtailor • Arlington (VA)

On-site
USD 140,000 - 190,000
Data Engineer
Data Engineer

Jobtailor • Vienna (VA)

On-site
USD 130,000 - 180,000
Databricks Engineer
Databricks Engineer

iLink Digital • Milpitas (CA), Northern (KY)

On-site
USD 120,000 - 160,000
Data Engineer I, AWS, Databricks
Data Engineer I, AWS, Databricks

Jobtailor • United States

On-site
USD 120,000 - 170,000
Staff Data Engineer
Staff Data Engineer

Jobtailor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Data Engineer
Senior Data Engineer

Jobtailor • Newport News (VA)

On-site
USD 120,000 - 180,000