GDIT is seeking a Databricks Engineer to help build and operate modern data solutions supporting the NIA Data Enclave. In this role, you will design, develop, and optimize scalable data pipelines that enable researchers and analysts to securely work with sensitive federal and non-federal datasets.
You will work at the intersection of data engineering, cloud technology, and mission-focused analytics, partnering with data scientists, researchers, software engineers, architects, and platform teams to turn complex data into reliable, governed, and accessible analytical resources.
This is an opportunity to apply your Databricks and Spark expertise to a mission where data quality, security, scalability, and reliability matter.
How You'll Make an Impact
As a Databricks Engineer, you will:
- Design, build, and optimize scalable ETL/ELT pipelines using Databricks, PySpark, SQL, and Delta Lake.
- Develop and maintain Databricks notebooks, jobs, and workflows that support high-volume analytical workloads.
- Build reliable data ingestion, transformation, validation, and integration processes.
- Help migrate and modernize existing data workloads for improved scalability, performance, and maintainability.
- Optimize Spark workloads through effective partitioning, caching, joins, file management, and other performance-tuning techniques.
- Implement automated testing, data quality checks, monitoring, logging, and operational processes.
- Support CI/CD and infrastructure automation for Databricks workloads using Git and tools such as Azure DevOps or GitHub Actions.
- Configure and optimize Databricks compute, clusters, runtimes, and job execution.
- Work within secure, role-based cloud environments and help implement appropriate data governance and access controls.
- Troubleshoot production issues, perform root-cause analysis, and continuously improve the reliability of data services.
- Collaborate with data scientists, researchers, analysts, and other engineers to deliver high-quality analytical datasets.
What You'll Need to Succeed
- Bachelor's degree in computer science, software engineering, data engineering, or a related technical field.
- 5+ years of data engineering experience, including significant hands-on experience with Databricks.
- Strong experience with PySpark, SQL, Apache Spark, Delta Lake, and Databricks.
- Experience developing production-grade data pipelines and workflows.
- Experience working with cloud-based data platforms and storage such as AWS S3, Azure Data Lake Storage, or Google Cloud Storage.
- Experience with Git and CI/CD practices for deploying and managing data engineering workloads.
- Understanding of distributed data processing, data modeling, data quality, and pipeline performance optimization.
- Experience troubleshooting and supporting production data workloads.
- Understanding of cloud security concepts such as role-based access control, identity management, least-privilege access, and data protection.
- Strong communication skills and the ability to collaborate effectively with technical and mission-focused stakeholders.
Preferred Qualifications
- Experience with Databricks Unity Catalog and enterprise data governance.
- Experience working in FISMA Moderate/High or other regulated environments.
- Experience with AWS, Azure, and/or GCP in a multi-cloud environment.
- Familiarity with CMS, federal, healthcare, biomedical, or other sensitive datasets.
- Experience with secure data enclaves, restricted-access environments, or federated data platforms.
- Experience with Terraform or other infrastructure-as-code technologies.
- Experience with Databricks Lakeflow Declarative Pipelines / Delta Live Tables.
- Experience implementing data lineage, metadata management, monitoring, and audit-ready logging.
- Familiarity with Databricks APIs, SDKs, or automation frameworks.
Why GDIT?
At GDIT, you'll have the opportunity to apply modern data engineering technologies to meaningful mission challenges. You'll work with a collaborative team of engineers, architects, researchers, and data professionals while helping build secure and scalable capabilities for an environment where trustworthy data can directly enable better research and decision-making.
If you're a hands-on data engineer who enjoys solving complex data problems and wants to apply your Databricks expertise to an important mission, we'd like to hear from you.