Employment Type
Full-Time
Clearance Requirement
Active DOE Q Clearance (or Ability to Obtain)
Position Overview
TradeWind is seeking an experienced Senior Databricks Engineer to augment the Client's enterprise data engineering team during the implementation of a new ERP platform.
This is an active, hands-on engineering position where the candidate will directly develop code, build and optimize data pipelines, and work agilely across a variety of day-to-day engineering tasks to support technical execution within the team. The candidate will jump in alongside an existing technical contractor team and internal staff to develop, refine, and operationalize five inaugural domain lakehouse pipelines.
As a collaborative and adaptable technical contributor, the candidate will lead development across the Silver through Platinum medallion layers, implement dimensional data models, establish automated testing and quality assurance frameworks, maintain technical documentation (e.g., Architecture Decision Records [ADRS] in GitHub), and mentor incoming staff on daily operations.
The Statement of Work may evolve as project requirements progress. The candidate will work under the technical and operational direction of the Senior Manager, Data & Analytics and the Data Architecture Workstream Lead to maintain alignment with critical ERP project milestones and enterprise architecture standards.
Key Responsibilities
- The candidate will deliver hands-on technical execution across five core areas:
- Medallion Pipeline Development & Ingestion/Distribution
- Hands-on development, tune, and operationalize high-throughput batch and streaming ETL/ELT pipelines in PySpark and SQL across the Medallion Architecture (Bronze -Silver -Gold/Platinum).
- Ingest complex, enterprise transactional data from the ERP platform and a legacy data warehouse into Azure Data Lake Storage (ADLS Gen2) and Delta Lake.
- Implement governed data distribution patterns optimized for downstream applications, analytics, reporting layers, and Power BI semantic models.
- Dimensional Data Modeling (Gold/Platinum Layers)
- Lead the implementation of enterprise dimensional models, star schemas, slowly changing dimensions (SCDs), and curated analytical aggregates.
- Collaborate with data architects and business analysts to translate functional data mappings into performant, query-optimized physical data structures.
- Pipeline Quality Assurance, Testing, & Observability
- Design and implement automated data quality validations, schema enforcement, and end-to-end integration tests across development, test, and production environments.
- Establish automated monitoring, alerting, error-handling routines, and pipeline telemetry to guarantee data accuracy, operational resilience, and SLA compliance.
- Operationalization, CI/CD, & Security
- Operationalize and optimize Databricks Workflows and Jobs for cost efficiency, scalability, and enterprise fault tolerance.
- Standardize deployment workflows using Git/GitHub CI/CD patterns (such as Databricks Asset Bundles [DAB]).
- Apply fine-grained access controls, object governance, and lineage tracking leveraging Databricks Unity Catalog.
- Documentation, Architecture Records, & Team Enablement
- Document pipeline designs, runbooks, and Architecture Decision Records (ADRs) directly within GitHub.
- Mentor and train incoming Databricks Engineers to transfer platform domain knowledge and ensure smooth handover of day-to-day lakehouse operations.
Required Qualifications
- S. Citizenship: The candidate must be strictly a United States Citizen.
- Information Security: The candidate must execute agreements to safeguard sensitive data in accordance with lab requirements and federal stipulations.
- Experience: 7+ years of data engineering/platform engineering experience, with 3-5+ years focused on production cloud data architectures.
- Databricks Depth: 5+ years of production experience with Azure Databricks, Delta Lake, PySpark, Spark SQL, Workflows, and Unity Catalog.
- Hands-on Execution: Proven ability to write clean, modular, and maintainable production code in Python/PySpark and SQL while agilely adapting to diverse technical tasks.
- Data Modeling: Demonstrated expertise in dimensional data modeling (Kimball methodology), star schemas, fact/dimension structures, and Platinum layer curation.
- Transactional System Integration: Proven background ingesting, transforming, and modeling large-scale, complex transactional or ERP data structures.
- DevOps & QA: Proficient in CI/CD pipelines (e.g., Databricks Asset Bundles, GitHub Actions), automated testing frameworks, and managing ADR documentation.
- Teamwork & Mentorship: Exceptional interpersonal skills, an amenable and collaborative team-first mindset, and a proven ability to mentor and upskill fellow engineers.
Preferred Qualifications
- Legacy Data Warehousing & SQL: Demonstrated familiarity or hands-on experience with legacy data warehouses, SQL Server databases, data marts, and migrating legacy SQL/ETL workloads to Databricks.
- Enterprise & Regulated Environments: Prior experience delivering data pipelines in highly regulated or security-conscious environments.
- GenAI Acceleration: Experience leveraging GenAI/LLM developer tools (e.g., GitHub Copilot, Databricks Assistant) to accelerate development and test creation.
- Certifications: Databricks Certified Data Engineer Associate/Professional or Microsoft Certified: Azure Data Engineer Associate.
Salary and Benefits Information
Expected Salary Range: $125k - $200k / year (depending on experience)
TradeWind Services offers a selection of exceptional benefits to employees and their families. We can customize your benefits to fit your specific needs. Benefits include:
- Long- and Short-term Disability Insurance