Principal Data Engineer

codametrix

United States

On-site

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

CodaMetrix is seeking a Principal Data Engineer to lead its Databricks-based data platform, driving scalable streaming and batch pipelines for 30+ healthcare customers. You will own architecture, governance, and cost optimization while enabling ML Engineering, Analytics, and Product teams.

Ideal candidates bring 8+ years in data engineering, deep Databricks experience, and strong Python/SQL skills. The role requires leadership, architecture reviews, and cross-team collaboration to raise data

Qualifications

  • 8+ years of data engineering with progressive responsibility.
  • 5+ years hands-on with Databricks platform (Unity Catalog, Delta Lake, Structured Streaming).
  • Expert in PySpark and Python; working knowledge of Scala.
  • Extensive Terraform IaC experience across multi-environment deployments.
  • Experience with Apache Kafka / AWS MSK and CI/CD pipelines (Jenkins, GitHub Actions).
  • Strong SQL skills and understanding of lakehouse patterns; HIPAA/SOC 2 awareness.

Responsibilities

  • Own the technical strategy and roadmap for the data platform and its integration with ML/Analytics.
  • Design, build, and maintain scalable streaming and batch data pipelines on Databricks.
  • Manage infrastructure with Terraform; implement CI/CD using Jenkins and GitHub Actions.
  • Enforce data governance, security, and cost optimization across environments.
  • Mentor engineers and collaborate with cross-functional teams to drive platform standards.

Skills

PySpark
Python
Scala
Terraform
Apache Kafka / AWS MSK
Jenkins
GitHub Actions
SQL

Education

Bachelor's degree in Computer Science or related field

Tools

Databricks platform
Unity Catalog
Delta Lake
Structured Streaming
Spark SQL
Terraform for IaC
Jenkins CI/CD
GitHub Actions
AWS (S3, IAM)

Job description

CodaMetrix is revolutionizing Revenue Cycle Management with its AI-powered autonomous coding solution, a multi-specialty AI-platform that translates clinical information into accurate sets of medical codes. CodaMetrix's autonomous coding drives efficiency under fee-for-service and value-based care models and supports improved patient care. We are passionate about getting physicians and healthcare providers away from the keyboard and back to clinical care.

Overview

The Principal Data Engineer is a member of the Data Platform team, reporting to the Director of Machine Learning Engineering and Data. The Data Platform team is responsible for executing the data strategy for the organization, ensuring high-quality external data is ingested into the Lakehouse and realized in powerful insights for internal and external customers, while ensuring ML/AI, BI and customer success teams have the data they need to develop and train their models, build insightful dashboards and design semantic layer. As a Principal Data Engineer (L4), you will serve as a key technical leader for CodaMetrix's Databricks-based data platform, supporting streaming and batch workloads across 30+ healthcare customers. You will own the platform's architecture, evolution, and operational excellence, including infrastructure-as-code, CI/CD automation, disaster recovery, cost optimization, data contracts, and build-vs-buy decisions, while enabling ML Engineering, Analytics, DevOps, and Product teams and influencing technical standards beyond the immediate team. This role operates with minimal oversight and is expected to influence technical standards beyond the immediate team.

Responsibilities

Own the technical strategy and roadmap - Own the vision, architecture, and roadmap for the CodaMetrix data platform, ensuring scalability, reliability, regulatory alignment, and operational excellence. Lead key technology decisions, disaster recovery design, architecture reviews, and data engineering standards across teams.

Platform Engineering at Scale - Design, build, and maintain scalable streaming and batch data platforms Databricks using PySpark, Unity Catalog, Delta Lake, and Structured Streaming. Own Terraform-based infrastructure across environments, including jobs, catalogs, schemas, permissions, compute policies, volumes, and external locations. Build Jenkins CI/CD workflows for automated testing, tagging, and deployments, while optimizing training pipelines, materialized views, retention policies, and production performance.

Data Governance & Security - Implement and evolve least-privilege access controls across Databricks using Unity Catalog grants, YAML-driven group policies, and schema-level restrictions. Ensure HIPAA and SOC 2 compliance through PHI masking, audit logging, environment-level data segmentation, user provisioning, compute policies, and cost attribution.

Cost Optimization & Operational Excellence - Drive platform cost reduction through compute policy tuning, serverless optimization, reserved pools, materialized view improvements, and remediation of underutilized resources. Monitor Databricks/AWS spend (using tools like CloudZero and AWS Cost Explorer), own cost attribution and budgeting, resolve production incidents, and maintain runbooks to ensure platform SLAs for uptime and performance.

Cross-Functional Enablement & Mentorship - Enable ML, BI, Analytics, DevOps, and customer success and implementation teams with training data pipelines, model-ready datasets, feature store architecture, optimized views, dashboards, Tableau refreshes, infrastructure changes, and tenant onboarding. Mentor engineers through code and design reviews, establish quality standards, and serve as a subject matter expert for data platform engineering across the organization.

Requirements
Must-Haves
  • A degree in Computer Science or a related field (Bachelor's, Master's, or Ph.D.), or an equivalent combination of education and demonstrable professional experience
  • 8+ years of data engineering experience with progressive responsibility
  • 5+ years hands‑on experience with the Databricks platform (Unity Catalog, Delta Lake, Structured Streaming, Spark SQL, cluster management, platform administration)
  • Expert proficiency in PySpark and Python; working knowledge of Scala
  • Expert experience with Terraform for infrastructure-as-code (state management, modular structures, multi-environment deployments, YAML-driven configuration)
  • Expert experience with Apache Kafka / AWS MSK (streaming ingestion, SASL/IAM auth, topic management, cluster migrations)
  • Deep understanding of medallion/lakehouse architecture patterns (bronze/silver/gold, SCD2, materialized views, slowly changing dimensions)
  • Proven track record building and maintaining CI/CD pipelines (Jenkins, GitHub Actions) for data platform deployments
  • Expert SQL skills - complex CTEs, window functions, performance optimization of large‑scale queries across petabyte-scale datasets
  • Strong AWS experience (S3, IAM,
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Data Platform Engineer — Databricks & Data Lakehouse
Lead Data Platform Engineer — Databricks & Data Lakehouse

PVH (Tommy Hilfiger/Calvin Klein) • United States

On-site
USD 140,000 - 210,000
Data Architect
Data Architect

Pho Prime, LLC • Shelton (CT)

On-site
USD 120,000 - 190,000
Mobility Allowance
Databricks Engineer
Databricks Engineer

CMT Services, Inc. • Adelphi (MD)

On-site
USD 100,000 - 130,000
Principal Data Engineer
Principal Data Engineer

Medical Guardian • Pennsylvania

On-site
USD 150,000 - 210,000
Health Plan
Paid Time Off
Short Term Disability & Life Insurance
+1
Principal Data Engineer
Principal Data Engineer

Medical Guardian LLC • Philadelphia

On-site
USD 150,000 - 210,000
Health Care Plan (Medical, Dental &amp
Paid Time Off (Vacation, Sick Time Off
Company Paid Short Term Disability and
+1
Databricks Data Platform Architect
Databricks Data Platform Architect

UsefulBI Corporation • Raleigh (NC)

On-site
USD 130,000 - 180,000
Data Engineer, Principal
Data Engineer, Principal

United States Digital Space LLC • United States

Hybrid
Data Solution Architect
Data Solution Architect

Falcon Smart IT (FalconSmartIT) • San Mateo (CA)

On-site
USD 180,000 - 240,000
Data Engineer, Principal
Data Engineer, Principal

Blue Shield of CA • Milwaukee (WI)

On-site
USD 170,000 - 230,000
Sr. Data Engineer, Data Platform
Sr. Data Engineer, Data Platform

Mirion Technologies • United States

On-site
USD 120,000 - 160,000