Senior Databricks Engineer - Revenue Cycle Management (RCM)

Aarista Technologies

Houston (TX)

On-site

USD 120,000 - 190,000

Full time

9 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Aarista Technologies in Houston seeks an experienced Data Engineer to design and optimize production data pipelines on Azure Databricks, handling healthcare data with Delta Lake and time travel features.

You will build MongoDB sync pipelines for real-time access and develop Scala-based reporting engines, ensuring robust data quality and HIPAA-compliant processing across layers.

Qualifications

  • 6–7 years of data engineering experience with 3+ years on Databricks/Spark.
  • Proficient in PySpark (primary) and Scala/Spark; Python 3.11+.

Responsibilities

  • Design and maintain production data pipelines on Azure Databricks.
  • Implement Delta Lake with ACID, time travel, and partitioning.
  • Build MongoDB sync pipelines for real-time access.
  • Develop Scala-based reporting and orchestration engines.
  • Ensure HIPAA-compliant data processing across layers.

Skills

Databricks/Spark
PySpark
Scala/Spark
Python 3.11+
Data pipelines
Data quality frameworks

Tools

MongoDB
SQL Server
Azure Pipelines
CI/CD

Job description

Key Responsibilities:
  • Design & maintain production data pipelines using PySpark and Scala on Azure Databricks, processing healthcare EDI transactions (837 Claims, 835 Remittances, 277/999 Acknowledgments)
  • Implement and optimize Delta Lake tables with ACID transactions, Z-ordering, partitioning strategies, and data versioning (time travel) for petabyte-scale healthcare data
  • Build MongoDB sync pipelines to deliver Gold-layer data into operational MongoDB collections for real-time application access
  • Develop/Maintain Scala-based reporting and workflow orchestration engines, including dynamic report generation, AP analysis, and notification triggers
  • Integrate with Azure ecosystem: Azure Blob Storage, Azure Data Lake, Azure AD, Azure DevOps CI/CD pipelines, and Databricks Unity Catalog
  • Manage multi-tenant data architectures with per-client configurations, secrets management (Databricks secrets scopes), and feature flag-driven rollouts
  • Build and maintain SQL Server integrations via JDBC for workflow metadata, feature flags, report configurations, and state management
  • Implement data quality frameworks with validation layers at each pipeline stage (patient info, diagnosis codes, CPT codes, facility mappings, rendering providers)
  • Write ad-hoc data correction, migration, and backfill scripts for production data reconciliation
  • Participate in code reviews, PR validation pipelines, and engineering best practices (commit/lint, conventional commits)
Required Skills & Qualifications:
  • Core: 6–7 years of data engineering experience with 3+ years on Databricks/Spark
  • Languages: Proficient in PySpark (primary) and Scala/Spark; Python 3.11+
  • Data Lake: Deep expertise in Delta Lake — medallion architecture, ACID transactions, schema evolution, Z-ordering, partition pruning
  • Databases: Production experience with MongoDB (aggregation pipelines, indexing, bulk sync) and SQL Server (JDBC, stored procedures)
  • EDI/Healthcare: Familiarity with HIPAA EDI standards (X12 837, 835, 277, 999) or willingness to learn quickly
  • DevOps: Azure Pipelines (YAML), CI/CD for Databricks artifacts, PR validation
  • Data Quality: Experience implementing validation frameworks, data reconciliation, and error tracking
Preferred Qualifications:
  • Experience in healthcare/RCM domain — claims lifecycle, payer/provider workflows, CPT/ICD-10 coding would be a plus
  • Hands-on with Databricks Unity Catalog for data governance
  • Exposure to AI/ML integration in data pipelines — LLM orchestration, model observability (Langfuse or similar)
  • Experience with feature flag systems for gradual data pipeline rollouts
  • Performance tuning Spark jobs at scale (shuffle optimization, broadcast joins, adaptive query execution)
What You'll Impact:
  • Process thousands of healthcare claims daily across multiple clients
  • Reduce claim processing time through pipeline optimization
  • Enable AI-powered medical coding that improves coder productivity
  • Deliver real-time operational dashboards for revenue cycle teams
  • Ensure HIPAA-compliant data processing at every layer
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Azure Databricks Engineer
Azure Databricks Engineer

Insight Global • Woonsocket (RI)

On-site
USD 120,000 - 160,000
Azure Databricks Engineer (Payer Data)
Azure Databricks Engineer (Payer Data)

Insight Global • Woonsocket (RI)

On-site
USD 130,000 - 180,000
Senior Data Engineer
Senior Data Engineer

On-Demand Group • BLOOMINGTON (MN)

Hybrid
USD 103,000 - 124,000
Healthcare benefits
Hybrid/remote options
Data Engineer
Data Engineer

NationsBenefits • Plantation (FL)

On-site
USD 95,000 - 125,000
DataBricks Data Engineer
DataBricks Data Engineer

Prodapt • Irving (TX)

On-site
USD 140,000 - 190,000
1104 | Senior Data Engineer (Databricks)
1104 | Senior Data Engineer (Databricks)

Intetics • Town of Poland (NY)

On-site
USD 140,000 - 190,000
1104 | Senior Data Engineer (Databricks)
1104 | Senior Data Engineer (Databricks)

Intetics • Spain (TX)

On-site
USD 120,000 - 160,000
Data Engineer
Data Engineer

InfoVision Inc. • Detroit (MI)

On-site
USD 105,000 - 155,000
Databricks Developer – Medicare
Databricks Developer – Medicare

Big Resourcing • Northern (KY)

Hybrid
USD 140,000 - 190,000
Director, Business Intelligence & Data Engineering
Director, Business Intelligence & Data Engineering

Salud Healthcare • Town of Florida (NY)

On-site
USD 180,000 - 240,000