Lead Data Engineer – Azure Databricks

Solvimate Pvt. Ltd.

Cary (NC)

Hybrid

USD 140,000 - 145,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Solvimate Pvt. Ltd. is building an enterprise-scale data and AI platform with a hybrid role in Cary, NC.

You will shape the Lakehouse foundation using Azure Databricks and related Azure tools, delivering end-to-end data engineering solutions and AI-assisted capabilities. You will design and deploy production pipelines, governance, and scalable architectures, while mentoring engineers and collaborating with stakeholders.

Qualifications

  • 12+ years of professional experience in Data Engineering or related fields.
  • Expert-level proficiency in Python, Scala, and PySpark.
  • Strong SQL and data modeling experience, including dimensional and normalized models.
  • Extensive hands-on experience with Azure Databricks and Delta Lake.
  • Experience with Databricks Jobs & Workflows, Unity Catalog, cluster management and performance tuning.
  • Hands-on Azure Data Factory and metadata-driven pipeline frameworks.
  • Experience with Azure Event Hubs, Kafka, or Spark Structured Streaming.
  • Proven experience designing enterprise-scale Lakehouse/Medallion architectures.
  • At least 3 years of hands-on experience designing and deploying production LLM/AI systems.
  • Strong understanding of RAG, embeddings, vector/hybrid retrieval, prompt engineering, and agentic workflows.
  • Experience with LangChain, LlamaIndex, or LangGraph.
  • Experience with Azure OpenAI, OpenAI, or Databricks Model Serving.
  • Experience implementing AI evaluation frameworks including golden datasets, regression testing, and human-in-the-loop validation.
  • Strong understanding of Azure security: Entra ID, RBAC, Managed Identities, Key Vault, ACLs, Private Endpoints.

Responsibilities

  • Design and implement scalable Lakehouse/Medallion Architecture (Bronze, Silver, Gold).
  • Develop production-grade pipelines using Python, Scala, PySpark, Spark Structured Streaming, Delta Lake.
  • Build metadata-driven ingestion frameworks for batch, CDC, and streaming sources.
  • Work across Azure Databricks, ADLS Gen2, Azure Data Factory, Azure Event Hubs, and Kafka.
  • Design Delta Lake tables with partitioning, schema evolution, and data retention; enforce data contracts.
  • Optimize Spark workloads, clusters, pipelines, and overall platform performance.
  • Establish engineering standards: coding practices, automated testing, code/PR reviews.
  • Implement CI/CD with Azure DevOps, Terraform, and Databricks Asset Bundles.
  • Develop AI-assisted ingestion, canonical mapping, data quality, reconciliation workflows.
  • Design production-grade LLM, RAG, and agentic AI workflows including embeddings and vector search.
  • Build semantic-layer and text-to-SQL capabilities for enterprise data access.
  • Implement data governance, lineage, access control, and PII protection using Unity Catalog.
  • Monitor data quality, schema drift, anomalies, reconciliation issues, and pipeline health.
  • Mentor data engineers and present designs to stakeholders.

Skills

Python
Scala
PySpark
SQL
Azure Databricks
Delta Lake
Unity Catalog
Spark Structured Streaming
Azure Data Factory
ETL / Data Pipelines
LLM / AI systems
LangChain / LlamaIndex
Text-to-SQL

Tools

Azure Databricks
Delta Lake
ADLS Gen2
Azure Data Factory
Azure Event Hubs
Kafka
Databricks Asset Bundles
Terraform
Azure OpenAI / OpenAI
Databricks Jobs & Workflows
Unity Catalog

Job description

Solvimate is building an enterprise-scale data and AI platform initiative, and this hybrid role in Cary, NC focuses on delivering modern data engineering solutions end to end. In this position, you will help shape the technical foundation of a Lakehouse platform using Azure Databricks and related Azure data and AI tooling, with opportunities to drive architecture, optimization, governance, and production delivery.

Salary: USD 140,000 - 145,000 per year
Experience: 12+ years
Location: Cary, NC (hybrid)

Responsibilities
  • Design and implement a scalable Lakehouse / Medallion Architecture using Bronze, Silver, and Gold layers.
  • Develop production-grade pipelines with Python, Scala, PySpark, Spark Structured Streaming, and Delta Lake.
  • Build metadata-driven and parameterized ingestion frameworks for batch, database extracts, CDC, and streaming sources.
  • Work across Azure Databricks, ADLS Gen2, Azure Data Factory, Azure Event Hubs, and Kafka.
  • Design Delta Lake tables including partitioning, schema evolution, and data retention processes, along with data contracts.
  • Troubleshoot and optimize Spark workloads, clusters, pipelines, and overall platform performance.
  • Establish engineering standards: coding practices, automated testing, and code or PR reviews.
  • Implement CI/CD with Azure DevOps, Terraform, and Databricks Asset Bundles.
  • Develop AI-assisted capabilities for ingestion, canonical mapping, data quality, and reconciliation.
  • Design production-grade LLM, RAG, and agentic AI workflows, including embeddings, vector search, hybrid retrieval, prompt engineering, and tool-calling architectures.
  • Build semantic-layer and text-to-SQL capabilities to support enterprise data access.
  • Implement data governance, lineage, access control, and PII protection using Unity Catalog.
  • Monitor data quality, schema drift, anomalies, reconciliation issues, and pipeline health.
  • Mentor data engineers with technical guidance on architecture, implementation, and best practices.
  • Present and defend technical designs to engineering teams, architects, and business stakeholders.
Requirements
  • 12–18 years of professional experience in Data Engineering, Data Platforms, or related fields.
  • Expert-level proficiency in Python, Scala, and PySpark.
  • Strong SQL and data modeling experience, including dimensional and normalized data models.
  • Extensive hands-on experience with Azure Databricks and Delta Lake.
  • Strong knowledge of Databricks Jobs & Workflows, Unity Catalog, cluster management, and performance tuning.
  • Experience with ADLS Gen2.
  • Strong experience with Azure Data Factory and metadata-driven pipeline frameworks.
  • Experience with Azure Event Hubs, Kafka, or Spark Structured Streaming.
  • Proven experience designing and delivering enterprise-scale Lakehouse / Medallion architectures.
  • At least 3 years of hands-on experience designing and deploying production LLM/AI systems.
  • Strong understanding of RAG, embeddings, vector/hybrid retrieval, prompt engineering, and agentic workflows.
  • Experience with LangChain, LlamaIndex, or LangGraph.
  • Experience with at least one AI platform such as Azure OpenAI, OpenAI, or Databricks Model Serving.
  • Experience implementing AI evaluation frameworks including golden datasets, regression testing, and human-in-the-loop validation.
  • Strong understanding of Azure security: Entra ID, RBAC, Managed Identities, Key Vault, ACLs, and Private Endpoints.
  • Experience with Azure DevOps, Terraform, CI/CD, and automated testing.
  • Excellent communication, documentation, problem-solving, and stakeholder management skills.
Preferred Qualifications
  • Experience with Knowledge Graphs and Ontologies, including RDF/SPARQL, Neo4j, or graph-based data modeling.
  • Experience developing enterprise-scale Text-to-SQL or semantic-layer solutions.
  • Experience with ML-based anomaly detection for transactional or time-series data.
  • Previous experience in Financial Services or Insurance.
  • Experience with LLMOps / MLOps, model and prompt versioning, cost management, and observability.
  • Experience with dbt, Great Expectations, or similar data quality frameworks.
  • Databricks Data Engineer Professional certification.
  • Azure certifications such as Azure DP-203, DP-700, or AZ-305.
Application Requirements
  • Full Name
  • Current Location
  • Contact Number
  • Email Address
  • Work Authorization – US Citizen
  • LinkedIn Profile
  • Availability to Start
  • Interview Availability for the Next 3 Days
  • Updated Resume
What We Are Looking For

Solvimate needs a senior engineer with strong architecture skills and hands-on coding ability. The role involves ownership from technical design through implementation, testing, deployment, and production support. Candidates who have not written or reviewed production code recently may not be suitable for this position.

Technologies: Azure Databricks, PySpark, Scala, Python, Delta Lake, Azure data services, Lakehouse / Medallion Architecture, Spark Structured Streaming, Azure Data Lake Storage Gen2 (ADLS Gen2), Azure Data Factory, Azure Event Hubs, Kafka, Azure DevOps, Terraform, Databricks Asset Bundles, LLM, RAG, Unity Catalog, embeddings, vector search, hybrid retrieval, prompt engineering, tool-calling architectures, text-to-SQL, LangChain, LlamaIndex, LangGraph, Azure OpenAI, OpenAI, Databricks Model Serving, Entra ID, RBAC, Managed Identities, Key Vault, ACLs, Private Endpoints, Databricks Jobs & Workflows, CI/CD, automated testing.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Data Engineer (Hands-On)
Lead Data Engineer (Hands-On)

Cephas Consultancy Services Private Limited • Cary (NC)

Hybrid
USD 150,000 - 210,000
Lead Data Engineer (Hands-On)
Lead Data Engineer (Hands-On)

Brickstech • Cary (NC)

Hybrid
USD 150,000 - 185,000
Lead Data Engineer
Lead Data Engineer

SIDRAM TECHNOLOGIES • Cary (NC)

Hybrid
USD 160,000 - 230,000
Sr. Azure Data Engineer
Sr. Azure Data Engineer

Noblesoft Technologies Inc. • Greenfield (IN)

On-site
USD 120,000 - 180,000
Senior Data Engineer with Databricks Exp. - 100% Remote
Senior Data Engineer with Databricks Exp. - 100% Remote

SDH Systems • United States

Remote
USD 140,000 - 190,000
Senior Azure Databricks Data Engineer
Senior Azure Databricks Data Engineer

EXL • New York (NY)

Hybrid
USD 140,000 - 190,000
Senior Databricks Data Engineer
Senior Databricks Data Engineer

CRED • College Station (TX)

On-site
USD 120,000 - 160,000
Incentive Program
Accrued Time Off
401(k) with match
+4
Lead Data Engineer
Lead Data Engineer

Recru • Spring (TX)

On-site
USD 130,000 - 170,000
Data Engineer
Data Engineer

Solvd • United States

Remote
USD 110,000 - 160,000
Platform Engineer Consultant
Platform Engineer Consultant

Deloitte France • St. Louis (MO)

On-site
USD 140,000 - 190,000
Travel opportunities
Client project exposure