Lead Data Engineer

Reuben Cooley Inc.

Cary (NC)

Presencial

USD 120.000 - 180.000

Jornada completa

hace 9 horas
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

No envíes un currículum genérico — crea un currículum y una carta de presentación adaptados a este puesto concreto.

Supera los filtros ATS

Descripción de la vacante

Reuben Cooley Inc. seeks a senior data engineer to architect and deliver production-grade pipelines using Python, Scala, PySpark, Spark, and Delta Lake.

You will optimize batch and streaming workloads, ensure robust schemas and data contracts, and own Databricks platforms and governance. The role requires 12–18 years of relevant experience, strong security practices, and deep expertise in AI data paths, LangChain tooling, and model serving.

Formación

  • Expert-level proficiency in Python, Scala, and PySpark with production-grade pipelines.
  • Strong SQL and data modelling, schema design and data contracts.
  • Databricks Delta Lake, Unity Catalog, Jobs & Workflows, and cluster management.
  • 12–18 years of total experience in data engineering or related disciplines.
  • Proven delivery of a medallion/lakehouse architecture at enterprise scale.
  • Azure security and governance: Entra ID, RBAC, Key Vault, private endpoints, and PII handling.
  • CI/CD and infrastructure as code: Azure DevOps, Terraform, Databricks Asset Bundles, and automated testing of data pipelines.
  • Clear technical writing and the ability to present and defend a design to engineers and non-technical stakeholders.

Responsabilidades

  • Design and ship production-grade data pipelines using Python, Scala, and Spark.
  • Optimize large-scale batch and streaming workloads with Delta Lake and Spark technologies.
  • Define data contracts, schemas, and lineage to ensure reliable data assets.
  • Lead AI/ML data paths, including LLM-based pipelines, retrieval, and prompt engineering.

Conocimientos

Python
Scala
PySpark
SQL
Data Modeling
Schema Design
Delta Lake
Databricks
Unity Catalog
Jobs & Workflows
Spark

Herramientas

Databricks
Azure OpenAI
OpenAI
Terraform
Azure DevOps
LangChain
LlamaIndex
LangGraph
Model Serving

Descripción del empleo

  • Expert - level proficiency in Python, Scala, and PySpark, with a strong track record of designing and delivering production-ready, modular, and well-tested solutions; developing and troubleshooting Spark workloads; and optimizing large-scale batch and streaming data pipelines using Delta Lake and Spark technologies.
  • Strong SQL and data modelling dimensional and normalised; schema design and data contract definition.
  • Databricks expertise Delta Lake, Unity Catalog, Jobs & Workflows, cluster and pool management, performance tuning, Model Serving.
  • 3+ years designing and shipping LLM-based systems in production: RAG pipelines, agentic / tool-calling workflows, structured output, chunking and embedding strategy, vector and hybrid retrieval, and prompt engineering.
  • Evaluation discipline — golden datasets, regression suites, accuracy and hallucination tracking, human-in-the-loop feedback loops; you measure AI quality, not assert it.
  • Hands-on experience with LangChain, LlamaIndex, or LangGraph, plus at least one provider stack (Azure OpenAI, OpenAI, or Databricks Model Serving).
  • Metadata-driven thinking — schema inference, data profiling, lineage, catalogs, and configuration-driven frameworks that onboard the next source without new code.
  • 12–18 years of total experience in data engineering, data platform delivery, or related disciplines.
  • Proven delivery of a medallion/lakehouse architecture at enterprise scale — not just familiarity with the concept.
  • Azure security and governance: Entra ID, managed identities, RBAC, POSIX ACLs on ADLS Gen2, Key Vault, private endpoints, and PII handling.
  • CI/CD and infrastructure as code: Azure DevOps, Terraform, Databricks Asset Bundles, and automated testing of data pipelines.
  • Clear technical writing and the ability to present and defend a design to both engineers and non-technical stakeholders.
Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Lead Data Engineer (Hands-On)
Lead Data Engineer (Hands-On)

Cephas Consultancy Services Private Limited • Cary (NC)

Híbrido
USD 150.000 - 210.000
Data Engineer
Data Engineer

InfoVision, Inc. • EE. UU.

Presencial
USD 110.000 - 150.000
Lead Data Engineer
Lead Data Engineer

Bitwise • Richmond (VA)

Presencial
USD 140.000 - 200.000
Databricks Developer
Databricks Developer

Delan Associates, Inc • Louisville (KY)

Presencial
USD 100.000 - 150.000
Lead Data Engineer (Hands-On)
Lead Data Engineer (Hands-On)

Brickstech • Cary (NC)

Híbrido
USD 150.000 - 185.000
Data Engineer
Data Engineer

InfoVision Inc. • Detroit (MI)

Presencial
USD 105.000 - 155.000
Senior Data Engineer - Databricks
Senior Data Engineer - Databricks

UNAVAILABLE • McLean (VA)

Presencial
USD 120.000 - 160.000
Azure Databricks Data Architect
Azure Databricks Data Architect

Ascendum Solutions • Cincinnati (OH)

Presencial
USD 140.000 - 170.000
Senior Data Engineer with Databricks Exp. - 100% Remote
Senior Data Engineer with Databricks Exp. - 100% Remote

SDH Systems • EE. UU.

A distancia
USD 140.000 - 190.000
Data Engineer - Databricks
Data Engineer - Databricks

UNAVAILABLE • McLean (VA)

Presencial
USD 120.000 - 170.000