Lead Databricks ML & AI Ops Engineer

KData Inc.

United States

Remote

USD 180,000 - 240,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

KData Inc. seeks a Lead Databricks ML & AI Ops Engineer to architect, scale, and govern the ML/AI platforms powering our enterprise data products.

You will define the technical vision for end-to-end MLOps on Databricks, Delta Lake, and Unity Catalog. As the primary technical authority, you will lead a team of senior engineers, collaborate with data science leadership, and align with product teams to deliver robust AI features.

Qualifications

  • 8+ years of software, data, or DevOps engineering experience, with 5+ years in production ML/AI systems.
  • Deep knowledge of the Databricks ecosystem: Workflows, Unity Catalog, Delta Live Tables.

Responsibilities

  • Define the architectural roadmap for the enterprise ML platform.
  • Establish engineering standards for CI/CD, testing, and governance.
  • Lead, mentor, and coach senior engineers.
  • Engage with Databricks account teams and represent the company at events.
  • Architect reusable batch and streaming feature pipelines with Delta Live Tables and Spark Structured Streaming.
  • Standardize cluster configurations for cost and performance optimization.

Skills

Databricks expertise
Distributed systems
Python
DevOps / GitOps
Cloud platforms
Kubernetes
Spark internals
MLOps

Tools

Databricks Workflows
Delta Live Tables
Unity Catalog
MLflow
LangChain
Mosaic AI

Job description

This is a remote position. We are seeking a Lead Databricks ML & AI Ops Engineer to architect, scale, and govern the machine learning and AI platforms that power our enterprise data products. In this role, you will define the technical vision and engineering strategy for our end-to-end MLOps lifecycle—spanning from exploratory data science to global-scale production deployment on Databricks, Delta Lake, and Unity Catalog. As a Lead Engineer , you will serve as the primary architect and technical authority for our ML platform. You will lead and mentor a talented team of senior engineers , collaborate closely with data science executives, and align with product teams to deliver robust, high-performance AI features. This role requires a balance of advanced distributed computing expertise, cutting-edge Generative AI platform design, and a strong DevOps/GitOps mindset.

Key Responsibilities
Technical Leadership & Strategy
  • Platform Vision: Define the architectural roadmap for the enterprise ML platform, steering the migration away from legacy systems (e.g., Apache Airflow) to modern Databricks Workflows and Asset Bundles.
  • Standards & Governance: Establish, document, and enforce global engineering standards for code quality, CI/CD pipelines, and automated testing. Lead enterprise-wide ML governance and data security strategies utilizing Unity Catalog.
  • Team Mentorship: Lead, mentor, and coach senior and mid-level engineers. Conduct advanced design and code reviews to foster a culture of technical excellence.
  • Vendor & Community Engagement: Act as the primary technical point of contact for Databricks account teams. Represent the company at external conferences and contribute to tech blogs.
ML & LLM Platform Architecture
  • Enterprise Pipelines: Architect reusable, highly efficient batch and streaming feature pipelines using Delta Live Tables, Spark Structured Streaming, and Databricks Feature Store.
  • Compute Optimization: Standardize cluster configurations (including multi-node GPU training) to optimize complex AutoML and deep learning workloads (PyTorch, Hugging Face) for cost and performance.
  • Generative AI Infrastructure: Design secure, production-grade frameworks for Large Language Model (LLM) operations (LLMOps), including robust RAG pipelines, agentic workflows, and Vector Search indexing using Mosaic AI and LangChain.
  • Model Lifecycle Management: Oversee the global setup of MLflow, defining enterprise staging gates, automated retraining loops, and instant rollback strategies.
Infrastructure & Operations (Ops)
  • Production Serving: Architect high-availability, low-latency model serving frameworks utilizing Databricks Model Serving, Mosaic AI Gateway, and containerized deployments via Kubernetes/FastAPI.
  • Infrastructure-as-Code: Lead the standardization of Infrastructure-as-Code (IaC) using Terraform or Pulumi to automate environment provisioning across multiple cloud regions.
  • Observability & SLAs: Establish advanced, proactive monitoring for data drift, model performance decay, hallucination rates, and pipeline SLAs using Databricks Lakehouse Monitoring, Prometheus, and Grafana.
Requirements
Required Qualifications
  • Experience: 8+ years of software, data, or DevOps engineering experience, with at least 5 years dedicated specifically to production-grade ML/AI systems. Proven experience in a technical lead or architectural capacity.
  • Databricks Mastery: Deep, expert-level knowledge of the Databricks ecosystem, including Workflows, Delta Live Tables, Unity Catalog, Mosaic AI, and Databricks Asset Bundles.
  • Distributed Systems: Advanced understanding of Spark internals (DAG optimization, shuffle tuning, memory management) at an enterprise, multi-terabyte scale.
  • Core Engineering: Exceptional Python skills (writing highly optimized, type-annotated, and modular code) alongside deep familiarity with frameworks like PyTorch, scikit-learn, and XGBoost.
  • DevOps & GitOps: Strong background in enterprise GitOps workflows (e.g., trunk-based development), Docker containerization, Kubernetes orchestration, and complex GitHub Actions CI/CD pipelines.
  • Cloud Architecture: Extensive hands‑on experience provisioning and securing cloud infrastructure on AWS, Azure, or GCP.
Preferred Qualifications
  • Certification: Databricks Certified Machine Learning Professional and/or Databricks Certified Enterprise Architect.
  • Advanced AI/Governance: Direct experience implementing Responsible AI frameworks, model cards, bias auditing, and cost-tracking guardrails for LLMs.
  • Open Source: Active contributor to open-source ML, MLOps, or data engineering projects (e.g., MLflow, Delta Lake, LangChain).
  • Data Mesh: Experience designing Lakehouse architectures within a decentralized Data Mesh organizational framework.
Technical Environment
  • Platform & Storage: Databricks (AWS/Azure/GCP), Delta Lake, Unity Catalog, S3/ADLS Gen2.
  • Orchestration & CI/CD: Databricks Workflows, Databricks Asset Bundles, GitHub Actions, Terraform.
  • ML & LLM Frameworks: MLflow, PyTorch, Hugging Face, Mosaic AI, LangChain, Databricks Vector Search, GPT-4 / Claude APIs.
  • Serving & Observability: Databricks Model Serving, FastAPI, Docker, Kubernetes, Databricks Lakehouse Monitoring, Prometheus, Grafana.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

MLOps Lead Engineer
MLOps Lead Engineer

Tiger Analytics • St. Louis (MO), Northern (KY)

On-site
USD 140,000 - 190,000
ML Architect Databricks
ML Architect Databricks

Ktek Talent Solutions • United States

Remote
USD 150,000 - 210,000
MLOps Engineer (Databricks Specialist)
MLOps Engineer (Databricks Specialist)

KData Inc. • United States

Remote
USD 120,000 - 160,000
Remote work
Senior AI/ML Platform Engineer
Senior AI/ML Platform Engineer

TalentBridge • Denver (CO)

On-site
USD 140,000 - 210,000
Remote Lead Databricks ML & AI Ops Engineer
Remote Lead Databricks ML & AI Ops Engineer

KData Inc. • United States

Remote
USD 180,000 - 240,000
Lead Data Engineer
Lead Data Engineer

Accrescent Group • Cary (NC)

On-site
USD 140,000 - 180,000
Databricks SME
Databricks SME

Scicom Infrastructure Services, Inc. • United States

On-site
USD 140,000 - 210,000
Data Engineer
Data Engineer

InfoVision, Inc. • United States

On-site
USD 110,000 - 150,000
Databricks SME
Databricks SME

Scicominfra • Atlanta (GA)

On-site
USD 180,000 - 240,000
Databricks Engineer
Databricks Engineer

CMT Services, Inc. • Adelphi (MD)

On-site
USD 100,000 - 130,000