Lead Data and ML Infrastructure Engineer

Harrison Clarke

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Meaningful equity
Autonomy and ownership

Job summary

Harrison Clarke in San Francisco is seeking a Lead Data & ML Infrastructure Engineer to own the production backbone for AI systems at scale. The role focuses on pipelines, platforms, and observability that convert experiments into reliable production results.

You will design systems where models monitor themselves, adapt to drift in real time, and decisions shape AI capabilities. High autonomy and impact are expected.

Qualifications

  • 5+ years of production ML and data infrastructure experience.
  • Experienced in modern data infra stacks: Spark, Delta Lake, Kafka, Flink, Airflow.
  • Proficient with ML lifecycle platforms (MLflow, Kubeflow, SageMaker).
  • Strong in observability, data quality, and pipeline SLAs.

Responsibilities

  • Architect and build next-gen ML pipelines that self-heal and auto-retrain on drift.
  • Design the data infrastructure layer for an AI-native stack with real-time and batch processing.
  • Build production-grade ML observability for model health and drift.
  • Push real-time feature serving and streaming computation.
  • Own infrastructure reliability: graceful degradation and zero-downtime deployments.
  • Bridge to LLM/agentic infra: model serving, inference optimization, retrieval-augmented generation.

Skills

Distributed systems
Production ML
Observability
Technical communication

Tools

Spark
Delta Lake
Kafka
Flink
Airflow
MLflow
Kubeflow
SageMaker

Job description

Build the Engine Behind Frontier Intelligent Systems - San Francisco

The next generation of AI isn't bottlenecked by models, it's bottlenecked by infrastructure. The companies that win will be the ones whose ML systems retrain themselves, detect their own failures, and scale without human intervention. We're hiring the engineer who builds that.

We're looking for a Lead Data & ML Infrastructure Engineer to own the production backbone that powers AI systems at scale. This is the layer between research breakthroughs and real-world impact: the pipelines, platforms, and observability systems that turn experimental models into reliable, self-improving production systems.

This is a hands-on, high-autonomy role at the frontier of what production ML infrastructure looks like. You'll be designing systems where models monitor their own performance, pipelines adapt to shifting data distributions in real time, and infrastructure decisions directly shape what our AI is capable of.

What You'll Do
  • Architect and build next-generation ML pipelines that go beyond train-deploy-monitor — systems that self-heal, auto-retrain on drift signals, and continuously validate their own outputs.
  • Design the data infrastructure layer for an AI-native stack — high-throughput feature pipelines, real-time and batch processing, intelligent storage tiering, and data versioning at scale.
  • Build production-grade ML observability that treats model health as a first-class signal — semantic drift detection, feature attribution monitoring, pipeline SLA enforcement, and automated anomaly response.
  • Push the boundary on feature infrastructure — real-time feature serving, streaming feature computation, and feature platforms that unify offline training and online inference with zero skew.
  • Own infrastructure reliability at scale — designing for graceful degradation, self-healing pipelines, automated failover, and zero-downtime deployments across distributed ML systems.
  • Build the bridge to LLM and agentic infrastructure — model serving for large language models, inference optimization, context management, retrieval-augmented generation pipelines, and evaluation frameworks for non-deterministic systems.
You Are
  • A builder who operates at the intersection of distributed systems and machine learning
  • Deeply experienced in production ML and data infrastructure at scale (typically 5+ years).
  • Fluent in modern data infrastructure — Spark, Delta Lake, Kafka, Flink, Airflow, or equivalent. You think in terms of data flow topology, not just individual tools.
  • Experienced with ML lifecycle platforms — MLflow, Kubeflow, SageMaker, or systems you built yourself.
  • An engineer who designs for observability from the start — you've built systems that monitor data quality, feature drift, model health, and pipeline SLAs, and you believe operational intelligence is as important as model intelligence.
  • A strong technical communicator who can partner with ML researchers, product engineers, and leadership — translating infrastructure capabilities into AI capabilities.
Nice to Have
  • Experience building real-time feature stores or streaming feature computation systems.
  • Hands-on work with LLM serving infrastructure — model routing, batched inference, KV-cache optimization, quantization, or multi-model orchestration.
  • Background in building evaluation and regression frameworks for non-deterministic AI systems (LLMs, agents, generative models).
  • Experience with knowledge graph infrastructure, vector databases, or retrieval-augmented generation at scale.
Tech Stack (Indicative, Not Prescriptive)
What We Offer
  • Competitive compensation + meaningful equity
  • A seat at the frontier, you will build infrastructure for AI systems that don't exist yet at most companies
  • A team that treats ML infrastructure as a first-class engineering discipline, not a support function
  • High autonomy, real ownership, and the mandate to make architecture decisions that shape the product
  • The chance to define what production-grade AI looks like for the next decade
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • San Francisco (CA)

On-site
USD 180,000 - 300,000
Cash and equity
Ownership & autonomy
Visa sponsorship
+5
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • Seattle (WA)

On-site
USD 180,000 - 300,000
Health insurance
Vision insurance
Dental insurance
+3
Software Engineer
Software Engineer

Venture Up • San Francisco (CA)

On-site
USD 350,000 - 600,000
Equity
Office in San Francisco FiDi
On-site with optional remote Sunday (½
Backend Software Engineer (ML Infra)
Backend Software Engineer (ML Infra)

Rockstar • San Francisco (CA)

On-site
USD 100,000 - 130,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Clera • San Mateo (CA)

On-site
USD 180,000 - 240,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
Machine Learning Engineer
Machine Learning Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 150,000 - 190,000
ML Engineer, Cloud Platform
ML Engineer, Cloud Platform

PriorLabs GmbH • New York (NY)

On-site
USD 140,000 - 190,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Acceler8 Talent • San Francisco (CA)

Hybrid
USD 233,000 - 275,000
ML Engineer – AI-Powered Automation & Workflow Intelligence
ML Engineer – AI-Powered Automation & Workflow Intelligence

Blue-Signal-Search • San Francisco (CA)

On-site
USD 130,000 - 160,000
Competitive compensation package
Significant equity upside
Collaborative in-person work environment