Build the Engine Behind Frontier Intelligent Systems - San Francisco
The next generation of AI isn't bottlenecked by models, it's bottlenecked by infrastructure. The companies that win will be the ones whose ML systems retrain themselves, detect their own failures, and scale without human intervention. We're hiring the engineer who builds that.
We're looking for a Lead Data & ML Infrastructure Engineer to own the production backbone that powers AI systems at scale. This is the layer between research breakthroughs and real-world impact: the pipelines, platforms, and observability systems that turn experimental models into reliable, self-improving production systems.
This is a hands-on, high-autonomy role at the frontier of what production ML infrastructure looks like. You'll be designing systems where models monitor their own performance, pipelines adapt to shifting data distributions in real time, and infrastructure decisions directly shape what our AI is capable of.
What You'll Do
- Architect and build next-generation ML pipelines that go beyond train-deploy-monitor — systems that self-heal, auto-retrain on drift signals, and continuously validate their own outputs.
- Design the data infrastructure layer for an AI-native stack — high-throughput feature pipelines, real-time and batch processing, intelligent storage tiering, and data versioning at scale.
- Build production-grade ML observability that treats model health as a first-class signal — semantic drift detection, feature attribution monitoring, pipeline SLA enforcement, and automated anomaly response.
- Push the boundary on feature infrastructure — real-time feature serving, streaming feature computation, and feature platforms that unify offline training and online inference with zero skew.
- Own infrastructure reliability at scale — designing for graceful degradation, self-healing pipelines, automated failover, and zero-downtime deployments across distributed ML systems.
- Build the bridge to LLM and agentic infrastructure — model serving for large language models, inference optimization, context management, retrieval-augmented generation pipelines, and evaluation frameworks for non-deterministic systems.
You Are
- A builder who operates at the intersection of distributed systems and machine learning
- Deeply experienced in production ML and data infrastructure at scale (typically 5+ years).
- Fluent in modern data infrastructure — Spark, Delta Lake, Kafka, Flink, Airflow, or equivalent. You think in terms of data flow topology, not just individual tools.
- Experienced with ML lifecycle platforms — MLflow, Kubeflow, SageMaker, or systems you built yourself.
- An engineer who designs for observability from the start — you've built systems that monitor data quality, feature drift, model health, and pipeline SLAs, and you believe operational intelligence is as important as model intelligence.
- A strong technical communicator who can partner with ML researchers, product engineers, and leadership — translating infrastructure capabilities into AI capabilities.
Nice to Have
- Experience building real-time feature stores or streaming feature computation systems.
- Hands-on work with LLM serving infrastructure — model routing, batched inference, KV-cache optimization, quantization, or multi-model orchestration.
- Background in building evaluation and regression frameworks for non-deterministic AI systems (LLMs, agents, generative models).
- Experience with knowledge graph infrastructure, vector databases, or retrieval-augmented generation at scale.
Tech Stack (Indicative, Not Prescriptive)
What We Offer
- Competitive compensation + meaningful equity
- A seat at the frontier, you will build infrastructure for AI systems that don't exist yet at most companies
- A team that treats ML infrastructure as a first-class engineering discipline, not a support function
- High autonomy, real ownership, and the mandate to make architecture decisions that shape the product
- The chance to define what production-grade AI looks like for the next decade