Member of Technical Staff

Harrison Clarke

San Francisco (CA)

On-site

USD 180,000 - 280,000

Full time

37 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Harrison Clarke partners with a well-funded frontier AI lab to pursue a Machine Learning Infrastructure Engineer who will design and build the core infrastructure powering next-generation large-scale model training and deployment.

You will own the stack for training orchestration, distributed compute, and model serving, collaborating with researchers to deliver reliable, scalable platforms that accelerate AI research and production.

Qualifications

  • 5+ years of software engineering with ML infra focus.
  • Experience with ML tooling and distributed systems.
  • Expertise in Python and GPU programming.

Responsibilities

  • Design foundational ML infrastructure for training orchestration, distributed compute frameworks, GPU/TPU cluster management, and experiment tracking systems.
  • Optimize large-scale training pipelines for throughput, fault tolerance, checkpointing, and resource utilization across multi-node, multi-accelerator environments.
  • Develop and maintain model serving infrastructure with low-latency inference, packaging, deployment pipelines, and autoscaling.
  • Build internal platforms and tooling for self-service compute, data pipelines, and reproducible experiment workflows.
  • Collaborate closely with research teams to translate bleeding-edge ML requirements into reliable, scalable systems.
  • Drive infrastructure reliability with monitoring, observability, capacity planning, and incident response for mission-critical ML workloads.

Skills

Python
Distributed systems
ML infrastructure
GPUs/accelerators
Kubernetes
Large-scale training
Experiment tracking

Tools

PyTorch
JAX
DeepSpeed
Megatron
CUDA
NCCL
InfiniBand

Job description

We're partnering with a well-funded frontier AI lab to find a Machine Learning Infrastructure Engineer who will design and build the foundational ML/AI infrastructure powering the next generation of large-scale model training and deployment.

This is a rare opportunity to join a team at the absolute cutting edge — working on systems that underpin frontier research and production AI at massive scale.

About the Role

As an ML Infrastructure Engineer, you'll own and evolve the core infrastructure stack that researchers and engineers depend on daily. You'll work at the intersection of distributed systems, high-performance computing, and machine learning — building platforms that enable training runs across thousands of accelerators, efficient model serving, and seamless experimentation at scale.

What You'll Do
  • Design and build foundational ML infrastructure — training orchestration, distributed compute frameworks, GPU/TPU cluster management, and experiment tracking systems
  • Optimize large-scale training pipelines — improving throughput, fault tolerance, checkpointing, and resource utilization across multi-node, multi-accelerator environments
  • Develop and maintain model serving infrastructure — low-latency inference systems, model packaging, deployment pipelines, and autoscaling
  • Build internal platforms and tooling — enabling researchers to iterate faster with self-service compute, data pipelines, and reproducible experiment workflows
  • Collaborate closely with research teams — translating bleeding-edge ML requirements into reliable, scalable systems
  • Drive infrastructure reliability — monitoring, observability, capacity planning, and incident response for mission-critical ML workloads
What You Bring
  • 5+ years of software engineering experience, with significant time spent on ML infrastructure, distributed systems, or platform engineering
  • Expertise in Python
  • Hands-on experience with large-scale distributed training frameworks (PyTorch, JAX, DeepSpeed, Megatron, or equivalent)
  • Strong background in GPU/accelerator programming, CUDA, and high-performance networking (NCCL, InfiniBand, RoCE)
  • Production experience with Kubernetes, container orchestration, and cloud infrastructure (AWS, GCP, or Azure)
  • Familiarity with ML experiment tracking, data pipelines, and model lifecycle management tooling
  • Experience building systems that operate at thousands-of-GPUs scale is a strong plus
  • Excellent problem-solving instincts and a bias toward shipping reliable, well-instrumented systems
Nice to Have
  • Experience at a frontier lab, hyperscaler, or research-heavy AI organization
  • Background in HPC, scientific computing, or large-scale simulation infrastructure
  • Contributions to open-source ML infrastructure projects
  • Familiarity with custom accelerator hardware (TPUs, Trainium, custom ASICs)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • San Francisco (CA)

On-site
USD 180,000 - 300,000
Cash and equity
Ownership & autonomy
Visa sponsorship
+5
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • Seattle (WA)

On-site
USD 180,000 - 300,000
Health insurance
Vision insurance
Dental insurance
+3
Machine Learning Infrastructure Engineer
Machine Learning Infrastructure Engineer

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 150,000 - 450,000
Comprehensive medical, dental, and vision
401(k) program
Generous PTO
+3
ML Infrastructure Engineer
ML Infrastructure Engineer

Strativ Group • Menlo Park (CA)

On-site
USD 250,000 - 320,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
ML Infra Engineer, Modeling
ML Infra Engineer, Modeling

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Infrastructure Engineer
AI Infrastructure Engineer

Fuel Talent • Seattle (WA)

On-site
USD 180,000 - 210,000
AI Infrastructure Engineer
AI Infrastructure Engineer

dicedemo • Boston (CT)

On-site
USD 130,000 - 170,000
Member of Technical Staff - ML Infra
Member of Technical Staff - ML Infra

Kindredventures • San Francisco (CA)

On-site
USD 160,000 - 220,000
ML Infra Engineer
ML Infra Engineer

Monograph • San Francisco (CA)

On-site
USD 120,000 - 160,000