ML Infrastructure Engineer

OP Recruiting

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

On-site work
Health insurance
Wellness stipend
Flexible PTO
Daily catered lunches
Commuter reimbursement

Job summary

OP Recruiting in San Francisco, CA is hiring a Machine Learning Infrastructure Engineer to own compute, training, and execution frameworks powering next-gen models. This role focuses on scaling distributed systems across large GPU clusters on-site.

You will collaborate with core platform teams, build reliable tooling, and drive performance benchmarks, while contributing to production-grade research-to-deployment workflows.

Qualifications

  • Experience designing and scaling ML infrastructure.
  • Familiarity with GPU clusters and distributed systems.
  • Strong programming in Python/C++ and tooling.
  • Ability to work with cross-functional teams.

Responsibilities

  • Architect high-throughput training and inference pipelines for rapid experimentation and low latency.
  • Benchmark system stacks to identify bottlenecks and optimize resources.
  • Evaluate model optimization developments and integrate into production.
  • Engineer multi-node, observable environments for parallelized training.
  • Maximize hardware efficiency across large GPU clusters.
  • Build internal tooling and telemetry to move research to production.

Job description

Job Title: Machine Learning Infrastructure Engineer
Location: San Francisco, CA Metro Area (100% On-Site)

About The Opportunity

An ultra-high-growth artificial intelligence platform company is seeking a Machine Learning Infrastructure Engineer to help architect the compute, training, and execution frameworks powering next-generation model performance. Backed by top-tier venture capital firms and serving elite technology enterprises, this team is scaling rapidly to solve complex unstructured data challenges. In this role, you will hold direct ownership over scaling distributed systems across massive hardware clusters while collaborating directly with core platform teams.

Responsibilities
  • Architect and sustain high-throughput execution and training pipelines optimized for rapid experimentation, flexibility, and ultra-low latency serving.
  • Establish comprehensive benchmarking suites across system stacks to identify performance bottlenecks and optimize resource efficiency.
  • Evaluate cutting-edge developments in model optimization and integrate state-of-the-art research into production systems.
  • Engineer resilient, highly observable multi-node environments that support seamless parallelized training tasks.
  • Maximize hardware efficiency and operational reliability while running large-scale workloads across extensive GPU clusters.
  • Build developer abstractions, internal tooling, and telemetry frameworks that streamline the transition from research prototype to production.
Preferred Qualifications (Nice-to-Have)
  • Prior experience working within an early-stage or rapidly scaling tech startup.
  • Meaningful open-source contributions to established distributed training or inference engines.
  • Practical experience managing multi-node inference setups across hundreds or thousands of GPUs.
  • Strong passion for aligning deep technical excellence with direct business results.
Compensation & Benefits
  • Work Arrangement: Full-time, 5 days per week on-site in San Francisco, CA.
  • Health & Wellness: Fully covered medical, dental, and vision coverage, plus a $150 monthly stipend for wellness and fitness expenses.
  • Time Off & Flexibility: Flexible PTO policy and adaptable parental leave programs.
  • Perks: Daily catered lunches in the office and full commuter reimbursement.
  • Equal Opportunity Employer: All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or protected veteran status.

Accepting Candidates

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000
Software Engineer, Data Infrastructure
Software Engineer, Data Infrastructure

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision insurance
Unlimited paid time off (PTO)
Paid parental leave
+1
ML Infra Engineer
ML Infra Engineer

Monograph • San Francisco (CA)

On-site
USD 120,000 - 160,000
AI Infrastructure Engineer, Model Serving Platform
AI Infrastructure Engineer, Model Serving Platform

Segment (Twilio) • San Francisco (CA)

On-site
USD 175,000 - 220,000
Comprehensive health coverage
Retirement benefits
Learning and development stipend
+2
Founding Machine Learning Infrastructure Engineer
Founding Machine Learning Infrastructure Engineer

Model AI • Palo Alto (CA)

On-site
USD 120,000 - 180,000
Research Engineer, Infrastructure, Inference
Research Engineer, Infrastructure, Inference

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Infrastructure Engineer
Infrastructure Engineer

David Joseph & Company • San Francisco (CA)

On-site
USD 150,000 - 250,000
Equity up to 1%
Visa sponsorship
SF hacker house housing & meals
Research Engineer, Infrastructure, Training Systems
Research Engineer, Infrastructure, Training Systems

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
ML Infra Architect – On-Site in SF
ML Infra Architect – On-Site in SF

OP Recruiting • San Francisco (CA)

On-site
USD 180,000 - 260,000
On-site work
Health insurance
Wellness stipend
+3