ML Infra Engineer: Scale, GPU Performance & Reliability

Jobtailor

San Francisco (CA)

Hybrid

USD 140,000 - 210,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Jobtailor is seeking a highly capable ML Infrastructure Engineer to build and operate systems for large-scale model training and evaluation in our SF office. You will enhance reliability, throughput, and resource efficiency, develop shared inference platforms, optimize GPU scheduling, and collaborate with researchers and engineers to drive end-to-end performance.

Candidates should bring strong software fundamentals and experience with distributed systems; hybrid work from the US office is

Qualifications

  • Strong software engineering fundamentals.
  • Experience building or operating large-scale distributed systems.
  • Experience in ML infrastructure, inference systems, GPU performance, or infrastructure tooling.
  • Highly self-motivated and comfortable taking ownership of open-ended problems.
  • Enjoy debugging across system boundaries.
  • Use measurements to guide improvements in performance and reliability.
  • Ability to work from the US office three days per week.
  • Authorization to work in the country where the job is located.
  • Must disclose whether employment visa sponsorship is required.

Responsibilities

  • Build and operate infrastructure for large-scale training and evaluation
  • Improve reliability, throughput, and resource efficiency
  • Develop shared inference and grading platforms with automated capacity management, health monitoring, and performance visibility
  • Improve compute scheduling and resource allocation to reduce idle GPU time
  • Help workloads recover quickly from failures
  • Diagnose bottlenecks across training, inference, and orchestration
  • Work across teams to improve end-to-end performance
  • Build self-service tools, automated validation, and observability
  • Help researchers launch experiments, diagnose issues, and compare results with less manual intervention
  • Own projects from identifying bottlenecks and designing solutions through deployment and operation
  • Collaborate closely with researchers and engineering teams

Skills

Distributed Systems
ML Infrastructure
GPU Performance
Performance Measurement
Debugging Across Boundaries
Ownership
Self-Motivated
Collaboration
Compute Scheduling
Resource Allocation
Observability
Automated Validation
Capacity Management
Health Monitoring

Job description

Jobtailor is seeking a highly capable ML Infrastructure Engineer to build and operate systems for large-scale model training and evaluation in our SF office. You will enhance reliability, throughput, and resource efficiency, develop shared inference platforms, optimize GPU scheduling, and collaborate with researchers and engineers to drive end-to-end performance.

Candidates should bring strong software fundamentals and experience with distributed systems; hybrid work from the US office is

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infra Engineer: Scale & Optimize Large-Scale Training
ML Infra Engineer: Scale & Optimize Large-Scale Training

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Platform Engineer — Scale GPU-Driven Research Infra
ML Platform Engineer — Scale GPU-Driven Research Infra

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
ML Inference Infrastructure Engineer — Scale & GPU
ML Inference Infrastructure Engineer — Scale & GPU

Talanto • Northern (KY)

Hybrid
USD 221,000 - 260,000
Generous Time Off
Comprehensive Health Plans
Paid Parental Leave
+8
ML Platform Engineer: Build GPU-Scale Infra & Research
ML Platform Engineer: Build GPU-Scale Infra & Research

Triwill Group • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
ML Infrastructure Engineer: Scale & Performance
ML Infrastructure Engineer: Scale & Performance

Physical Intelligence • San Francisco (CA)

On-site
USD 150,000 - 230,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
ML Infra Engineer — GPU Clusters & Distributed Systems
ML Infra Engineer — GPU Clusters & Distributed Systems

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
+1
ML Infra Engineer — Scale GPU ML Platform & Equity
ML Infra Engineer — Scale GPU ML Platform & Equity

Socket.dev • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+5
ML Infra Engineer: Scale Training & Inference (Hybrid)
ML Infra Engineer: Scale Training & Inference (Hybrid)

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1