AI Infrastructure Engineer

Jobtailor

San Francisco (CA)

Hybrid

USD 140,000 - 210,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Jobtailor is seeking a highly capable ML Infrastructure Engineer to build and operate systems for large-scale model training and evaluation in our SF office. You will enhance reliability, throughput, and resource efficiency, develop shared inference platforms, optimize GPU scheduling, and collaborate with researchers and engineers to drive end-to-end performance.

Candidates should bring strong software fundamentals and experience with distributed systems; hybrid work from the US office is

Qualifications

  • Strong software engineering fundamentals.
  • Experience building or operating large-scale distributed systems.
  • Experience in ML infrastructure, inference systems, GPU performance, or infrastructure tooling.
  • Highly self-motivated and comfortable taking ownership of open-ended problems.
  • Enjoy debugging across system boundaries.
  • Use measurements to guide improvements in performance and reliability.
  • Ability to work from the US office three days per week.
  • Authorization to work in the country where the job is located.
  • Must disclose whether employment visa sponsorship is required.

Responsibilities

  • Build and operate infrastructure for large-scale training and evaluation
  • Improve reliability, throughput, and resource efficiency
  • Develop shared inference and grading platforms with automated capacity management, health monitoring, and performance visibility
  • Improve compute scheduling and resource allocation to reduce idle GPU time
  • Help workloads recover quickly from failures
  • Diagnose bottlenecks across training, inference, and orchestration
  • Work across teams to improve end-to-end performance
  • Build self-service tools, automated validation, and observability
  • Help researchers launch experiments, diagnose issues, and compare results with less manual intervention
  • Own projects from identifying bottlenecks and designing solutions through deployment and operation
  • Collaborate closely with researchers and engineering teams

Skills

Distributed Systems
ML Infrastructure
GPU Performance
Performance Measurement
Debugging Across Boundaries
Ownership
Self-Motivated
Collaboration
Compute Scheduling
Resource Allocation
Observability
Automated Validation
Capacity Management
Health Monitoring

Job description

  • Build and operate infrastructure for large-scale training and evaluation
  • Improve reliability, throughput, and resource efficiency
  • Develop shared inference and grading platforms with automated capacity management, health monitoring, and performance visibility
  • Improve compute scheduling and resource allocation to reduce idle GPU time
  • Help workloads recover quickly from failures
  • Diagnose bottlenecks across training, inference, and orchestration
  • Work across teams to improve end-to-end performance
  • Build self-service tools, automated validation, and observability
  • Help researchers launch experiments, diagnose issues, and compare results with less manual intervention
  • Own projects from identifying bottlenecks and designing solutions through deployment and operation
  • Collaborate closely with researchers and engineering teams
Requirements
  • Strong software engineering fundamentals
  • Experience building or operating large-scale distributed systems
  • Experience in ML infrastructure, inference systems, GPU performance, or infrastructure tooling
  • Highly self-motivated and comfortable taking ownership of open-ended problems
  • Enjoy debugging across system boundaries
  • Use measurements to guide improvements in performance and reliability
  • Ability to work from the US office three days per week
  • Authorization to work in the country where the job is located
  • Must disclose whether employment visa sponsorship is required
Core Competencies

Demonstrates expertise in building and operating large-scale distributed systems, with a focus on ML infrastructure and GPU performance optimization. Capable of diagnosing bottlenecks and improving system reliability and throughput through effective collaboration and ownership of projects.

Highest-signal resume keywords
  • Large-Scale Distributed Systems
  • ML Infrastructure
  • GPU Performance Optimization
  • Performance Measurement and Improvement
  • Debugging Across System Boundaries
Hard Skills
  • Software Engineering Fundamentals
  • Infrastructure Tooling
  • Capacity Management
  • Health Monitoring
  • Performance Visibility
  • Compute Scheduling
  • Resource Allocation
  • Automated Validation
  • Observability
  • Experiment Launching
Soft Skills
  • Self-Motivated
  • Ownership of Problems
  • Collaboration
  • Problem-Solving
Industry Keywords
  • Training and Evaluation Infrastructure
  • Resource Efficiency
  • Idle GPU Time Reduction
  • Workload Recovery
  • End-to-End Performance
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Software Engineer, AI Infrastructure – LVM Inference & Evaluation
Software Engineer, AI Infrastructure – LVM Inference & Evaluation

Jobtailor • Redwood City (CA)

On-site
USD 180,000 - 240,000
VP – AI Infrastructure Engineering
VP – AI Infrastructure Engineering

Jobtailor • Bellevue (WA)

On-site
USD 200,000 - 350,000
Principal ML Infrastructure Engineer (Relocation Available)
Principal ML Infrastructure Engineer (Relocation Available)

Franklin Fitch • Dallas (TX)

On-site
USD 100,000 - 140,000
AI Infrastructure Engineer
AI Infrastructure Engineer

Fuel Talent • Seattle (WA)

On-site
USD 180,000 - 210,000
Member of Technical Staff, AI Infrastructure
Member of Technical Staff, AI Infrastructure

Goaly • Menlo Park (CA), Northern (KY)

Hybrid
USD 150,000 - 180,000
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Calance • Costa Mesa (CA)

Hybrid
USD 180,000 - 240,000
Principal Infrastructure Engineer, AI Cluster Performance & Validation
Principal Infrastructure Engineer, AI Cluster Performance & Validation

Nscale • New York (NY), San Francisco (CA), Seattle (WA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000
Member of Technical Staff
Member of Technical Staff

Harrison Clarke • San Francisco (CA)

On-site
USD 180,000 - 280,000