AI Infrastructure Engineer — Scale ML Training & Inference

Triwill Group

San Francisco, Northern (CA, KY)

Hybrid

USD 140,000 - 190,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

OpenAI is seeking an AI Systems Engineer to scale infrastructure behind training and evaluation workflows. You will own projects from bottleneck identification to deployment, combining distributed systems engineering with performance tuning and collaboration with researchers.

You will build shared inference and grading platforms, improve scheduling, and develop observability tools to help researchers run experiments efficiently and reliably, with a strong culture of safety and human-centric

Qualifications

  • Strong foundations in software engineering and distributed systems.
  • Experience building or operating large-scale ML infrastructure.
  • Familiarity with GPU performance optimization and observability tooling.
  • Ability to own projects end-to-end from design to deployment.

Responsibilities

  • Build and operate infrastructure for large-scale training and evaluation, improving reliability, throughput, and resource efficiency.
  • Develop shared inference and grading platforms with automated capacity management, health monitoring, and visibility into performance.
  • Improve compute scheduling and resource allocation to reduce idle GPU time and speed up workload recovery from failures.
  • Diagnose bottlenecks across training, inference, and orchestration, coordinating with teams to boost end-to-end performance.
  • Build self-service tools and automated validation to empower researchers to launch experiments with less manual intervention.

Skills

Distributed systems
Performance optimization
ML infrastructure
Inference systems
GPU performance
Infrastructure tooling
Capacity management
Health monitoring
Compute scheduling
Resource allocation
Debugging
Observability

Job description

OpenAI is seeking an AI Systems Engineer to scale infrastructure behind training and evaluation workflows. You will own projects from bottleneck identification to deployment, combining distributed systems engineering with performance tuning and collaboration with researchers.

You will build shared inference and grading platforms, improve scheduling, and develop observability tools to help researchers run experiments efficiently and reliably, with a strong culture of safety and human-centric

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure Engineer: Scale Training & Systems
AI Infrastructure Engineer: Scale Training & Systems

Precision Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
AI Systems Engineer - Scalable Training Infra
AI Systems Engineer - Scalable Training Infra

OpenAI • San Francisco (CA)

On-site
USD 160,000 - 210,000
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
GenAI ML Systems Engineer: Scalable Training & Inference
GenAI ML Systems Engineer: Scalable Training & Inference

Meta • Menlo Park (CA)

On-site
USD 180,000 - 300,000
AI Inference Infrastructure Engineer
AI Inference Infrastructure Engineer

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, vision benefits
Unlimited PTO
Parental leave
+1
Senior Systems Engineer, AI Inference Platform
Senior Systems Engineer, AI Inference Platform

Slope • San Francisco (CA)

On-site
USD 180,000 - 260,000
Staff AI Platform Engineer: Scale ML Infra
Staff AI Platform Engineer: Scale ML Infra

DAT Freight Solutions • Seattle (WA)

Hybrid
USD 198,000 - 246,000
Medical Insurance
Dental Insurance
Vision Insurance
+6
Senior AI Infrastructure Engineer
Senior AI Infrastructure Engineer

AI Breaking Wire • Menlo Park (CA), Northern (KY)

Hybrid
USD 200,000 - 350,000
RSUs
Health benefits
Parental leave
+1
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
RL Training Infra Engineer: Scale, Debug, & Optimize
RL Training Infra Engineer: Scale, Debug, & Optimize

OpenAI • San Francisco (CA)

On-site
USD 120,000 - 160,000