Senior AI Infrastructure Lead - GPU Clusters & Model Serving

Outsourceit

San Francisco (CA)

On-site

USD 120,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Outsourceit is seeking an AI Infrastructure Lead to design and operate innovative GPU infrastructure for enterprise AI workloads. This role requires a minimum commitment of 6 months and involves working closely with the CTO and a team of engineers.

The ideal candidate will have proven experience with large-scale GPU systems, distributed training, and containerised workloads, alongside expertise in AWS, GCP, or Azure services. Full-time availability and participation in an on-call rotation are required.

Qualifications

  • Proven experience designing and operating large-scale GPU infrastructure and model serving systems.
  • Deep knowledge of distributed training, inference optimisation, and containerised workloads.
  • Hands-on expertise with AWS, GCP, or Azure AI/ML services and Kubernetes.

Responsibilities

  • Own the design and operation of the GPU cluster management layer.
  • Lead the model serving pipeline and low-latency routing system.
  • Make architectural decisions that affect thousands of enterprise clients.

Skills

Large-scale GPU infrastructure design
Model serving systems
Distributed training
Inference optimisation
Containerised workloads
AWS services
GCP services
Azure AI/ML services
Kubernetes
Monitoring
Alerting
Incident response

Job description

Outsourceit is seeking an AI Infrastructure Lead to design and operate innovative GPU infrastructure for enterprise AI workloads. This role requires a minimum commitment of 6 months and involves working closely with the CTO and a team of engineers.

The ideal candidate will have proven experience with large-scale GPU systems, distributed training, and containerised workloads, alongside expertise in AWS, GCP, or Azure services. Full-time availability and participation in an on-call rotation are required.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure Lead
AI Infrastructure Lead

Outsourceit • San Francisco (CA)

On-site
USD 120,000 - 170,000
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
AI Infrastructure Architect — Scalable GPU Compute
AI Infrastructure Architect — Scalable GPU Compute

EngineersOfAI • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Senior AI Infrastructure Architect (GPU & Serving)
Senior AI Infrastructure Architect (GPU & Serving)

Makers Fund • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Equity
Health benefits
Monthly stipends
+1
Remote Engineering Manager, AI GPU Infrastructure
Remote Engineering Manager, AI GPU Infrastructure

5C • United States

On-site
USD 180,000 - 220,000
Senior AI Training Infra Engineer - Scale GPU Clusters
Senior AI Training Infra Engineer - Scale GPU Clusters

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Medical insurance
401(k) with company match
Paid holidays
Senior AI Infrastructure Engineer - GPU Compute
Senior AI Infrastructure Engineer - GPU Compute

Unchain Data • United States

On-site
USD 120,000 - 160,000
Staff Compute Infra Engineer - GPU & AI Systems
Staff Compute Infra Engineer - GPU & AI Systems

xAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Senior AI Infrastructure Engineer | Scale GPU Clusters
Senior AI Infrastructure Engineer | Scale GPU Clusters

Fuel Talent LLC • Seattle (WA)

Hybrid
USD 126,000 - 189,000
Senior AI Infrastructure Engineer — Scale GPU Clusters
Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2