AI Infrastructure Lead

Outsourceit

San Francisco (CA)

On-site

USD 120,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Outsourceit is seeking an AI Infrastructure Lead to design and operate innovative GPU infrastructure for enterprise AI workloads. This role requires a minimum commitment of 6 months and involves working closely with the CTO and a team of engineers.

The ideal candidate will have proven experience with large-scale GPU systems, distributed training, and containerised workloads, alongside expertise in AWS, GCP, or Azure services. Full-time availability and participation in an on-call rotation are required.

Qualifications

  • Proven experience designing and operating large-scale GPU infrastructure and model serving systems.
  • Deep knowledge of distributed training, inference optimisation, and containerised workloads.
  • Hands-on expertise with AWS, GCP, or Azure AI/ML services and Kubernetes.

Responsibilities

  • Own the design and operation of the GPU cluster management layer.
  • Lead the model serving pipeline and low-latency routing system.
  • Make architectural decisions that affect thousands of enterprise clients.

Skills

Large-scale GPU infrastructure design
Model serving systems
Distributed training
Inference optimisation
Containerised workloads
AWS services
GCP services
Azure AI/ML services
Kubernetes
Monitoring
Alerting
Incident response

Job description

Parallax Systems is building the next generation of inference infrastructure for enterprise AI workloads. We need an AI Infrastructure Lead to own the design and operation of our GPU cluster management layer, model serving pipeline, and low-latency routing system.

You will work directly with the CTO and a team of four senior engineers. This is a high-ownership role in a fast-moving environment — you will make architectural decisions that affect thousands of enterprise clients.

Core Requirements
  • Proven experience designing and operating large-scale GPU infrastructure and model serving systems.
  • Deep knowledge of distributed training, inference optimisation, and containerised workloads.
  • Hands‑on expertise with AWS, GCP, or Azure AI/ML services and Kubernetes.
  • Strong background in monitoring, alerting, and incident response for critical AI systems.

Rolling contract with 6‑month minimum commitment.

Work load 40 HRS/WK

Full‑time availability required. On‑call rotation included.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Lead - GPU Clusters & Model Serving
Senior AI Infrastructure Lead - GPU Clusters & Model Serving

Outsourceit • San Francisco (CA)

On-site
USD 120,000 - 170,000
Software Engineer, AI Infra
Software Engineer, AI Infra

Makers Fund • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Equity
Health benefits
Monthly stipends
+1
AI Infrastructure / ML Infrastructure Engineer
AI Infrastructure / ML Infrastructure Engineer

DeWinter Group • Campbell (CA)

On-site
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
Staff Software Engineer (AI Infrastructure)
Staff Software Engineer (AI Infrastructure)

DeepRec.ai • Palo Alto (CA)

On-site
USD 180,000 - 320,000
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Software Engineer, AI Infrastructure
Software Engineer, AI Infrastructure

Harell Data • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Customer Solution Architect - Systems Integrator
Customer Solution Architect - Systems Integrator

Hamilton Barnes Associates Limited • New York (NY)

On-site
USD 225,000 - 275,000
RSU equity
20% bonus