Autonomous AI Infrastructure Engineer: GPU Fleet Mastery

Together

San Francisco (CA)

On-site

USD 190,000 - 270,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Startup equity
Competitive benefits

Job summary

Together AI is building and operating one of the world’s largest GPU fleets for frontier model training and inference in San Francisco. We seek engineers who treat infrastructure as software and automate everything to scale massively.

Responsibilities include designing fleet automation systems, AI Infrastructure Agents, and Fleet Intelligence to monitor, diagnose, and optimize performance. This role demands a strong automation mindset and collaboration across teams.

Qualifications

  • 3+ years building distributed systems, infrastructure platforms, or large-scale backend software.
  • Strong software engineering skills in Python, Go, or Rust.
  • Experience building platforms, automation systems, or developer infrastructure.
  • Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies.
  • Strong systems thinking with the ability to understand problems across hardware and software.
  • A passion for solving complex infrastructure challenges through software.
  • An automation-first mindset—if a task is repeated, your instinct is to build a system to eliminate it.

Responsibilities

  • Design and build fleet automation systems that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention.
  • Build AI Infrastructure Agents that automate deployment, root-cause failures, incident triage, and autonomous remediation.
  • Develop Fleet Intelligence platforms that monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures.
  • Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators.
  • Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads.
  • Build internal platforms and developer tools to manage infrastructure through software, not manual operations.
  • Continuously improve deployment velocity, reliability, and operational efficiency through automation.
  • Partner with hardware, networking, platform, and AI teams to push AI infrastructure limits.

Skills

Python
Go
Rust
Distributed systems

Tools

Kubernetes
Terraform
Ansible
Linux

Job description

Together AI is building and operating one of the world’s largest GPU fleets for frontier model training and inference in San Francisco. We seek engineers who treat infrastructure as software and automate everything to scale massively.

Responsibilities include designing fleet automation systems, AI Infrastructure Agents, and Fleet Intelligence to monitor, diagnose, and optimize performance. This role demands a strong automation mindset and collaboration across teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Fleet Infra Engineer — Scale, Automation & Kubernetes
GPU Fleet Infra Engineer — Scale, Automation & Kubernetes

OpenAI • New York (NY)

Hybrid
USD 180,000 - 240,000
Relocation assistance
Hybrid work model
Staff AI Infra Engineer: GPU Fleet Reliability Leader
Staff AI Infra Engineer: GPU Fleet Reliability Leader

Luma AI • United States

Remote
USD 210,000 - 320,000
Generative AI Infra Engineer (GPU Fleet)
Generative AI Infra Engineer (GPU Fleet)

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
Relocation assistance
Health, dental, and vision insurance (
Team events & offsites
+1
Senior AI Infrastructure Engineer — Scale GPU Clusters
Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
AI Infrastructure Architect — Scalable GPU Compute
AI Infrastructure Architect — Scalable GPU Compute

EngineersOfAI • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Senior AI Infra Architect for Multi-GPU Training + Equity
Senior AI Infra Architect for Multi-GPU Training + Equity

NVIDIA Gruppe • California (MO)

On-site
USD 224,000 - 431,000
Equity
Benefits
Software Engineer, Fleet Infrastructure
Software Engineer, Fleet Infrastructure

OpenAI • New York (NY)

Hybrid
USD 180,000 - 240,000
Relocation assistance
Hybrid work model
Infrastructure Engineer — AI Platform & GPU Fleet
Infrastructure Engineer — AI Platform & GPU Fleet

fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation assistance to San Francisco
Health, dental and vision insurance (U
Regular team events and offsites
ML Infra Engineer: GPU Fleet & Inference Orchestrator
ML Infra Engineer: GPU Fleet & Inference Orchestrator

Generalist • San Francisco (CA)

On-site
USD 120,000 - 160,000
AI Infrastructure & Automation Engineer
AI Infrastructure & Automation Engineer

Nscale • New York (NY)

On-site
USD 140,000 - 210,000
Competitive package
Equity
Growth opportunities