GPU Fleet Orchestrator for AI Infra

Advanced Micro Devices

San Jose (CA)

Hybrid

USD 150,000 - 210,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Advanced Micro Devices is seeking an experienced software engineer to build Fleet Manager, a secure control plane for operating large-scale AMD GPU infrastructure.

You will design and develop distributed services spanning Kubernetes, GPU scheduling, inference infrastructure, and developer‑facing APIs and tools. The role combines systems software, cloud infrastructure, and accelerated computing.

Qualifications

  • Strong systems‑software development experience in Rust, C++, Go or a comparable language.
  • Experience designing and operating distributed systems, control planes, schedulers, or cloud infrastructure.
  • Strong understanding of concurrency, asynchronous programming, state machines, and failure recovery.
  • Experience building reliable services using REST, streaming, WebSocket, or gRPC APIs.

Responsibilities

  • Design and develop Fleet Manager’s distributed control‑plane services, APIs, schedulers, inference gateway, and command‑line tools.
  • Build reliable orchestration for GPU training, inference, custom jobs, and interactive development workloads.
  • Develop scalable scheduling and admission‑control capabilities, including priority, fairness, topology‑aware placement, quotas, backfilling, and multi‑node workload coordination.
  • Develop durable reconciliation, lifecycle management, retries, idempotency, and recovery across PostgreSQL and external execution systems.
  • Integrate Fleet Manager with Kubernetes and technologies such as Kueue, JobSet, container runtimes, storage systems, and observability platforms.

Skills

Rust
C++
Go
Distributed systems
Kubernetes
PostgreSQL
REST APIs
gRPC
Security
Multi-tenant

Education

Bachelor’s or Master’s degree in Computer Science / Engineering

Tools

Kubernetes
Kueue
JobSet
PostgreSQL

Job description

Advanced Micro Devices is seeking an experienced software engineer to build Fleet Manager, a secure control plane for operating large-scale AMD GPU infrastructure.

You will design and develop distributed services spanning Kubernetes, GPU scheduling, inference infrastructure, and developer‑facing APIs and tools. The role combines systems software, cloud infrastructure, and accelerated computing.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Fleet Orchestrator for AI Infrastructure
GPU Fleet Orchestrator for AI Infrastructure

AMD • San Jose (CA)

Hybrid
USD 180,000 - 240,000
GPU Fleet Infra Engineer — Scale, Automation & Kubernetes
GPU Fleet Infra Engineer — Scale, Automation & Kubernetes

OpenAI • New York (NY)

Hybrid
USD 180,000 - 240,000
Relocation assistance
Hybrid work model
Autonomous AI Infrastructure Engineer: GPU Fleet Mastery
Autonomous AI Infrastructure Engineer: GPU Fleet Mastery

Together • San Francisco (CA)

On-site
USD 190,000 - 270,000
Health insurance
Startup equity
Competitive benefits
AI Infrastructure Engineer: GPU Fleet & Automation
AI Infrastructure Engineer: GPU Fleet & Automation

fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation assistance to San Francisco
Health, dental, and vision insurance (
Team events and offsites
Generative AI Infra Engineer (GPU Fleet)
Generative AI Infra Engineer (GPU Fleet)

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
Relocation assistance
Health, dental, and vision insurance (
Team events & offsites
+1
Software Development Engineer — GPU Fleet Management & AI Infrastructure
Software Development Engineer — GPU Fleet Management & AI Infrastructure

AMD • San Jose (CA)

Hybrid
USD 180,000 - 240,000
Engineering Manager, AI Cloud GPU Fleet
Engineering Manager, AI Cloud GPU Fleet

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
ML Infra Engineer: GPU Fleet & Inference Orchestrator
ML Infra Engineer: GPU Fleet & Inference Orchestrator

Generalist • San Francisco (CA)

On-site
USD 120,000 - 160,000
AI Infra Engineer: GPU Fleet Automation
AI Infra Engineer: GPU Fleet Automation

Fal.ai Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
Relocation assistance
Health, dental, and vision insurance (
Regular team events and offsites
GPU Fleet Automation Engineer (Hybrid)
GPU Fleet Automation Engineer (Hybrid)

Tower Research Capital • New York (NY)

Hybrid
USD 200,000 - 300,000
Generous PTO
Hybrid work
Free meals