GPU Fleet Orchestrator for AI Infrastructure

AMD

San Jose (CA)

Hybrid

USD 180,000 - 240,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

AMD is seeking an experienced software engineer to build Fleet Manager, a secure control plane for operating large-scale GPU infrastructure. You will design production systems spanning distributed control planes, Kubernetes, GPU scheduling, and developer-facing APIs and tools.

The role sits at the intersection of systems software, cloud infrastructure, and accelerated computing. You will influence how engineers train, serve, debug, and operate workloads on AMD GPUs.

Qualifications

  • Strong systems-software development in Rust, C++, or Go.
  • Experience designing and operating distributed systems or cloud infrastructure.
  • Strong understanding of concurrency, state machines, and failure recovery.
  • Experience building reliable services with REST, streaming, WebSocket, or gRPC APIs.
  • Familiarity with Kubernetes internals, controllers, or scheduling.
  • Knowledge of batch schedulers like Kueue, Slurm, or JobSet.

Responsibilities

  • Design and develop Fleet Manager’s distributed control-plane services, APIs, schedulers, and tools.
  • Build orchestration for GPU training, inference, and interactive workloads.
  • Develop scalable scheduling, quotas, and multi-node workload coordination.
  • Implement durable reconciliation, retries, and recovery across PostgreSQL and external systems.
  • Integrate Fleet Manager with Kubernetes and related runtimes and storage.
  • Evolve Fleet Manager toward multi-environment support (Kubernetes, Slurm, Spur).
  • Develop GPU health, diagnostics, and controlled remediation features.
  • Build secure multi-tenant infrastructure with strong authentication and least-privilege defaults.
  • Improve AI inference reliability, including routing, streaming, and metering.
  • Define and maintain stable APIs, data models, and operational procedures.
  • Diagnose complex failures across distributed services, GPUs, and networks.
  • Develop automated tests and production-readiness checks.
  • Collaborate with architecture, driver, security, and ML teams on current/future GPUs.
  • Participate in bring-up of new GPU, system, cluster stacks.
  • Provide technical leadership through design reviews and mentoring.

Skills

Rust
C++
Go
Distributed systems
APIs

Education

Bachelor’s or Master’s in CS/CE

Tools

Kubernetes
PostgreSQL
REST
gRPC

Job description

AMD is seeking an experienced software engineer to build Fleet Manager, a secure control plane for operating large-scale GPU infrastructure. You will design production systems spanning distributed control planes, Kubernetes, GPU scheduling, and developer-facing APIs and tools.

The role sits at the intersection of systems software, cloud infrastructure, and accelerated computing. You will influence how engineers train, serve, debug, and operate workloads on AMD GPUs.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Fleet Orchestrator for AI Infra
GPU Fleet Orchestrator for AI Infra

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 150,000 - 210,000
AI Infrastructure Engineer: GPU Fleet & Automation
AI Infrastructure Engineer: GPU Fleet & Automation

fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation assistance to San Francisco
Health, dental, and vision insurance (
Team events and offsites
Autonomous AI Infrastructure Engineer: GPU Fleet Mastery
Autonomous AI Infrastructure Engineer: GPU Fleet Mastery

Together • San Francisco (CA)

On-site
USD 190,000 - 270,000
Health insurance
Startup equity
Competitive benefits
Software Development Engineer — GPU Fleet Management & AI Infrastructure
Software Development Engineer — GPU Fleet Management & AI Infrastructure

AMD • San Jose (CA)

Hybrid
USD 180,000 - 240,000
GPU Fleet Infra Engineer — Scale, Automation & Kubernetes
GPU Fleet Infra Engineer — Scale, Automation & Kubernetes

OpenAI • New York (NY)

Hybrid
USD 180,000 - 240,000
Relocation assistance
Hybrid work model
Generative AI Infra Engineer (GPU Fleet)
Generative AI Infra Engineer (GPU Fleet)

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
Relocation assistance
Health, dental, and vision insurance (
Team events & offsites
+1
Engineering Manager, AI Cloud GPU Fleet
Engineering Manager, AI Cloud GPU Fleet

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
GPU Fleet Automation Engineer (Hybrid)
GPU Fleet Automation Engineer (Hybrid)

Tower Research Capital • New York (NY)

Hybrid
USD 200,000 - 300,000
Generous PTO
Hybrid work
Free meals
AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 180,000 - 260,000
GPU AI Platform Engineer - Kubernetes & DevOps
GPU AI Platform Engineer - Kubernetes & DevOps

AMD • San Jose (CA)

On-site
USD 140,000 - 170,000
AMD benefits