Senior AI Infrastructure Engineer | Scale GPU Clusters

Fuel Talent LLC

Seattle (WA)

Hybrid

USD 126,000 - 189,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Fuel Talent LLC in Seattle, WA is seeking a Senior Software Engineer to design and operate the infrastructure powering large-scale AI training. You will build and optimize orchestration and resource scheduling across GPU clusters, collaborating with researchers and engineers to scale performance and reliability.

The role combines hands-on software development with systems engineering, requiring Go and/or Python expertise, strong Linux knowledge, and experience with containers.

Qualifications

  • 8+ years of experience building business-critical software and running large-scale compute infrastructure.
  • Proficient in Go and/or Python and Linux systems.
  • Experience with container technologies such as Docker and orchestration tools like Kubernetes.
  • Strong distributed systems design, debugging, and performance optimization.

Responsibilities

  • Design and deliver infrastructure for large-scale AI training workloads.
  • Build and improve workload scheduling, orchestration, and execution systems.
  • Automate infrastructure management and reduce manual overhead.
  • Collaborate with researchers and engineers to optimize GPU resource utilization.

Skills

Go
Python
Linux
Docker
Distributed systems
GPU compute

Education

Bachelor's degree in Computer Science or equivalent

Tools

Kubernetes
Slurm
NCCL
InfiniBand
WEKA
Ceph

Job description

Fuel Talent LLC in Seattle, WA is seeking a Senior Software Engineer to design and operate the infrastructure powering large-scale AI training. You will build and optimize orchestration and resource scheduling across GPU clusters, collaborating with researchers and engineers to scale performance and reliability.

The role combines hands-on software development with systems engineering, requiring Go and/or Python expertise, strong Linux knowledge, and experience with containers.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 160,000
GPU Infrastructure Engineer for Scalable AI/ML Platform
GPU Infrastructure Engineer for Scalable AI/ML Platform

RXinsider LTD. • Seattle (WA)

On-site
USD 115,000 - 235,000
Medical insurance
Dental insurance
Vision insurance
+2
Hybrid Staff AI Training Infrastructure Engineer
Hybrid Staff AI Training Infrastructure Engineer

Designworks Talent LLC • Bellevue (KY)

Hybrid
USD 150,000 - 210,000
Medical, dental, vision insurance
401(k) with company match
Paid holidays
Senior AI Infra Engineer — HPC & Scheduler
Senior AI Infra Engineer — HPC & Scheduler

Ai2 • Seattle (WA)

On-site
USD 126,000 - 189,000
Medical, dental, and vision insurance
401(k) plan enrollment
Monthly stipends for commuting and fitness
+1
Senior AI Infrastructure Lead - GPU Clusters & Model Serving
Senior AI Infrastructure Lead - GPU Clusters & Model Serving

Outsourceit • San Francisco (CA)

On-site
USD 120,000 - 170,000
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior AI Infrastructure Architect (GPU & Serving)
Senior AI Infrastructure Architect (GPU & Serving)

Makers Fund • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Equity
Health benefits
Monthly stipends
+1
AI Infrastructure Architect: Kubernetes & GPU Scaling
AI Infrastructure Architect: Kubernetes & GPU Scaling

NVIDIA • United States

Remote
USD 272,000 - 431,000
AI Infrastructure Architect — Scalable GPU Compute
AI Infrastructure Architect — Scalable GPU Compute

EngineersOfAI • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Senior GPU Systems Engineer: Scale AI Clusters & HPC
Senior GPU Systems Engineer: Scale AI Clusters & HPC

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000