Lead GPU Infrastructure Engineer (Hybrid; AWS, IaC, Scheduling)

Electronic Arts (EA)

Redwood City (CA)

Hybrid

USD 193,000 - 297,000

Full time

10 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Paid time off
Sick time
401(k)

Job summary

Electronic Arts is seeking a Lead Infrastructure Engineer to own the GPU fleet and set technical direction for GPU operations. This hybrid role requires three days per week in Redwood City, Montreal, or Vancouver, reporting to the Head of Data and Infrastructure.

You will build the scheduling layer from zero, diagnose GPU and node failures, and drive capacity planning with AWS support. You will also create runbooks and onboarding docs to enable scalable researcher support.

Qualifications

  • 8+ years of experience operating production infrastructure with deep AWS depth.
  • Experience scheduling, diagnosing, and managing GPUs in AWS.
  • Experience operating GPU fleets at 1000+ GPU scale.
  • Scripting and automation with Python, PowerShell, bash, or equivalent.
  • Infrastructure as code (Terraform or equivalent).
  • Familiarity with GPU scheduling/orchestration (Slurm, Kubernetes with Kueue or Volcano, Ray, dStack or SkyPilot).
  • Observability practice including Grafana, Prometheus, or equivalent.

Responsibilities

  • Own GPU fleet operations across the AWS estate.
  • Build the scheduling layer from zero.
  • Diagnose GPU and node failures fast and drive hardware replacement through AWS support and capacity-block channels.
  • Run researcher support as a first-class product including holding office hours, owning the support channel, and driving recurring causes out of existence with self-service tooling, preflight checks, and documentation.
  • Instrument the fleet including utilization, queue depth, job success rate, and cost per experiment metrics.
  • Partner with external compute and lab partnerships as a technical contact, and with EA's central infrastructure groups on shared services and escalation.
  • Author runbooks, decision records, and onboarding docs.

Skills

AWS GPU management
Python scripting
IaC with Terraform
GPU scheduling/orchestration
Observability (Grafana/Prometheus)
Kubernetes

Tools

Terraform
Grafana
Prometheus
Kubernetes
Slurm
Ray

Job description

Electronic Arts is seeking a Lead Infrastructure Engineer to own the GPU fleet and set technical direction for GPU operations. This hybrid role requires three days per week in Redwood City, Montreal, or Vancouver, reporting to the Head of Data and Infrastructure.

You will build the scheduling layer from zero, diagnose GPU and node failures, and drive capacity planning with AWS support. You will also create runbooks and onboarding docs to enable scalable researcher support.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Infrastructure Engineer
Lead Infrastructure Engineer

Electronic Arts (EA) • Redwood City (CA)

On-site
USD 193,000 - 297,000
Paid time off
Sick time
401(k)
GPU Infrastructure Engineer | Datacenter Ops
GPU Infrastructure Engineer | Datacenter Ops

Prime Intellect • United States

Remote
USD 150,000 - 300,000
Senior GPU Data Center Engineer
Senior GPU Data Center Engineer

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Lead GPU Cloud Fleet Engineering
Lead GPU Cloud Fleet Engineering

Lambda Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Health, dental, and vision coverage
401k with 2% company match (USA)
Flexible paid time off
GPU Fleet Automation Engineer (Hybrid)
GPU Fleet Automation Engineer (Hybrid)

Tower Research Capital • New York (NY)

Hybrid
USD 200,000 - 300,000
Generous PTO
Hybrid work
Free meals
Senior InfraOps Engineer — GPU Infra, Hybrid, Equity
Senior InfraOps Engineer — GPU Infra, Hybrid, Equity

Lightning AI • New York (NY)

Hybrid
USD 160,000 - 200,000
Health coverage
Equity/RSUs
401(k) matching
+1
Senior GPU Infrastructure Support Engineer
Senior GPU Infrastructure Support Engineer

Nscale • San Francisco (CA)

On-site
USD 120,000 - 170,000
Equity
Remote-friendly team
Flexible workplace
GPU Cloud Infrastructure Engineer
GPU Cloud Infrastructure Engineer

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
GPU Infrastructure Operations Lead
GPU Infrastructure Operations Lead

Primeintellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior GPU Infra Engineer — Remote
Senior GPU Infra Engineer — Remote

Nscale • Seattle (WA)

On-site
USD 120,000 - 170,000
Remote-first culture
Equity plan
Flexible workplace