Lead Infrastructure Engineer - GPU Fleet & Cloud Ops

Electronic Arts (EA)

Montreal (administrative region)

Hybrid

CAD 170,000 - 243,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Vacation: 3 weeks
Health benefits
Retirement plan
Sick time: 10 days

Job summary

Electronic Arts seeks a Lead Infrastructure Engineer to own the GPU fleet used by researchers, set technical direction for GPU operations, and lead Infrastructure as Code initiatives. The role spans Montreal, Redwood City, or Vancouver, with a hybrid schedule and reporting to the Head of Data and Infrastructure.

The ideal candidate has 8+ years of production infra experience in AWS, plus expertise in GPU scaling, automation, and observability. This is a senior, hands-on leadership role.

Qualifications

  • 8+ years of production infrastructure experience with deep AWS depth.
  • Experience scheduling, diagnosing, and managing GPUs in AWS specifically.
  • Experience operating GPU fleets at 1000+ GPU scale.
  • Scripting and automation with Python, PowerShell, bash, or equivalent.
  • Infrastructure as code experience (Terraform or equivalent).
  • Familiarity with a GPU scheduling/orchestration layer (Slurm, Kubernetes with Kueue/Volcano, Ray, dStack or SkyPilot).
  • Observability practice including Grafana, Prometheus, or equivalent.

Responsibilities

  • Own GPU fleet operations across our AWS estate.
  • Build the scheduling layer from zero.
  • Diagnose GPU and node failures fast and completely and drive hardware evidence and replacement through AWS support and capacity-block channels.
  • Run researcher support as a first-class product including office hours, support channel, and self-service tooling.
  • Instrument the fleet including utilization, queue depth, job success rate, and cost per experiment metrics.
  • Partner with external compute and lab partnerships as a technical contact, and with EA's central infrastructure groups on shared services and escalation.
  • Author runbooks, decision records, and onboarding docs.

Skills

AWS GPU fleets
GPU orchestration
Python scripting
PowerShell scripting
Terraform
Grafana
Prometheus
Kubernetes

Tools

Slurm
Kueue
Volcano
Ray
dStack
SkyPilot

Job description

Electronic Arts seeks a Lead Infrastructure Engineer to own the GPU fleet used by researchers, set technical direction for GPU operations, and lead Infrastructure as Code initiatives. The role spans Montreal, Redwood City, or Vancouver, with a hybrid schedule and reporting to the Head of Data and Infrastructure.

The ideal candidate has 8+ years of production infra experience in AWS, plus expertise in GPU scaling, automation, and observability. This is a senior, hands-on leadership role.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead GPU Infrastructure Engineer - Hybrid & Cloud
Lead GPU Infrastructure Engineer - Hybrid & Cloud

Electronic Arts • Vancouver

Hybrid
CAD 150,000 - 190,000
Paid time off
Health insurance
Dental/Vision
+3
Lead Infrastructure Engineer
Lead Infrastructure Engineer

Electronic Arts • Vancouver

On-site
CAD 150,000 - 190,000
Paid time off
Health insurance
Dental/Vision
+3
Développeur.se principal.e de l’infrastructure / Lead Infrastructure Developer
Développeur.se principal.e de l’infrastructure / Lead Infrastructure Developer

Electronic Arts (EA) • Montreal (administrative region)

Hybrid
CAD 170,000 - 243,000
Vacation: 3 weeks
Health benefits
Retirement plan
+1
Développeur.se principal.e de l’infrastructure / Lead Infrastructure Developer
Développeur.se principal.e de l’infrastructure / Lead Infrastructure Developer

Electronic Arts (EA) • Vancouver

Hybrid
CAD 170,000 - 243,000
Paid time off
Health insurance
Retirement plan
Senior Backend Engineer – Scalable Cloud Services
Senior Backend Engineer – Scalable Cloud Services

Electronic Arts (EA) • Montreal (administrative region)

Hybrid
CAD 141,000 - 204,000
3 weeks vacation
Extended health/dental/vision
Remote Infrastructure/GPU Cluster/Platform Operations Lead
Remote Infrastructure/GPU Cluster/Platform Operations Lead

Bilinguallink • Brantford

Remote
CAD 140,000 - 190,000
Remote Infrastructure/GPU Cluster/Platform Operations Lead
Remote Infrastructure/GPU Cluster/Platform Operations Lead

Bilinguallink • Lower Sackville

Remote
CAD 120,000 - 160,000
Senior Site Reliability Engineer – Cloud & Automation
Senior Site Reliability Engineer – Cloud & Automation

Electronic Arts (EA) • Toronto

On-site
CAD 122,000 - 160,000
Senior Engineering Manager - AI-Driven, Multiteam Delivery
Senior Engineering Manager - AI-Driven, Multiteam Delivery

Kibbi • Vancouver

On-site
CAD 141,000 - 204,000
Vacation 3 weeks/year
Sick time 10 days/year
Health/dental/vision coverage
+1
Systems Software Engineer: Engine & AI Tools (Onsite Vancouver)
Systems Software Engineer: Engine & AI Tools (Onsite Vancouver)

Electronic Arts • Vancouver

On-site
CAD 105,000 - 143,000
3 weeks vacation
Sick time 10 days/yr
Health/dental/vision coverage
+1