Lead Infrastructure Engineer

Electronic Arts

Vancouver

Hybrid

CAD 150,000 - 190,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Paid time off
Health insurance
Dental/Vision
Retirement plan
Life insurance
Disability insurance

Job summary

Electronic Arts is seeking a Lead Infrastructure Engineer to own the GPU fleet our researchers train on, including capacity, scheduling, diagnostics, and support. You will set technical direction for GPU operations and infrastructure architecture, and lead Infrastructure as Code setup, permissions, and debugging.

This is a hybrid role, working three days per week in Redwood City, Montreal, or Vancouver. You will report to the Head of Data and Infrastructure and collaborate with EA central infra

Qualifications

  • 8+ years of experience operating production infrastructure with deep AWS depth (EC2 GPU fleets, EKS, IAM and cross-account security, VPC and networking, S3/FSx).
  • Experience scheduling, diagnosing, and managing GPUs in AWS specifically.
  • Experience operating GPU fleets at 1000+ GPU scale.
  • Scripting and automation with Python, PowerShell, bash, or equivalent.
  • Infrastructure as code (Terraform or equivalent).
  • Familiarity with GPU scheduling/orchestration layers (Slurm, Kubernetes with Kueue/Volcano, Ray, dStack, SkyPilot).
  • Observability with Grafana/Prometheus or equivalent.

Responsibilities

  • Own GPU fleet operations across AWS estate.
  • Build the scheduling layer from zero.
  • Diagnose GPU and node failures and drive hardware remediation with AWS support.
  • Run researcher support as a product with office hours and self-service tooling.
  • Instrument the fleet with utilization, queue depth, job success rate, and cost per experiment metrics.
  • Partner with external compute and lab partnerships and EA central infra groups.
  • Author runbooks, decision records, and onboarding docs.

Skills

AWS GPU
GPU orchestration
Scripting
Terraform
Kubernetes
Observability
IAM & Security
Networking

Tools

Grafana
Prometheus
AWS IAM

Job description

  • Redwood City
  • California
  • United States of America
  • Montreal
  • Canada

Role ID

216124

Worker Type

Regular Employee

Studio/Department

Other

Work Model

Hybrid

Description & Requirements

Electronic Arts creates next-level entertainment experiences that inspire players and fans around the world. Here, everyone is part of the story. Part of a community that connects across the globe. A place where creativity thrives, new perspectives are invited, and ideas matter. A team where everyone makes play happen.

At EA, we believe games are powerful because they bring together multiple ways people engage: play, watch, create, and connect. And increasingly, the biggest entertainment platforms aren't just places to consume content - they're places where communities build.

Creator-made content is already a proven part of EA's history - from community creation tools in Battlefield to The Gallery in The Sims 4. We believe new creative technologies and tools will expand how players engage with and contribute to our experiences, supported by thoughtful product design, safety systems, and global reach. Our focus is on enabling more players to participate in creative expression by making creation easier, safer, and more rewarding.

As Lead Infrastructure Engineer, you will own the GPU fleet our researchers train on including capacity, scheduling, diagnostics, and support. You will set technical direction for GPU operations and infrastructure architecture. You will additionally lead Infrastructure as Code setup, granting permissions, and debugging infrastructure problems.

This is a hybrid role, working three days per week in Redwood City, Montreal, or Vancouver.

You will report to the Head of Data and Infrastructure.

Responsibilities:
  • You will own GPU fleet operations across our AWS estate.
  • You will build the scheduling layer from zero.
  • You will diagnose GPU and node failures fast and completely and drive hardware evidence and replacement through AWS support and capacity-block channels.
  • You will run researcher support as a first-class product including holding office hours, owning the support channel, and driving the recurring causes out of existence with self-service tooling, preflight checks, and documentation.
  • You will instrument the fleet including utilization, queue depth, job success rate, and cost per experiment metrics.
  • You will partner with our external compute and lab partnerships as a technical contact, and with EA's central infrastructure groups on shared services and escalation.
  • You will author runbooks, decision records, and onboarding docs.
Qualifications:
  • 8+ years of experience operating production infrastructure, with deep, current, hands-on AWS depth - EC2 GPU fleets, EKS, IAM and cross-account security, VPC and networking, S3 and FSx.
  • Experience scheduling, diagnosing, and managing GPUs in AWS specifically.
  • Experience operating GPU fleets at 1000+ GPU scale.
  • Expertise in scripting and automation with Python, PowerShell, bash, or equivalent.
  • Expertise in infrastructure as code (Terraform or equivalent).
  • Familiarity with a GPU scheduling or orchestration layer (like Slurm, Kubernetes with Kueue or Volcano, Ray, dStack or SkyPilot).
  • Observability practice including Grafana, Prometheus, or equivalent.

Pay Transparency - North America

COMPENSATION AND BENEFITS

  • paid time off (3 weeks per year to start)
  • 80 hours per year of sick time
  • 16 paid company holidays per year
  • 10 weeks paid time off to bond with baby
  • medical/dental/vision insurance
  • life insurance
  • disability insurance
  • 401(k)
  • vacation (3 weeks per year to start)
  • 10 days per year of sick time
  • paid top-up to EI/QPIP benefits up to 100% of base salary when you welcome a new child (12 weeks for maternity, and 4 weeks for parental/adoption leave)
  • extended health/dental/vision coverage
  • life insurance
  • disability insurance
  • retirement plan to regular full-time employees

PAY RANGES

Pay is just one part of the overall compensation at EA.

About Electronic Arts

We're proud to have an extensive portfolio of games and experiences, locations around the world, and opportunities across EA. We value adaptability, resilience, creativity, and curiosity. From leadership that brings out your potential, to creating space for learning and experimenting, we empower you to do great work and pursue opportunities for growth.

We adopt a holistic approach to our benefits programs, emphasizing physical, emotional, financial, career, and community wellness to support a balanced life. Our packages are tailored to meet local needs and may include healthcare coverage, mental well-being support, retirement savings, paid time off, family leaves, complimentary games, and more. We nurture environments where our teams can always bring their best to what they do.

Electronic Arts is an equal opportunity employer. All employment decisions are made without regard to race, color, national origin, ancestry, sex, gender, gender identity or expression, sexual orientation, age, genetic information, religion, disability, medical condition, pregnancy, marital status, family status, veteran status, or any other characteristic protected by law. We will also consider employment qualified applicants with criminal records in accordance with applicable law. EA also makes workplace accommodations for qualified individuals with disabilities as required by applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Développeur.se principal.e de l’infrastructure / Lead Infrastructure Developer
Développeur.se principal.e de l’infrastructure / Lead Infrastructure Developer

Electronic Arts • Vancouver, Montreal (administrative region)

Hybrid
CAD 170,000 - 243,000
Développeur.se principal.e de l’infrastructure / Lead Infrastructure Developer
Développeur.se principal.e de l’infrastructure / Lead Infrastructure Developer

Electronic Arts (EA) • Vancouver

Hybrid
CAD 170,000 - 243,000
Paid time off
Health insurance
Retirement plan
Développeur.se principal.e de l’infrastructure / Lead Infrastructure Developer
Développeur.se principal.e de l’infrastructure / Lead Infrastructure Developer

Electronic Arts (EA) • Montreal (administrative region)

Hybrid
CAD 170,000 - 243,000
Vacation: 3 weeks
Health benefits
Retirement plan
+1
Senior Software Engineer (Devops)
Senior Software Engineer (Devops)

Electronic Arts (EA) • Vancouver

Hybrid
CAD 122,000 - 171,000
Développeur.se - opérations liées à l’apprentissage automatique / MLOps Engineer
Développeur.se - opérations liées à l’apprentissage automatique / MLOps Engineer

Electronic Arts • Montreal (administrative region), Vancouver

Hybrid
CAD 141,000 - 204,000
Vacation 3 weeks (US/Canada)
Sick time 10 days
Top-up EI/QPIP
+3
Site Reliability Engineer
Site Reliability Engineer

Electronic Arts • Vancouver

On-site
CAD 105,000 - 143,000
DevOps Engineer - EA SPORTSTM TECHNOLOGY
DevOps Engineer - EA SPORTSTM TECHNOLOGY

Electronic Arts • Vancouver

On-site
CAD 122,000 - 171,000
Technical Program Manager II (temporary)
Technical Program Manager II (temporary)

Electronic Arts • Vancouver

Hybrid
CAD 92,000 - 130,000
3 weeks vacation
Extended health/dental/vision
Life insurance
Advanced Rendering Software Engineer - FC
Advanced Rendering Software Engineer - FC

EA SPORTS • Vancouver

On-site
CAD 122,000 - 171,000
Vacation 3 weeks
Health/dental/vision
Retirement plan
+1
Senior Software Engineer - Advanced Technology Group
Senior Software Engineer - Advanced Technology Group

EA SPORTS • Vancouver

On-site
CAD 141,000 - 204,000
Vacation time
Comprehensive health coverage
Retirement plan
+1