Senior AI Compute Infra Engineer (Hybrid)

Arm Limited

Seattle (WA)

Hybrid

USD 209,000 - 283,000

Full time

13 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Relocation package
Recruitment accommodations

Job summary

Arm Limited is seeking an experienced AI Compute Infra engineer to design, build, and operate large-scale infrastructure for AI training, evaluation, and inference. You will work across Kubernetes clusters, accelerator enablement, workload scheduling, and high-performance networking, partnering with AI researchers to improve reliability and scalability.

The role involves supporting GPU and accelerator workloads, integrating drivers, and automating provisioning and monitoring.

Qualifications

  • 5+ years of experience building or operating cloud, compute, HPC, or distributed infrastructure in a production environment.
  • Programming experience in Go, Python, or another systems language, with an interest in developing reliable infrastructure software.
  • Practical knowledge of Kubernetes, containers, Linux, networking, and storage.
  • Experience supporting GPU, accelerator, or distributed machine-learning workloads.
  • An ability to troubleshoot complex systems and communicate clearly with engineers from different technical backgrounds.

Responsibilities

  • Build and operate Kubernetes clusters while improving workload scheduling, topology-aware placement, capacity use, and recovery.
  • Enable new CPU and GPU systems by integrating and validating drivers, networking, storage, monitoring, and health checks.
  • Investigate performance and reliability issues across applications, cloud infrastructure, clusters, and hardware, then turn findings into lasting improvements.
  • Partner with AI teams to understand their workloads and automate cluster provisioning, upgrades, monitoring, and maintenance around their needs.

Skills

Kubernetes
Containers
Linux
Networking
Storage
Go
Python
Distributed infra

Tools

NVIDIA CUDA
NVLink
NCCL
EFA
DCGM
AWS EKS
Terraform
Argo CD
Helm
Prometheus
Grafana

Job description

Arm Limited is seeking an experienced AI Compute Infra engineer to design, build, and operate large-scale infrastructure for AI training, evaluation, and inference. You will work across Kubernetes clusters, accelerator enablement, workload scheduling, and high-performance networking, partnering with AI researchers to improve reliability and scalability.

The role involves supporting GPU and accelerator workloads, integrating drivers, and automating provisioning and monitoring.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Compute Infra Engineer
Senior AI Compute Infra Engineer

Arm • Seattle (WA)

Hybrid
USD 209,000 - 283,000
Senior AI Compute Infrastructure Architect
Senior AI Compute Infrastructure Architect

Arm Limited • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Relocation package
Visa sponsorship
Lead AI Compute Infra Engineer — Relocation & Visa
Lead AI Compute Infra Engineer — Relocation & Visa

Arm • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Senior AI Infra Engineer: GPU Scale & Resilience (Hybrid)
Senior AI Infra Engineer: GPU Scale & Resilience (Hybrid)

Calance • Costa Mesa (CA)

Hybrid
USD 180,000 - 240,000
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Staff Software Engineer, AI Compute Infrastructure
Staff Software Engineer, AI Compute Infrastructure

Arm • Seattle (WA)

Hybrid
USD 209,000 - 283,000
Senior AI Compute Platform Engineer (Kubernetes)
Senior AI Compute Platform Engineer (Kubernetes)

Arm Limited • Seattle (WA)

Hybrid
USD 209,000 - 283,000
AI Infrastructure Architect — Scalable GPU Compute
AI Infrastructure Architect — Scalable GPU Compute

EngineersOfAI • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Senior AI Infrastructure Engineer — Scale GPU Clusters
Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
Senior GPU Infra Architect for Scalable AI Compute
Senior GPU Infra Architect for Scalable AI Compute

AI Chopping Block • Costa Mesa (CA), Northern (KY)

Hybrid
USD 166,000 - 220,000