Senior AI Compute Infrastructure Architect

Arm Limited

Seattle (WA)

Hybrid

USD 263,000 - 355,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Relocation package
Visa sponsorship

Job summary

Arm is seeking a Principal Engineer on the AI Compute Infra team to design, build, and operate large-scale AI infrastructure for training, fine-tuning, evaluation, and inference. You will lead Kubernetes clusters, accelerator enablement, and workload scheduling alongside researchers and engineers.

You will improve reliability, performance, scalability and developer productivity, driving end-to-end operational excellence and impactful outcomes in a collaborative, fast-paced setting.

Qualifications

  • 8+ years of experience building or operating cloud, compute, HPC, or distributed infrastructure in production.
  • Programming in Go or Python or another systems language with focus on reliable infra software.
  • Practical knowledge of Kubernetes, containers, Linux, networking, and storage.
  • Experience supporting GPU, accelerator, or distributed ML workloads.
  • Strong troubleshooting and cross-team communication skills.

Responsibilities

  • Build and operate Kubernetes clusters with improved scheduling and capacity use.
  • Enable new CPU/GPU systems by integrating drivers, networking, storage, and health checks.
  • Investigate performance and reliability issues and turn findings into improvements.
  • Partner with AI teams to automate provisioning, upgrades, monitoring, and maintenance.

Skills

Go
Python
Systems language
Kubernetes
Linux
Networking
Storage
GPU/accelerators
Troubleshooting

Tools

NVIDIA CUDA
Prometheus
Grafana
Terraform
Argo CD
Helm
AWS EKS

Job description

Arm is seeking a Principal Engineer on the AI Compute Infra team to design, build, and operate large-scale AI infrastructure for training, fine-tuning, evaluation, and inference. You will lead Kubernetes clusters, accelerator enablement, and workload scheduling alongside researchers and engineers.

You will improve reliability, performance, scalability and developer productivity, driving end-to-end operational excellence and impactful outcomes in a collaborative, fast-paced setting.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal AI Inference Cloud Architect
Principal AI Inference Cloud Architect

Arm Limited • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Senior Cloud AI Inference Engineer
Senior Cloud AI Inference Engineer

Arm Limited • Seattle (WA)

Hybrid
USD 209,000 - 283,000
Relocation package
Visa sponsorship
Principal AI Inference Runtime Architect
Principal AI Inference Runtime Architect

Arm Limited • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Hybrid working
Accommodations during recruitment
Senior AI Infrastructure Engineer
Senior AI Infrastructure Engineer

AI Breaking Wire • Menlo Park (CA), Northern (KY)

Hybrid
USD 200,000 - 350,000
RSUs
Health benefits
Parental leave
+1
Principal Software Engineer, AI Compute Infrastructure
Principal Software Engineer, AI Compute Infrastructure

Arm Limited • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Relocation package
Visa sponsorship
Senior Staff AI Infra Engineer — Lead Scalable Compute
Senior Staff AI Infra Engineer — Lead Scalable Compute

Anthropic • San Francisco (CA)

Hybrid
USD 405,000 - 485,000
AI Infrastructure Architect — Scalable GPU Compute
AI Infrastructure Architect — Scalable GPU Compute

EngineersOfAI • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Senior AI Infra Engineer: Large-Scale GPU & HPC
Senior AI Infra Engineer: Large-Scale GPU & HPC

Anduril Industries • Costa Mesa (CA)

On-site
USD 166,000 - 220,000
Equity grants
Benefits package
Senior AI Inference Runtime Engineer - Distributed
Senior AI Inference Runtime Engineer - Distributed

Arm Limited • Seattle (WA)

Hybrid
USD 209,000 - 283,000
Senior AI Infra Engineer: Kubernetes, GPUs & Scale
Senior AI Infra Engineer: Kubernetes, GPUs & Scale

Seekr • Reston (VA)

Hybrid
USD 180,000 - 240,000
Unlimited PTO
14 paid company holidays
Hybrid work offices: Reston, VA &
+3