Senior AI Infrastructure Engineer

KLA-Belgium

Ann Arbor (MI)

On-site

USD 130,000 - 190,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Medical
Dental
Vision
401(K) including company matching
ESPP

Job summary

KLA in the United States, Michigan (Ann Arbor) is seeking a Sr. Software Engineer to design, build, and operate production-grade AI infrastructure for large-scale GPU training and inference. You will own cluster architecture, multi-node orchestration, and performance engineering in production environments.

The role emphasizes Kubernetes, distributed systems, and open-source collaboration, with a strong emphasis on Python/Linux skills and capacity to deploy multi-node AI platforms.

Qualifications

  • Bachelor's Degree and eight (8) years of Software Engineering experience.
  • Four (4) years in software, cloud, platform, HPC, or infrastructure engineering, including two (2) years supporting distributed AI/ML workloads.
  • Proven hands-on experience building GPU clusters from the ground up and operating them at production scale.
  • Deep Kubernetes expertise, including operators, CRDs, Helm, networking, storage, scheduling, and cluster lifecycle management.
  • Production experience with Ray and at least two of the following: NVIDIA NeMo RL, vLLM, SGLang, or NVIDIA Dynamo.
  • Strong knowledge of NVIDIA GPUs, CUDA, NCCL, GPU Operator, DCGM, MIG, RDMA, and multi-node collective communication.
  • Experience debugging training and inference performance across GPU, CPU, memory, network, and storage layers.
  • Proficiency in Python and Linux, plus experience with containers, CI/CD, GitOps, infrastructure-as-code, and platform observability.
  • Demonstrated open-source contribution, maintainership, or meaningful participation in an AI infrastructure project.

Responsibilities

  • Design and deploy scalable, multi-node GPU clusters on Kubernetes.
  • Build distributed training and reinforcement learning platforms using Ray and NVIDIA NeMo RL, supporting PyTorch/JAX.
  • Deploy and optimize high-throughput LLM inference using vLLM, SGLang, NVIDIA Dynamo.
  • Implement GPU scheduling, quotas, isolation, autoscaling, health monitoring, multi-tenant management.
  • Profile and troubleshoot GPU workloads using Nsight, CUDA, DCGM, and related tools.
  • Automate cluster provisioning, upgrades, workload deployment, and recoveries via IaC and GitOps.
  • Contribute to open-source AI infrastructure communities and upstream best practices.

Skills

Python programming
Linux proficiency
Open-source collaboration
Problem-solving

Education

Bachelor's Degree in Software Engineering or related field

Tools

Kubernetes
Ray
NVIDIA NeMo RL
vLLM
SGLang
NVIDIA Dynamo
PyTorch
JAX
CUDA
NVIDIA Nsight

Job description

KLA in the United States, Michigan (Ann Arbor) is seeking a Sr. Software Engineer to design, build, and operate production-grade AI infrastructure for large-scale GPU training and inference. You will own cluster architecture, multi-node orchestration, and performance engineering in production environments.

The role emphasizes Kubernetes, distributed systems, and open-source collaboration, with a strong emphasis on Python/Linux skills and capacity to deploy multi-node AI platforms.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI/ML Engineer for Production-Ready Systems
Senior AI/ML Engineer for Production-Ready Systems

KLA • Ann Arbor Charter Township (MI)

Hybrid
USD 129,600 - 220,300
Medical, dental, and vision benefits
401(K) with company matching
Paid time off and holidays
Senior AI/ML Production Systems Engineer
Senior AI/ML Production Systems Engineer

KLA-Belgium • Ann Arbor (MI)

On-site
USD 130,000 - 220,000
Medical
Dental
Vision
+4
Senior AI/ML Systems Engineer (Production-Ready)
Senior AI/ML Systems Engineer (Production-Ready)

KLA • Ann Arbor (MI)

On-site
USD 130,000 - 220,000
AI Systems Lead: Production ML & Data Ops
AI Systems Lead: Production ML & Data Ops

KLA • Chandler (AZ)

On-site
USD 112,000 - 190,000
Senior AI/ML Engineer — Hybrid, Production-Ready
Senior AI/ML Engineer — Hybrid, Production-Ready

KLA • Ann Arbor (MI)

Hybrid
USD 129,600 - 220,300
401(K) with company matching
Employee stock purchase program (ESPP)
Tuition reimbursement
+2
AI & HPC Infrastructure Lead
AI & HPC Infrastructure Lead

KLA • Ann Arbor Charter Township (MI)

On-site
USD 151,100 - 256,900
Medical benefits
401(K) with company matching
Tuition reimbursement program
+1
Lead Full-Stack AI Engineer
Lead Full-Stack AI Engineer

KLA • Ann Arbor (MI)

On-site
USD 112,000 - 190,000
Senior AI/ML Systems Engineer
Senior AI/ML Systems Engineer

KLA • Ann Arbor Charter Township (MI)

On-site
USD 130,000 - 220,000
Medical, dental, vision benefits
401(k) matching
Employee stock purchase program (ESPP)
+3
Senior AI Infrastructure Engineer — Kubernetes & GPU
Senior AI Infrastructure Engineer — Kubernetes & GPU

Seekr • Austin (TX)

Hybrid
USD 180,000 - 240,000
Equity ownership
Unlimited PTO
14 holidays
+4
Senior AI Infrastructure Engineer - GPU & Kubernetes
Senior AI Infrastructure Engineer - GPU & Kubernetes

HCL Technologies Limited • California (MO)

On-site
USD 120,000 - 180,000
401(k) retirement plan
Paid time off (PTO)
Paid holidays
+1