Sr. Technology Enablement Engineer

KLA-Belgium

Ann Arbor (MI)

On-site

USD 130,000 - 190,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Medical
Dental
Vision
401(K) including company matching
ESPP

Job summary

KLA in the United States, Michigan (Ann Arbor) is seeking a Sr. Software Engineer to design, build, and operate production-grade AI infrastructure for large-scale GPU training and inference. You will own cluster architecture, multi-node orchestration, and performance engineering in production environments.

The role emphasizes Kubernetes, distributed systems, and open-source collaboration, with a strong emphasis on Python/Linux skills and capacity to deploy multi-node AI platforms.

Qualifications

  • Bachelor's Degree and eight (8) years of Software Engineering experience.
  • Four (4) years in software, cloud, platform, HPC, or infrastructure engineering, including two (2) years supporting distributed AI/ML workloads.
  • Proven hands-on experience building GPU clusters from the ground up and operating them at production scale.
  • Deep Kubernetes expertise, including operators, CRDs, Helm, networking, storage, scheduling, and cluster lifecycle management.
  • Production experience with Ray and at least two of the following: NVIDIA NeMo RL, vLLM, SGLang, or NVIDIA Dynamo.
  • Strong knowledge of NVIDIA GPUs, CUDA, NCCL, GPU Operator, DCGM, MIG, RDMA, and multi-node collective communication.
  • Experience debugging training and inference performance across GPU, CPU, memory, network, and storage layers.
  • Proficiency in Python and Linux, plus experience with containers, CI/CD, GitOps, infrastructure-as-code, and platform observability.
  • Demonstrated open-source contribution, maintainership, or meaningful participation in an AI infrastructure project.

Responsibilities

  • Design and deploy scalable, multi-node GPU clusters on Kubernetes.
  • Build distributed training and reinforcement learning platforms using Ray and NVIDIA NeMo RL, supporting PyTorch/JAX.
  • Deploy and optimize high-throughput LLM inference using vLLM, SGLang, NVIDIA Dynamo.
  • Implement GPU scheduling, quotas, isolation, autoscaling, health monitoring, multi-tenant management.
  • Profile and troubleshoot GPU workloads using Nsight, CUDA, DCGM, and related tools.
  • Automate cluster provisioning, upgrades, workload deployment, and recoveries via IaC and GitOps.
  • Contribute to open-source AI infrastructure communities and upstream best practices.

Skills

Python programming
Linux proficiency
Open-source collaboration
Problem-solving

Education

Bachelor's Degree in Software Engineering or related field

Tools

Kubernetes
Ray
NVIDIA NeMo RL
vLLM
SGLang
NVIDIA Dynamo
PyTorch
JAX
CUDA
NVIDIA Nsight

Job description

To make electronics, you need chips, wafers, transistors, reticles, and... To make these, you must see, test and manufacture them at scale—faster and better than ever before. That's where KLA comes in. Whether you're early in your career or an experienced professional, you'll solve complex challenges, work alongside brilliant minds and help shape the future of technology. Group/Division KLA's IT group supports business growth and productivity by connecting people, process and technology around the world. We work to improve the technology that drives our business to thrive and focus on empowering employee use of technology. This integrated approach to customer service, creativity and technological excellence enhances employee productivity, business analytics and process excellence.

What You'll Do

In this role, you will play a key part in advancing business priorities by delivering high-impact work across your area of expertise. We are seeking a highly skilled Sr. Software Engineer with specialized expertise Design, build, and operate production-grade AI infrastructure for large-scale GPU training and inference. The role owns cluster architecture from ground zero, distributed workload orchestration, accelerator management, performance engineering, and open-source integration. TPU experience is optional.

Key Responsibilities
  • Design and deploy scalable, multi-node GPU clusters on Kubernetes, including compute, networking, storage, scheduling, security, and observability.
  • Build distributed training and reinforcement learning platforms using Ray and NVIDIA NeMo RL, supporting frameworks such as PyTorch and JAX.
  • Deploy and optimize high-throughput LLM inference using vLLM, SGLang, and NVIDIA Dynamo.
  • Implement GPU scheduling, quotas, isolation, autoscaling, health monitoring, and capacity management for multi-tenant environments.
  • Profile and troubleshoot GPU workloads using NVIDIA Nsight Systems, Nsight Compute, DCGM, CUDA, NCCL, and related diagnostics.
  • Automate cluster provisioning, upgrades, workload deployment, and operational recovery through infrastructure-as-code and GitOps practices.
  • Contribute to or actively participate in relevant open-source AI infrastructure communities and bring upstream best practices into the platform.
Minimum Qualifications
  • Bachelor's Degree and eight (8) years of Software Engineering experience
  • Four (4) years in software, cloud, platform, HPC, or infrastructure engineering, including two (2) years supporting distributed AI/ML workloads.
  • Proven hands‑on experience building GPU clusters from the ground up and operating them at production scale.
  • Deep Kubernetes expertise, including operators, CRDs, Helm, networking, storage, scheduling, and cluster lifecycle management.
  • Production experience with Ray and at least two of the following: NVIDIA NeMo RL, vLLM, SGLang, or NVIDIA Dynamo.
  • Strong knowledge of NVIDIA GPUs, CUDA, NCCL, GPU Operator, DCGM, MIG, RDMA, and multi‑node collective communication.
  • Experience debugging training and inference performance across GPU, CPU, memory, network, and storage layers.
  • Proficiency in Python and Linux, plus experience with containers, CI/CD, GitOps, infrastructure‑as‑code, and platform observability.
  • Demonstrated open‑source contribution, maintainership, or meaningful participation in an AI infrastructure project.
Preferred Qualifications
  • Experience with Google TPUs and TPU‑oriented frameworks or distributed workloads.
  • Experience with large language model training, fine‑tuning, RLHF or agentic reinforcement learning.
  • Knowledge of TensorRT‑LLM, Triton Inference Server, DeepSpeed, Megatron‑LM, or similar performance‑oriented frameworks.
  • Experience operating secure, multi‑tenant AI platforms in enterprise or regulated environments.
About KLA

We provide advanced inspection tools, metrology systems, process solutions, and computational analytics that make electronics possible, tackling complex challenges. From electron and photon optics to machine learning and data analytics, we seek perfection at the most fundamental level of matter in the universe. If you want to make electronics that push industries forward and make the world a better place, join us.

Total Rewards

Base Pay Range: $129,600.00 - $190,067.00 Annually

Primary Location: USA-MI-Ann Arbor-KLA

KLA’s total rewards package for employees may also include participation in performance incentive programs and eligibility for additional benefits including but not limited to:

  • medical
  • dental
  • vision
  • life
  • other voluntary benefits
  • 401(K) including company matching
  • employee stock purchase program (ESPP)
  • student debt assistance
  • tuition reimbursement program
  • development and career growth opportunities and programs
  • financial planning benefits
  • wellness benefits including an employee assistance program (EAP)
  • paid time off and paid company holidays
  • family care and bonding leave

Interns are eligible for some of the benefits listed.

Our pay ranges are determined by role, level, and location. The range displayed reflects the pay for this position in the primary location identified in this posting. Actual pay depends on several factors, including state minimum pay wage rates, location, job‑related skills, experience, and relevant education level or training. We are committed to complying with all applicable federal and state minimum wage requirements where applicable. If applicable, your recruiter can share more about the specific pay range for your preferred location during the hiring process.

Use of AI Statement

At KLA, our interviews seek to understand your individual skills, problem‑solving approach and authentic thinking. To ensure a fair and consistent evaluation, the use of AI, recording tools or other technologies to generate, suggest or provide responses during interviews—whether virtual or in person—is not permitted unless explicitly approved in advance as part of a reasonable accommodation or invited by the interviewer. Use of these tools may interfere with our ability to evaluate your individual qualifications and affect your candidacy.

KLA is committed to advancing innovation through responsible AI, and we value candidates who share this mindset.

Equal Opportunity Statement

KLA is proud to be an Equal Opportunity Employer. We will ensure that qualified individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us at talent.acquisition@kla.com or at +1-408-352-2808 to request accommodation.

For additional information, view the US Know Your Rights poster on the U.S. Equal Employment Opportunity Commission website.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Algorithm Software Architecture
Algorithm Software Architecture

KLA-Belgium • Milpitas (CA)

On-site
USD 186,000 - 273,000
Medical, dental, vision
Life insurance
401(k) with company match
+2
Engr, Software 1
Engr, Software 1

KLA-Belgium • Ann Arbor (MI)

On-site
USD 90,000 - 140,000
Medical
Dental
Vision
+3
Sr. AI/ML Engineer - Systems
Sr. AI/ML Engineer - Systems

KLA-Belgium • Ann Arbor (MI)

On-site
USD 130,000 - 220,000
Medical
Dental
Vision
+4
AI Solutions Engineer (NPI Function)
AI Solutions Engineer (NPI Function)

KLA-Belgium • Ann Arbor (MI)

On-site
USD 94,000 - 138,000
AI Solutions Engineer (NPI Function)
AI Solutions Engineer (NPI Function)

KLA • Ann Arbor (MI)

On-site
USD 94,000 - 138,000
AI Solutions Engineer (NPI Function)
AI Solutions Engineer (NPI Function)

KLA • Ann Arbor Charter Township (MI)

On-site
USD 94,000 - 138,000
Algorithm Engineering Intern (AI, Computer Vision & Software Engineering)
Algorithm Engineering Intern (AI, Computer Vision & Software Engineering)

KLA-Belgium • Milpitas (CA)

On-site
USD 66,000 - 94,000
HPC / AI Software Infrastructure Lead (E)
HPC / AI Software Infrastructure Lead (E)

KLA • Ann Arbor Charter Township (MI)

On-site
USD 151,100 - 256,900
Medical benefits
401(K) with company matching
Tuition reimbursement program
+1
Service Supply Chain AI Engineer
Service Supply Chain AI Engineer

KLA • Ann Arbor Charter Township (MI)

On-site
USD 90,000 - 133,000
Sr. Software Engineer - AI/ML
Sr. Software Engineer - AI/ML

KLA • Ann Arbor Charter Township (MI)

On-site
USD 129,600 - 220,300
Medical, dental, and vision benefits
401(K) with company matching
Paid time off and holidays