Senior AI Infrastructure Engineer - GPU & Kubernetes

HCL Technologies Limited

California (MO)

On-site

USD 120,000 - 180,000

Full time

34 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

401(k) retirement plan
Paid time off (PTO)
Paid holidays
Medical, dental, vision insurance

Job summary

HCL Technologies Limited is seeking an experienced AI Infrastructure Engineer (L3) to design, deploy, optimize, and support high-performance GPU-based AI infrastructure. You will manage large GPU clusters, support AI training/inference workloads, and drive platform reliability.

Applicants should have deep expertise in NVIDIA GPU platforms, Kubernetes, HPC environments, and cloud-native AI technologies. Collaboration with cloud, data science, and SRE teams is essential.

Qualifications

  • Bachelor's degree in CS/Engineering or related field.
  • 8-12 years of infrastructure/platform engineering experience.
  • 4-6 years supporting AI/ML environments and GPU-based platforms.
  • Experience operating production-scale AI infrastructure.

Responsibilities

  • Deploy and manage NVIDIA GPU infrastructure and AI accelerator platforms.
  • Administer Kubernetes GPU clusters with NVIDIA GPU Operator.
  • Install CUDA, cuDNN, TensorRT, firmware and drivers.
  • Manage Ceph, Lustre, BeeGFS, and NFS storage.
  • Support InfiniBand, RDMA, RoCE, NVLink and other high-speed networking.
  • Optimize Linux environments for AI and HPC workloads.
  • Support Kubeflow, MLflow, Ray, and Slurm.
  • Implement infrastructure as code with Terraform, Helm, and GitOps.
  • Monitor with Prometheus, Grafana, NVIDIA DCGM and OpenTelemetry.
  • Lead RCA and resolve GPU, networking, storage, and platform issues.
  • Collaborate with cloud, data science, MLOps, SRE, and engineering teams.

Skills

NVIDIA GPU platforms
GPU cluster administration
Kubernetes & cloud-native tech
CUDA & TensorRT
Distributed training frameworks
Linux performance tuning
Terraform & Helm
ArgoCD & automation
MLOps & AI infra
Troubleshooting & production support

Education

Bachelor's Degree in CS/Engineering or related field

Tools

NVIDIA GPU Operator
Ceph
Lustre
BeeGFS
NFS
Prometheus
Grafana
OpenTelemetry
Kubeflow
MLflow
Ray
Slurm
GitOps tools
Terraform
Helm

Job description

HCL Technologies Limited is seeking an experienced AI Infrastructure Engineer (L3) to design, deploy, optimize, and support high-performance GPU-based AI infrastructure. You will manage large GPU clusters, support AI training/inference workloads, and drive platform reliability.

Applicants should have deep expertise in NVIDIA GPU platforms, Kubernetes, HPC environments, and cloud-native AI technologies. Collaboration with cloud, data science, and SRE teams is essential.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Infrastructure Engineer
Infrastructure Engineer

HCL Technologies Limited • California (MO)

On-site
USD 120,000 - 180,000
401(k) retirement plan
Paid time off (PTO)
Paid holidays
+1
Senior AI Infra Engineer: GPU Compute on Kubernetes
Senior AI Infra Engineer: GPU Compute on Kubernetes

Harell Data • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Senior AI Infra Engineer: GPU Clusters & Kubernetes
Senior AI Infra Engineer: GPU Clusters & Kubernetes

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Health insurance
Flexible work hours
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Senior Cloud AI Infra Engineer: Go/C, Kubernetes, GPUs
Senior Cloud AI Infra Engineer: Go/C, Kubernetes, GPUs

Nvidia • Santa Clara (CA)

On-site
USD 240,000 - 340,000
Competitive salaries
Comprehensive benefits package
Equity
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Senior GPU Infrastructure Engineer — HPC & Clusters
Senior GPU Infrastructure Engineer — HPC & Clusters

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
AI Infrastructure Engineer — GPU, Kubernetes & Automation
AI Infrastructure Engineer — GPU, Kubernetes & Automation

MARS-TECHNOMINDS, INC • Town of Florida (NY)

On-site
USD 100,000 - 130,000
Senior AI Compute Infra Engineer (Hybrid)
Senior AI Compute Infra Engineer (Hybrid)

Arm Limited • Seattle (WA)

Hybrid
USD 209,000 - 283,000
Relocation package
Recruitment accommodations