AI Infrastructure Engineer - GPU & Kubernetes

HCLTech

California (MO)

On-site

USD 150,000 - 210,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical Insurance
Dental Insurance
Vision Insurance
401(k) Retirement Plan
Paid Time Off and Holidays

Job summary

HCLTech is seeking an experienced AI Infrastructure Engineer (L3) to design, deploy, and optimize large-scale GPU-based AI infrastructure. The role covers GPU clusters, CUDA/TensorRT stacks, storage, and high-speed networking across Linux environments in cloud-native contexts.

Ideal candidates will have extensive Kubernetes experience, proficiency with Terraform/Helm/ArgoCD, and a track record of production-grade AI platforms.

Qualifications

  • Bachelor's Degree in Computer Science, Engineering, or a related field.
  • 8-12 years of Infrastructure or Platform Engineering experience.
  • 4-6 years supporting AI/ML environments and GPU-based platforms.
  • Experience operating production-scale AI infrastructure.
  • Strong Linux administration and performance tuning skills.
  • Experience operating Kubernetes and cloud-native technologies.
  • Hands-on experience with Terraform, Helm, ArgoCD, and automation frameworks.

Responsibilities

  • Deploy and manage NVIDIA GPU infrastructure and AI accelerator platforms.
  • Administer Kubernetes GPU clusters using NVIDIA GPU Operator.
  • Install and maintain CUDA, cuDNN, TensorRT, firmware, and driver stacks.
  • Manage high-performance storage such as Ceph, Lustre, BeeGFS, and NFS.
  • Support InfiniBand, RDMA, RoCE, NVLink, and other high-speed networking.
  • Optimize Linux environments for AI and HPC workloads.
  • Support AI orchestration platforms like Kubeflow, MLflow, Ray, and Slurm.
  • Implement Infrastructure as Code using Terraform, Helm, and GitOps.
  • Monitor platform performance with Prometheus, Grafana, NVIDIA DCGM, and OpenTelemetry.
  • Lead root cause analysis and resolve GPU, networking, storage, and platform issues.
  • Collaborate with cloud, data science, MLOps, SRE, and engineering teams to deliver scalable AI platforms.

Skills

NVIDIA GPU platforms
Kubernetes
CUDA
TensorRT
HPC environments
Linux performance tuning
Terraform
Helm
ArgoCD
OpenTelemetry
Prometheus
NVIDIA DCGM

Education

Bachelor's Degree in Computer Science or Engineering

Tools

NVIDIA GPU Operator
Kubernetes
Terraform
Helm
ArgoCD
GitOps
Prometheus
Grafana
NVIDIA DCGM

Job description

HCLTech is seeking an experienced AI Infrastructure Engineer (L3) to design, deploy, and optimize large-scale GPU-based AI infrastructure. The role covers GPU clusters, CUDA/TensorRT stacks, storage, and high-speed networking across Linux environments in cloud-native contexts.

Ideal candidates will have extensive Kubernetes experience, proficiency with Terraform/Helm/ArgoCD, and a track record of production-grade AI platforms.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infrastructure Engineer
Infrastructure Engineer

ITCAPS LLC • St. Louis (MO)

On-site
USD 140,000 - 200,000
Infrastructure Engineer
Infrastructure Engineer

HCLTech • California (MO)

On-site
USD 150,000 - 210,000
Medical Insurance
Dental Insurance
Vision Insurance
+2
Senior Kubernetes Engineer: GPU AI Infra Platform
Senior Kubernetes Engineer: GPU AI Infra Platform

GTN Technical Staffing • Dallas (TX)

On-site
USD 150,000 - 210,000
AI Infrastructure Engineer — GPU Kubernetes for Production
AI Infrastructure Engineer — GPU Kubernetes for Production

vCluster • Germany (OH)

On-site
USD 150,000 - 200,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
Senior AI Infra Engineer: GPU Compute on Kubernetes
Senior AI Infra Engineer: GPU Compute on Kubernetes

Harell Data • Palo Alto (CA)

On-site
USD 180,000 - 260,000
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
AI Infra Engineer: Kubernetes on Bare Metal GPUs (Remote)
AI Infra Engineer: Kubernetes on Bare Metal GPUs (Remote)

vCluster • New York (NY)

Hybrid
USD 150,000 - 200,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Senior Cloud AI Infra Engineer: Go/C, Kubernetes, GPUs
Senior Cloud AI Infra Engineer: Go/C, Kubernetes, GPUs

Nvidia • Santa Clara (CA)

On-site
USD 240,000 - 340,000
Competitive salaries
Comprehensive benefits package
Equity
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2