Infrastructure Engineer

HCL Technologies Limited

California (MO)

On-site

USD 120,000 - 180,000

Full time

29 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

401(k) retirement plan
Paid time off (PTO)
Paid holidays
Medical, dental, vision insurance

Job summary

HCL Technologies Limited is seeking an experienced AI Infrastructure Engineer (L3) to design, deploy, optimize, and support high-performance GPU-based AI infrastructure. You will manage large GPU clusters, support AI training/inference workloads, and drive platform reliability.

Applicants should have deep expertise in NVIDIA GPU platforms, Kubernetes, HPC environments, and cloud-native AI technologies. Collaboration with cloud, data science, and SRE teams is essential.

Qualifications

  • Bachelor's degree in CS/Engineering or related field.
  • 8-12 years of infrastructure/platform engineering experience.
  • 4-6 years supporting AI/ML environments and GPU-based platforms.
  • Experience operating production-scale AI infrastructure.

Responsibilities

  • Deploy and manage NVIDIA GPU infrastructure and AI accelerator platforms.
  • Administer Kubernetes GPU clusters with NVIDIA GPU Operator.
  • Install CUDA, cuDNN, TensorRT, firmware and drivers.
  • Manage Ceph, Lustre, BeeGFS, and NFS storage.
  • Support InfiniBand, RDMA, RoCE, NVLink and other high-speed networking.
  • Optimize Linux environments for AI and HPC workloads.
  • Support Kubeflow, MLflow, Ray, and Slurm.
  • Implement infrastructure as code with Terraform, Helm, and GitOps.
  • Monitor with Prometheus, Grafana, NVIDIA DCGM and OpenTelemetry.
  • Lead RCA and resolve GPU, networking, storage, and platform issues.
  • Collaborate with cloud, data science, MLOps, SRE, and engineering teams.

Skills

NVIDIA GPU platforms
GPU cluster administration
Kubernetes & cloud-native tech
CUDA & TensorRT
Distributed training frameworks
Linux performance tuning
Terraform & Helm
ArgoCD & automation
MLOps & AI infra
Troubleshooting & production support

Education

Bachelor's Degree in CS/Engineering or related field

Tools

NVIDIA GPU Operator
Ceph
Lustre
BeeGFS
NFS
Prometheus
Grafana
OpenTelemetry
Kubeflow
MLflow
Ray
Slurm
GitOps tools
Terraform
Helm

Job description

HCLTech is a global technology company with over 220,000 professionals across 60 countries, delivering industry-leading capabilities in Digital, Engineering, Cloud, and AI. We help enterprises accelerate innovation through cutting-edge technologies and world-class talent.

Job Summary

We are seeking an experienced AI Infrastructure Engineer (L3) to design, deploy, optimize, and support high-performance AI and Machine Learning infrastructure. The ideal candidate will have deep expertise in GPU platforms, Kubernetes, HPC environments, distributed systems, and cloud-native AI technologies. This role involves managing large-scale GPU clusters, supporting AI training and inference workloads, troubleshooting complex infrastructure issues, and driving platform reliability.

Key Responsibilities
  • Deploy and manage NVIDIA GPU infrastructure (A100, H100, L40) and AI accelerator platforms.
  • Administer Kubernetes GPU clusters using NVIDIA GPU Operator and related technologies.
  • Install and maintain CUDA, cuDNN, TensorRT, firmware, and driver stacks.
  • Manage high-performance storage solutions such as Ceph, Lustre, BeeGFS, and NFS.
  • Support InfiniBand, RDMA, RoCE, NVLink, and other high-speed networking technologies.
  • Optimize Linux environments (RHEL, Ubuntu, Rocky Linux) for AI and HPC workloads.
  • Support AI orchestration platforms including Kubeflow, MLflow, Ray, and Slurm.
  • Implement Infrastructure as Code using Terraform, Helm, and GitOps tools.
  • Monitor platform performance with Prometheus, Grafana, NVIDIA DCGM, and OpenTelemetry.
  • Lead root cause analysis (RCA) and resolve critical GPU, networking, storage, and platform issues.
  • Collaborate with cloud, data science, MLOps, SRE, and engineering teams to deliver scalable AI platforms.
Required Skills
  • Strong experience with NVIDIA GPU platforms and GPU cluster administration.
  • Expertise in Kubernetes, containerization, and cloud-native technologies.
  • Hands-on experience with CUDA, TensorRT, NCCL, DeepSpeed, Horovod, and distributed training.
  • Strong Linux administration and performance tuning skills.
  • Experience with Terraform, Helm, ArgoCD, and automation frameworks.
  • Knowledge of AI infrastructure, MLOps, and large-scale distributed systems.
  • Excellent troubleshooting, debugging, and production support experience.
Preferred Certifications
  • NVIDIA Certified Associate – AI Infrastructure
  • NVIDIA Base Command Manager Certification
  • AWS Solutions Architect Associate
Qualifications
  • Bachelor's Degree in Computer Science, Engineering, or a related field.
  • 8-12 years of Infrastructure or Platform Engineering experience.
  • 4-6 years supporting AI/ML environments and GPU-based platforms.
  • Experience operating production-scale AI infrastructure.

Disclaimer

HCL is an equal opportunity employer, committed to providing equal employment opportunities to all applicants and employees regardless of race, religion, sex, color, age, national origin, pregnancy, sexual orientation, physical disability or genetic information, military or veteran status, or any other protected classification, in accordance with federal, state, and/or local law. Should any applicant have concerns about discrimination in the hiring process, they should provide a detailed report of those concerns to secure@hcltech.com for investigation.

Compensation and Benefits

A candidate’s pay within the range will depend on their work location, skills, experience, education, and other factors permitted by law. This role may also be eligible for performance-based bonuses subject to company policies. In addition, this role is eligible for the following benefits subject to company policies: medical, dental, vision, pharmacy, life, accidental death & dismemberment, and disability insurance; employee assistance program; 401(k) retirement plan; 10 days of paid time off per year (some positions are eligible for need-based leave with no designated number of leave days per year); and 10 paid holidays per year.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer - GPU & Kubernetes
Senior AI Infrastructure Engineer - GPU & Kubernetes

HCL Technologies Limited • California (MO)

On-site
USD 120,000 - 180,000
401(k) retirement plan
Paid time off (PTO)
Paid holidays
+1
AI and ML HPC Cluster Engineer, AI and ML HPC Cluster Engineer
AI and ML HPC Cluster Engineer, AI and ML HPC Cluster Engineer

NVIDIA • Colorado

On-site
USD 124,000 - 195,500
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA • United States

On-site
USD 176,000 - 334,000
Equity
Benefits
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Sr. AI Infrastructure Engineer
Sr. AI Infrastructure Engineer

ECLARO • Costa Mesa (CA)

On-site
USD 166,000 - 220,000
Senior Solutions Architect, AI Infrastructure Enterprise ISVs
Senior Solutions Architect, AI Infrastructure Enterprise ISVs

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 287,500
Equity
Benefits
Travel opportunities
Senior AI Compute Engineer - NVIS
Senior AI Compute Engineer - NVIS

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 148,000 - 287,500
Senior Solutions Architect, AI Infrastructure Enterprise ISVs
Senior Solutions Architect, AI Infrastructure Enterprise ISVs

NVIDIA • New York (NY)

On-site
USD 184,000 - 287,500
Equity
Benefits
Senior Solutions Architect, AI Infrastructure Enterprise ISVs
Senior Solutions Architect, AI Infrastructure Enterprise ISVs

NVIDIA • California (MO)

On-site
USD 184,000 - 287,500
Equity
Benefits