Senior Kubernetes Administrator – AI Infrastructure

HCLTech

United States

On-site

USD 78,000 - 148,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

HCLTech is seeking a Senior Kubernetes Administrator – AI Infrastructure to lead platform engineering for AI workloads in a remote setting. You will build, troubleshoot, and scale Kubernetes platforms supporting model development, training, and inference services.

Ideal candidates have 7+ years in infrastructure, deep Kubernetes expertise, experience with GPUs, Helm, GitOps, and strong Linux networking. Remote work with competitive compensation and benefits.

Qualifications

  • 7+ years in infrastructure engineering with hands-on Kubernetes administration.
  • Deep understanding of Kubernetes internals and cluster troubleshooting.
  • Experience with container runtimes, Helm, GitOps, and declarative operations.
  • Experience supporting GPU workloads on Kubernetes in lab/production.
  • Strong Linux administration and data center networking knowledge.
  • Ability to debug from symptom to root cause across node, pod, network, storage and control plane.
  • Scripting in Python, Bash, or Go.

Responsibilities

  • Build, administer, and troubleshoot Kubernetes platforms used for AI and data-intensive workloads.
  • Diagnose failures across control plane components, kubelet, CNI, CSI, ingress, and scheduling.
  • Support GPU-enabled Kubernetes environments, including device plugin behavior and drivers.
  • Improve platform reliability through automation, configuration standards, upgrade planning, and validation gates.
  • Investigate storage throughput, network policy, DNS, image pulls, autoscaling, and degraded node states.
  • Collaborate with Linux, network, validation, and SRE teams to resolve cross-layer issues.
  • Create reusable runbooks, dashboards, and health checks for day-2 operations.
  • Contribute to platform hardening and tenant readiness.

Skills

Kubernetes admin
Linux admin
Python
Bash
Go
GPU workloads
Cluster troubleshooting
GitOps
Helm
Declarative ops
Networking basics

Tools

Kubeflow
Argo
Prometheus
Grafana
Loki
Service mesh
Bare-metal Kubernetes
High-performance storage

Job description

HCLTech is looking for a highly talented and self- motivated Senior Kubernetes Administrator – AI Infrastructure

to join it in advancing the technological world through innovation and creativity.

Job ID: 158106

Position Type: Full-time

Location: Remote

Role/Responsibilities
Engagement summary

The Candidate will provide senior Kubernetes platform engineering services for AI infrastructure environments supporting model development, distributed training, inference services, and shared platform operations. The role requires strong cluster troubleshooting ability plus pragmatic platform engineering skill in mixed bare-metal and data center environments.

What this Candidate will be doing
  • Build, administer, and troubleshoot Kubernetes platforms used for AI and data-intensive workloads.
  • Diagnose failures across control plane components, kubelet, CNI, CSI, ingress, service discovery, scheduling, node lifecycle, container runtime, and resource isolation.
  • Support GPU-enabled Kubernetes environments, including device plugin behavior, driver dependencies, node health, and workload placement.
  • Improve platform reliability through automation, standardized configuration, upgrade planning, and cluster validation gates.
  • Investigate workload issues involving storage throughput, network policy, DNS, image pulls, autoscaling, pod eviction, and degraded node states.
  • Partner with Linux, network, validation, and SRE teams to resolve complex cross-layer failures affecting AI services.
  • Create reusable operational runbooks, dashboards, and health checks for day-2 support.
  • Contribute to platform hardening, tenant readiness, and service-level objectives.
What we need to see
  • 7+ years in infrastructure engineering, with deep hands-on Kubernetes administration experience.
  • Strong operational understanding of Kubernetes internals and cluster troubleshooting.
  • Experience with container runtimes, Helm, GitOps or declarative operations, and cluster lifecycle management.
  • Experience supporting GPU workloads on Kubernetes in lab, validation, or production settings.
  • Strong Linux administration foundation and understanding of data center network dependencies.
  • Ability to debug issues from symptom to root cause across node, pod, network, storage, and control plane layers.
  • Scripting and automation skill in Python, Bash, or Go.
Preferred experience
  • Experience with Kubeflow, Argo, Prometheus, Grafana, Loki, or service mesh technologies.
  • Familiarity with bare-metal Kubernetes and high-performance storage integration.
  • Exposure to regulated or high-change-control production environments.
Pay and Benefits

Pay Range Minimum: $78,000/Annum

Pay Range Maximum: $148,000/Annum

HCLTech is an equal opportunity employer, committed to providing equal employment opportunities to all applicants and employees regardless of race, religion, sex, color, age, national origin, pregnancy, sexual orientation, physical disability or genetic information, military or veteran status, or any other protected classification, in accordance with federal, state, and/or local law. Should any applicant have concerns about discrimination in the hiring process, they should provide a detailed report of those concerns to secure@hcltech.com for investigation.

Compensation and Benefits

A candidate’s pay within the range will depend on their work location, skills, experience, education, and other factors permitted by law. This role may also be eligible for performance-based bonuses subject to company policies. In addition, this role is eligible for the following benefits subject to company policies: medical, dental, vision, pharmacy, life, accidental death & dismemberment, and disability insurance; employee assistance program; 401(k) retirement plan; 10 days of paid time off per year (some positions are eligible for need-based leave with no designated number of leave days per year); and 10 paid holidays per year.

How You’ll Grow

At HCLTech, we offer continuous opportunities for you to find your spark and grow with us. We want you to be happy and satisfied with your role and to really learn what type of work sparks your brilliance the best. Throughout your time with us, we offer transparent communication with senior-level employees, learning and career development programs at every level, and opportunities to experiment in different roles or even pivot industries. We believe that you should be in control of your career with unlimited opportunities to find the role that fits you best.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Kubernetes Engineer for AI Infrastructure (Remote)
Senior Kubernetes Engineer for AI Infrastructure (Remote)

HCLTech • United States

On-site
USD 78,000 - 148,000
Senior Technical Lead – DevOps, Python, Kubernetes , Agentic AI Workflows
Senior Technical Lead – DevOps, Python, Kubernetes , Agentic AI Workflows

HCLTech • Denver (CO)

On-site
USD 78,000 - 115,000
Senior AI Engineer
Senior AI Engineer

HCLTech • New York (NY)

On-site
USD 120,000 - 180,000
Medical insurance
Dental insurance
Vision insurance
+5
Kubernetes Administrator – AI Infrastructure
Kubernetes Administrator – AI Infrastructure

Sira Consulting, an Inc 5000 company • United States

On-site
USD 140,000 - 210,000
Infrastructure Engineer
Infrastructure Engineer

HCLTech • California (MO)

On-site
USD 150,000 - 210,000
Medical Insurance
Dental Insurance
Vision Insurance
+2
Assosiate Program Manager - AI Infrastructure
Assosiate Program Manager - AI Infrastructure

HCLTech • Houston (TX)

On-site
USD 34,440,000 - 41,328,000
Medical benefits
401(k) retirement plan
Paid time off
+1
DevOps / Cloud Engg
DevOps / Cloud Engg

HCLTech • United States

Remote
USD 88,000 - 134,000
Medical insurance
401(k) retirement plan
Paid time off
AI Principal Architect
AI Principal Architect

HCLTech • New Jersey

Hybrid
USD 180,000 - 240,000
Performance-based bonuses
Medical, dental, vision
401(k) retirement plan
Forward Deployed Engineer Lead
Forward Deployed Engineer Lead

HCLTech • New Jersey

On-site
USD 164,000 - 263,000
Medical insurance
Dental insurance
Vision insurance
+4
Forward Deployment Engineer
Forward Deployment Engineer

HCLTech • Dallas (TX)

Hybrid
USD 150,000 - 210,000