Kubernetes Administrator – AI Infrastructure

Sira Consulting, an Inc 5000 company

United States

On-site

USD 140,000 - 210,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Sira Consulting is seeking a Senior Kubernetes Administrator / Platform Engineer to own and operate Kubernetes platforms for AI, GPU, and data-intensive workloads. The role emphasizes hands-on troubleshooting, automation, and Linux administration in a remote, US-based setting.

The ideal candidate has 7+ years in infra/platform engineering, deep Kubernetes knowledge, and experience with GPU workloads, Helm, and GitOps. Collaboration with SRE and Validation teams is key.

Qualifications

  • 7+ years of infra/platform engineering with hands-on Kubernetes.
  • Strong understanding of Kubernetes internals and troubleshooting.
  • Experience with container runtimes, Helm, GitOps, and cluster lifecycle management.
  • Experience supporting GPU workloads on Kubernetes in production or validation environments.
  • Strong Linux administration and data center network knowledge.
  • Ability to troubleshoot across nodes, pods, networking, storage, and control-plane layers.
  • Scripting/automation: Python, Bash, or Go.

Responsibilities

  • Build, administer, and troubleshoot Kubernetes platforms supporting AI and data-intensive workloads.
  • Troubleshoot control plane, kubelet, CNI, CSI, ingress, service discovery, scheduling, container runtime, and resource issues.
  • Support GPU-enabled Kubernetes environments, including device plugins and node health.
  • Improve platform reliability through automation, standardized configurations, upgrade planning, and cluster validation.
  • Troubleshoot storage, network policies, DNS, image pulls, autoscaling, pod eviction, and degraded nodes.
  • Collaborate with Linux, Network, SRE, and Validation teams to resolve cross-layer infrastructure issues.
  • Develop reusable runbooks, dashboards, health checks, and operational procedures for day-2 support.
  • Contribute to platform hardening, tenant readiness, monitoring, and SLAs.

Skills

Kubernetes administration
Cluster troubleshooting
Automation scripting
Linux administration
GPU workloads on Kubernetes
Python Bash Go
GitOps / declarative operations
Container runtimes

Tools

Kubernetes
Helm
GitOps
Kubectl
Container runtimes

Job description

Location: Remote

Job Summary

We are seeking a Senior Kubernetes Administrator / Platform Engineer with strong hands-on experience managing Kubernetes platforms for AI, GPU, and data-intensive workloads. The ideal candidate should have strong cluster troubleshooting, automation, containerization, and Linux administration skills.

Key Responsibilities
  • Build, administer, and troubleshoot Kubernetes platforms supporting AI and data-intensive workloads.
  • Troubleshoot control plane, kubelet, CNI, CSI, ingress, service discovery, scheduling, container runtime, and resource issues.
  • Support GPU-enabled Kubernetes environments, including device plugins, drivers, node health, and workload placement.
  • Improve platform reliability through automation, standardized configurations, upgrade planning, and cluster validation.
  • Troubleshoot issues involving storage, network policies, DNS, image pulls, autoscaling, pod eviction, and degraded nodes.
  • Collaborate with Linux, Network, SRE, and Validation teams to resolve cross-layer infrastructure issues.
  • Develop reusable runbooks, dashboards, health checks, and operational procedures for day-2 support.
  • Contribute to platform hardening, tenant readiness, monitoring, and service-level objectives.
Required Skills
  • 7+ years of infrastructure/platform engineering experience with deep hands-on Kubernetes administration.
  • Strong understanding of Kubernetes internals and cluster troubleshooting.
  • Experience with container runtimes, Helm, GitOps/declarative operations, and cluster lifecycle management.
  • Experience supporting GPU workloads on Kubernetes in production, lab, or validation environments.
  • Strong Linux administration skills and understanding of data center network dependencies.
  • Ability to troubleshoot issues across nodes, pods, networking, storage, and control-plane layers.
  • Strong scripting and automation skills using Python, Bash, or Go.
Preferred Skills
  • Experience with Kubeflow, Argo, Prometheus, Grafana, Loki, or service mesh technologies.
  • Familiarity with bare-metal Kubernetes and high-performance storage/networking.
  • Experience working in regulated or high-change-control production environments.
  • Knowledge of AI/ML infrastructure, GPU clusters, and distributed workloads.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Remote Senior Kubernetes Platform Engineer - AI & GPUs
Remote Senior Kubernetes Platform Engineer - AI & GPUs

Sira Consulting, an Inc 5000 company • United States

On-site
USD 140,000 - 210,000
Senior Kubernetes Administrator – AI Infrastructure
Senior Kubernetes Administrator – AI Infrastructure

HCLTech • United States

On-site
USD 78,000 - 148,000
Senior Kubernetes Engineer for AI Infrastructure (Remote)
Senior Kubernetes Engineer for AI Infrastructure (Remote)

HCLTech • United States

On-site
USD 78,000 - 148,000
Senior Kubernetes Engineer
Senior Kubernetes Engineer

GTN Technical Staffing • Dallas (TX)

On-site
USD 140,000 - 200,000
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Infra Engineer - SRE(Kubernetes)
Infra Engineer - SRE(Kubernetes)

GMI Cloud • United States

On-site
USD 100,000 - 130,000
Member of Technical Staff, Platform - AI Infrastructure
Member of Technical Staff, Platform - AI Infrastructure

Hamilton Barnes Associates Limited • United States

On-site
USD 213,000 - 288,000
Equity
Health care
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Systems and Platform Engineer
Senior Systems and Platform Engineer

MAXISIQ, Inc. • Bethesda (MD)

On-site
USD 170,000 - 210,000