Senior Kubernetes Engineer

GTN Technical Staffing

Dallas (TX)

On-site

USD 140,000 - 200,000

Full time

11 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

GTN Technical Staffing is seeking a Senior Kubernetes Engineer to design and scale GPU-accelerated compute platforms across on-prem and hybrid environments. You will own the architecture and operation of large Kubernetes clusters, enabling high-throughput, multi-tenant workloads with automation and platform evolution.

You will work with platform, HPC, and ML teams to deploy GPU-enabled clusters, extend Kubernetes via operators, and optimize NVIDIA tooling like GPU Operator, DCGM, and device

Qualifications

  • Strong experience operating Kubernetes in large-scale, production environments.
  • Hands-on experience with NVIDIA GPU ecosystem (GPU Operator, device plugins, MIG, DCGM).
  • Proficiency in Go or Python for Kubernetes operators and automation.
  • Deep understanding of Kubernetes internals: CRDs, controllers, RBAC, scheduling.

Responsibilities

  • Design, deploy, and operate large-scale Kubernetes clusters optimized for GPU workloads.
  • Architect container platforms for AI/ML, LLM training, and HPC use cases.
  • Extend Kubernetes with custom operators and controllers for automation.
  • Integrate NVIDIA ecosystem components and GPU scheduling strategies.
  • Improve cluster efficiency with scheduler extensions and observability.

Skills

Kubernetes
GPU ecosystem
Go/Python
CRDs & controllers
GitOps CI/CD
Terraform/Helm/Kustomize
RBAC & multi-tenancy
Monitoring/Observability
Networking/CNI

Tools

ArgoCD
FluxCD
Terraform
Helm
Kustomize
Prometheus
Grafana
DCGM

Job description

Type: Direct Hire

  • Competitive base salary + performance bonus
Overview

We are seeking a Senior Kubernetes Engineer to help design and scale a next-generation GPU-accelerated compute platform supporting AI, machine learning, and high-performance computing workloads. This role sits at the core of a rapidly expanding infrastructure environment, focused on building high-throughput, highly efficient container platforms across on-prem and hybrid environments.

You will play a key role in architecting and operating large-scale Kubernetes clusters optimized for GPU workloads, working closely with platform, HPC, and ML engineering teams to deliver reliable, multi-tenant compute at scale. This is a hands‑on engineering role with strong ownership across performance, automation, and platform evolution.

Key Responsibilities
  • Design, deploy, and operate large-scale Kubernetes clusters optimized for GPU-intensive workloads
  • Architect container platforms supporting AI/ML, LLM training, and HPC use cases
  • Extend Kubernetes through custom operators, controllers, and CRDs to support infrastructure automation
  • Integrate and optimize NVIDIA ecosystem components, including GPU Operator, DCGM, and device plugins
  • Implement GPU scheduling strategies, including MIG, sharing, and workload placement optimization
  • Enhance cluster efficiency using scheduler extensions such as kube-scheduler plugins, Slurm, or Volcano
Platform Performance & Reliability
  • Drive performance tuning across compute, networking, and storage layers for high-throughput workloads
  • Partner with HPC and ML teams to ensure scalability, reliability, and workload efficiency
  • Participate in production readiness, incident response, and continuous improvement initiatives
Observability & Automation
  • Implement monitoring and telemetry solutions using Prometheus, Grafana, DCGM Exporter, and OpenTelemetry
  • Build and maintain CI/CD pipelines for infrastructure using GitOps tools such as ArgoCD and FluxCD
  • Contribute to infrastructure-as-code using Terraform, Helm, and Kustomize
Security & Multi-Tenancy
  • Design and enforce secure multi-tenant environments with namespace isolation, RBAC, and policy controls
  • Implement governance frameworks using tools such as OPA or Gatekeeper
  • Ensure compliance with platform security and operational standards
Required Experience
  • Strong experience operating Kubernetes in large-scale, production environments
  • Hands‑on experience with NVIDIA GPU ecosystem, including GPU Operator, device plugins, MIG, and DCGM
  • Proficiency in Go or Python for building Kubernetes operators and automation tooling
  • Deep understanding of Kubernetes internals, including CRDs, controllers, RBAC, and scheduling
  • Experience supporting GPU-intensive workloads such as AI/ML training, LLMs, or scientific computing
  • Experience with GitOps, CI/CD pipelines, and infrastructure-as-code practices
  • Familiarity with container networking, including CNI plugins such as NVIDIA CNI or Multus
  • Experience with monitoring and observability tools for cluster and GPU performance

This is a high-impact opportunity to work at the forefront of AI infrastructure, helping build and scale the platforms that power next-generation compute.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Kubernetes Engineer
Senior Kubernetes Engineer

NMC2 • Dallas (TX)

On-site
USD 120,000 - 160,000
Senior Kubernetes Engineer for GPU AI Compute Platform
Senior Kubernetes Engineer for GPU AI Compute Platform

GTN Technical Staffing • Dallas (TX)

On-site
USD 140,000 - 200,000
HPC Solution Architect
HPC Solution Architect

Coda Search│Staffing • Dallas (TX)

Hybrid
USD 120,000 - 160,000
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Senior Kubernetes & GPU Infra Engineer for AI-scale Compute
Senior Kubernetes & GPU Infra Engineer for AI-scale Compute

Kindredventures • United States

On-site
USD 140,000 - 190,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Platform Engineer
Platform Engineer

AMroute LLC • St. Louis (MO)

On-site
USD 15,429,000 - 26,450,000
SRE / Platform Engineer, GPU Infrastructure
SRE / Platform Engineer, GPU Infrastructure

Bake AI • Hillsboro (OR)

On-site
USD 140,000 - 210,000
Senior Kubernetes Engineer — GPU/HPC Orchestration
Senior Kubernetes Engineer — GPU/HPC Orchestration

NorthMark Strategies • Dallas (TX)

On-site
USD 120,000 - 160,000
Company-Paid Lunch Stipend
100% Employer-Paid Medical
401(k) matching
+1