Senior Kubernetes Platform Developer – GPU & AI Infrastructure

GTN Technical Staffing

Dallas (TX)

Hybrid

USD 165,000 - 210,000

Full time

12 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Relocation available
Hybrid work model

Job summary

GTN Technical Staffing is seeking a Senior Kubernetes Platform Developer in Dallas to design and build software powering a GPU-accelerated compute platform for AI/ML and HPC workloads. This role focuses on Kubernetes-native software, extending Kubernetes with operators, CRDs, APIs, and scheduling capabilities, rather than traditional DevOps tasks.

The ideal candidate will deeply understand Kubernetes internals, build on top of Kubernetes, and contribute to large-scale GPU infrastructure.

Qualifications

  • Strong software development experience with Go, Python, or another modern programming language.
  • Hands-on experience building Kubernetes operators, controllers, CRDs, APIs, or other Kubernetes-native software.
  • Deep understanding of Kubernetes architecture, controllers, reconciliation, scheduling, RBAC, networking, and cluster lifecycle.
  • Experience developing platforms or distributed systems built on Kubernetes.
  • Experience with GPU infrastructure and NVIDIA technologies.
  • Experience supporting AI/ML, LLM, HPC, or other compute-intensive workloads.
  • Strong Linux and distributed systems knowledge.
  • Experience with Terraform, Helm, Kustomize, Argo CD, Flux, or similar tooling.
  • Ability to troubleshoot across Kubernetes, compute, networking, storage, GPUs, and applications.

Responsibilities

  • Develop Kubernetes-native software using Go, Python, or similar languages.
  • Build custom operators, controllers, CRDs, APIs, and platform services.
  • Extend Kubernetes to support GPU-intensive AI/ML and HPC workloads.
  • Develop automation for cluster provisioning, lifecycle management, scheduling, and infrastructure orchestration.
  • Build GPU scheduling, allocation, workload placement, and resource-isolation capabilities.
  • Integrate NVIDIA technologies including GPU Operator, device plugins, MIG, and DCGM.
  • Develop internal tools and APIs for provisioning and managing GPU compute resources.
  • Improve platform scalability, GPU utilization, workload performance, and reliability.
  • Integrate Kubernetes with high-performance networking, storage, and bare-metal infrastructure.
  • Build observability and automated remediation capabilities for distributed compute environments.

Skills

Go/Python
Kubernetes-native software
Kubernetes architecture
Distributed systems
NVIDIA GPU infra
AI/ML workloads
Linux systems
Terraform/Helm
Troubleshooting across stack

Tools

Terraform
Helm
Kustomize
Argo CD
Flux

Job description

Senior Kubernetes Platform Developer -- GPU & AI Infrastructure

Location: Dallas, TX preferred

Work Arrangement: Hybrid, 3 days onsite / 2 days remote

Remote Flexibility: Full remote may be considered for the right candidate

Relocation: Available

Employment Type: Direct Hire

Overview

We are seeking a Senior Kubernetes Platform Developer to design and build the software powering a next-generation GPU-accelerated compute platform supporting AI, machine learning, LLM, and HPC workloads.

This is a software development role focused on Kubernetes, not a traditional DevOps, SRE, or Kubernetes administration position.

The core focus is developing Kubernetes-native software including custom operators, controllers, CRDs, APIs, scheduling capabilities, and internal platform services used to orchestrate large-scale GPU infrastructure.

The ideal candidate is a strong developer who understands Kubernetes internals and has experience building software on top of Kubernetes, not simply deploying applications or maintaining clusters.

Key Responsibilities
  • Develop Kubernetes-native software using Go, Python, or similar languages.
  • Build custom operators, controllers, CRDs, APIs, and platform services.
  • Extend Kubernetes to support GPU-intensive AI/ML and HPC workloads.
  • Develop automation for cluster provisioning, lifecycle management, scheduling, and infrastructure orchestration.
  • Build GPU scheduling, allocation, workload placement, and resource-isolation capabilities.
  • Integrate NVIDIA technologies including GPU Operator, device plugins, MIG, and DCGM.
  • Develop internal tools and APIs for provisioning and managing GPU compute resources.
  • Improve platform scalability, GPU utilization, workload performance, and reliability.
  • Integrate Kubernetes with high-performance networking, storage, and bare-metal infrastructure.
  • Build observability and automated remediation capabilities for distributed compute environments.
Required Qualifications
  • Strong software development experience with Go, Python, or another modern programming language.
  • Hands-on experience building Kubernetes operators, controllers, CRDs, APIs, or other Kubernetes-native software.
  • Deep understanding of Kubernetes architecture, controllers, reconciliation, scheduling, RBAC, networking, and cluster lifecycle.
  • Experience developing platforms or distributed systems built on Kubernetes.
  • Experience with GPU infrastructure and NVIDIA technologies.
  • Experience supporting AI/ML, LLM, HPC, or other compute-intensive workloads.
  • Strong Linux and distributed systems knowledge.
  • Experience with Terraform, Helm, Kustomize, Argo CD, Flux, or similar tooling.
  • Ability to troubleshoot across Kubernetes, compute, networking, storage, GPUs, and applications.
Preferred Qualifications
  • Experience with NVIDIA GPU clusters.
  • Experience with Slurm, Volcano, kube-scheduler extensions, or custom scheduling.
  • Familiarity with CUDA, NCCL, PyTorch, or TensorFlow.
  • Experience with InfiniBand, RDMA, RoCE, or high-performance networking.
  • Experience with bare-metal Kubernetes.
  • Experience building internal developer platforms or self-service infrastructure.
  • Background in AI infrastructure, HPC, cloud infrastructure, or large‑scale distributed systems.
Ideal Candidate

The ideal candidate is a platform developer who builds Kubernetes-native systems.

This person should be comfortable writing operators, controllers, APIs, schedulers, and automation that extend Kubernetes and manage complex GPU infrastructure.

Candidates whose experience is primarily DevOps, CI/CD, Terraform administration, application deployment, or Kubernetes operations without substantial software development experience are unlikely to be the right fit.

Dallas-based candidates are preferred, but full remote may be considered for candidates with exceptional Kubernetes development and GPU infrastructure experience.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Kubernetes Developer – GPU & AI Infrastructure
Senior Kubernetes Developer – GPU & AI Infrastructure

GTN Technical Staffing • Town of Texas (WI), Fort Worth (TX)

Hybrid
USD 150,000 - 210,000
Relocation assistance
Hybrid work arrangement
Remote work flexibility
Kubernetes Platform Engineer - GPU & AI Infra
Kubernetes Platform Engineer - GPU & AI Infra

GTN Technical Staffing • Dallas (TX)

Hybrid
USD 165,000 - 210,000
Relocation available
Hybrid work model
Senior Kubernetes Platform Engineer
Senior Kubernetes Platform Engineer

The Brixton Group • Dallas (TX)

Hybrid
USD 138,000 - 165,000
Relocation assistance
Kubernetes Administrator – AI Infrastructure
Kubernetes Administrator – AI Infrastructure

Sira Consulting, an Inc 5000 company • United States

On-site
USD 140,000 - 210,000
Senior Kubernetes Engineer
Senior Kubernetes Engineer

NMC2 • Dallas (TX)

On-site
USD 120,000 - 160,000
Kubernetes-Native GPU AI Platform Engineer
Kubernetes-Native GPU AI Platform Engineer

GTN Technical Staffing • Town of Texas (WI), Fort Worth (TX)

Hybrid
USD 150,000 - 210,000
Relocation assistance
Hybrid work arrangement
Remote work flexibility
Senior Kubernetes Engineer – GPU & HPC Platform
Senior Kubernetes Engineer – GPU & HPC Platform

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 130,000 - 180,000
Company-Paid lunch stipend
Medical, dental, vision benefits
401(k) matching up to 6%
+1
Senior Platform Engineer
Senior Platform Engineer

STN Incorporated • United States

Hybrid
USD 140,000 - 180,000
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Senior Kubernetes Engineer
Senior Kubernetes Engineer

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 130,000 - 180,000
Company-Paid lunch stipend
Medical, dental, vision benefits
401(k) matching up to 6%
+1