Senior Kubernetes Engineer

NMC2

Dallas (TX)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading technology company in Dallas is searching for a Senior Kubernetes Engineer to design and manage GPU-accelerated container platforms. The ideal candidate will have extensive Kubernetes experience, expertise in NVIDIA technologies, and a strong background in Go or Python. Responsibilities include maintaining Kubernetes clusters, developing operators, and optimising GPU workloads for AI/ML and HPC applications. This role offers a unique opportunity to drive technological advancements and innovation in high-performance computing.

Qualifications

  • Extensive experience with Kubernetes in production-grade environments.
  • Proficiency in Go or Python for operator development.
  • Deep understanding of Kubernetes internals and custom controllers.

Responsibilities

  • Architecting and operating Kubernetes clusters for GPU workloads.
  • Developing and maintaining custom Kubernetes operators.
  • Optimising GPU utilisation and job placement.

Skills

Extensive experience with Kubernetes
Proficiency in Go or Python
Deep understanding of Kubernetes internals
Experience with GPU-intensive workloads
Hands-on experience with Helm
Familiarity with CNI plugins
Experience with monitoring GPU metrics

Tools

Kubernetes
NVIDIA GPU Operator
Helm
Terraform

Job description

Overview

NorthMark Compute & Cloud (NMC²) is backed by dedicated leadership and investment, with a clear mission as it operates at the bleeding edge of technology. Its goal is to scale and enhance the high-performance computing (HPC) and cloud infrastructure that supports its clients' research, production, and delivery, enabling breakthroughs that shape the industries of tomorrow. Its engineers build critical infrastructure to eliminate friction in scientific research, simulations, analysis, and decision-making, accelerating discovery and driving faster innovation.

The Position

We are seeking a highly skilled Senior Kubernetes Engineer to join our NMC2 office in Dallas. In this role, you will design, implement, and optimise GPU-accelerated container platforms at scale, enabling high-performance workloads (AI/ML, HPC, LLM training) across hybrid or on-prem environments. You will have deep expertise with both NVIDIA and Kubernetes ecosystems, including GPU scheduling, device plugins and custom operators.

Responsibilities
  • Architecting and operating Kubernetes clusters optimised for GPU workloads, leveraging NVIDIA GPU Operator, Network Operator and DCGM
  • Developing, deploying and maintaining custom Kubernetes operators and controllers to automate infrastructure services
  • Integrating NVIDIA device plugins, Multi-Instance GPU (MIG) and GPU sharing features into the scheduling layer
  • Optimising GPU utilisation and job placement through scheduler extensions, such as kube-scheduler plugins, Slurm and Volcano
  • Collaborating with HPC, ML and DevOps teams to ensure multi-tenant, high-throughput cluster performance
  • Driving observability and telemetry integrations using Prometheus, Grafana, DCGM Exporter and OpenTelemetry
  • Implementing secure multi-user and multi- namespace GPU isolation, with RBAC and policy enforcement, such as OPA or Gatekeeper
  • Maintaining CI/CD pipelines for Kubernetes infrastructure using GitOps, ArgoCD and FluxCD
  • Contributing to infrastructure-as-code, using Terraform, Helm, and Kustomize
  • Participating in performance tuning, incident response and production readiness reviews
Requirements
  • Extensive experience with Kubernetes in production-grade environments and working with NVIDIA and Kubernetes, including GPU Operator, device plugin, NVML, MIG and DCGM
  • Proficiency in Go or Python for operator development and Kubernetes controller logic
  • Deep understanding of Kubernetes internals, including CRDs, RBAC, custom controllers and scheduler extensions
  • Experience with GPU-intensive workloads, for example for LLMs, training pipelines and scientific computing
  • Hands-on experience with Helm, Kustomize and GitOps workflows
  • Familiarity with CNI plugins, especially NVIDIA CNI and Multus
  • Experience with monitoring GPU metrics and cluster health using Prometheus and DCGM Exporter
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Kubernetes Engineer
Senior Kubernetes Engineer

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 140,000 - 210,000
Senior Kubernetes Engineer
Senior Kubernetes Engineer

NorthMark Strategies • Dallas (TX)

On-site
USD 120,000 - 160,000
Company-Paid Lunch Stipend
100% Employer-Paid Medical
401(k) matching
+1
Senior Kubernetes Engineer — GPU HPC & AI Platforms
Senior Kubernetes Engineer — GPU HPC & AI Platforms

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 140,000 - 210,000
HPC Solution Architect
HPC Solution Architect

Coda Search│Staffing • Dallas (TX)

Hybrid
USD 120,000 - 160,000
Senior Kubernetes Engineer — GPU/HPC Orchestration
Senior Kubernetes Engineer — GPU/HPC Orchestration

NorthMark Strategies • Dallas (TX)

On-site
USD 120,000 - 160,000
Company-Paid Lunch Stipend
100% Employer-Paid Medical
401(k) matching
+1
Senior Kubernetes Engineer
Senior Kubernetes Engineer

NorthMark Strategies LLC • Dallas (TX)

On-site
USD 150,000 - 210,000
Company-Paid Lunch Stipend
100% Employer-Paid Medical
401(k) Company Match
+1
GPUaaS Kubernetes Platform Engineer
GPUaaS Kubernetes Platform Engineer

Veriipro • Irving (TX)

On-site
USD 140,000 - 180,000
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Software Engineer, Fleet Automation
Software Engineer, Fleet Automation

NMC2 • Dallas (TX), Northern (KY)

Hybrid
USD 120,000 - 150,000
Software Engineer, Fleet Automation
Software Engineer, Fleet Automation

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 110,000 - 160,000