Senior Kubernetes Engineer

NorthMark Compute & Cloud

Dallas (TX)

On-site

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NorthMark Compute & Cloud in Dallas seeks a Senior Kubernetes Engineer to design, implement, and optimise GPU-accelerated container platforms at scale, spanning hybrid or on-prem environments. The role requires deep NVIDIA and Kubernetes expertise, including GPU Operator, device plugins, MIG, and advanced scheduling.

You will drive multi-tenant, high-throughput workloads for AI/ML, HPC and LLM training, while enhancing observability and security with GitOps-driven delivery.

Qualifications

  • Experience operating Kubernetes in production environments with NVIDIA GPUs (Operator, device plugin, MIG).
  • Proficiency in Go or Python for operator/controllers development.
  • Deep understanding of Kubernetes internals: CRDs, RBAC, custom controllers and scheduler extensions.
  • Experience with GPU-intensive workloads (LLMs, HPC, ML pipelines).
  • Hands-on with Helm, Kustomize and GitOps workflows.

Responsibilities

  • Architect and operate GPU-optimised Kubernetes clusters using NVIDIA Operator, DCGM, and network tools.
  • Develop and maintain custom Kubernetes operators and controllers to automate infra services.
  • Integrate NVIDIA device plugins, MIG, and GPU sharing into scheduling layer.
  • Tune scheduling with kube-scheduler plugins, Slurm, Volcano for high throughput.

Skills

Kubernetes in prod
Go or Python
RBAC & security
GPU workloads
GitOps workflows

Education

Bachelor's degree in CS or related

Tools

NVIDIA GPU Operator
DCGM
Helm
Kustomize
ArgoCD
FluxCD
Terraform
Prometheus
Grafana

Job description

The Company
NorthMark Compute & Cloud (NMC²) is backed by dedicated leadership and investment, with a clear mission as it operates at the bleeding edge of technology. Its goal is to scale and enhance the high-performance computing (HPC) and cloud infrastructure that supports its clients' research, production, and delivery, enabling breakthroughs that shape the industries of tomorrow. Its engineers build critical infrastructure to eliminate friction in scientific research, simulations, analysis, and decision-making, accelerating discovery and driving faster innovation.
The Position
We are seeking a highly skilled Senior Kubernetes Engineer to join our NMC2 office in Dallas.
In this role, you will design, implement, and optimise GPU-accelerated container platforms at scale, enabling high-performance workloads (AI/ML, HPC, LLM training) across hybrid or on-prem environments.
You will have deep expertise with both NVIDIA and Kubernetes ecosystems, including GPU scheduling, device plugins and custom operators.

Responsibilities
  • Architecting and operating Kubernetes clusters optimised for GPU workloads, leveraging NVIDIA GPU Operator, Network Operator and DCGM
  • Developing, deploying and maintaining custom Kubernetes operators and controllers to automate infrastructure services
  • Integrating NVIDIA device plugins, Multi-Instance GPU (MIG) and GPU sharing features into the scheduling layer
  • Optimising GPU utilisation and job placement through scheduler extensions, such as kube-scheduler plugins, Slurm and Volcano
  • Collaborating with HPC, ML and DevOps teams to ensure multi-tenant, high-throughput cluster performance
  • Driving observability and telemetry integrations using Prometheus, Grafana, DCGM Exporter and OpenTelemetry
  • Implementing secure multi-user and multi-namespace GPU isolation, with RBAC and policy enforcement, such as OPA or Gatekeeper
  • Maintaining CI/CD pipelines for Kubernetes infrastructure using GitOps, ArgoCD and FluxCD
  • Contributing to infrastructure-as-code, using Terraform, Helm, and Kustomize
  • Participating in performance tuning, incident response and production readiness reviews
Requirements
  • Extensive experience with Kubernetes in production-grade environments and working with NVIDIA and Kubernetes, including GPU Operator, device plugin, NVML, MIG and DCGM
  • Proficiency in Go or Python for operator development and Kubernetes controller logic
  • Deep understanding of Kubernetes internals, including CRDs, RBAC, custom controllers and scheduler extensions
  • Experience with GPU-intensive workloads, for example for LLMs, training pipelines and scientific computing
  • Hands-on experience with Helm, Kustomize and GitOps workflows
  • Familiarity with CNI plugins, especially NVIDIA CNI and Multus
  • Experience with monitoring GPU metrics and cluster health using Prometheus and DCGM Exporter
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Kubernetes Engineer
Senior Kubernetes Engineer

NMC2 • Dallas (TX)

On-site
USD 120,000 - 160,000
Senior Kubernetes Engineer
Senior Kubernetes Engineer

NorthMark Strategies • Dallas (TX)

On-site
USD 120,000 - 160,000
Company-Paid Lunch Stipend
100% Employer-Paid Medical
401(k) matching
+1
Senior Kubernetes Engineer
Senior Kubernetes Engineer

GTN Technical Staffing • Dallas (TX)

Hybrid
USD 150,000 - 210,000
Performance bonus
Senior Kubernetes Engineer — GPU HPC & AI Platforms
Senior Kubernetes Engineer — GPU HPC & AI Platforms

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 140,000 - 210,000
HPC Solution Architect
HPC Solution Architect

Coda Search│Staffing • Dallas (TX)

Hybrid
USD 120,000 - 160,000
Senior Kubernetes Engineer — GPU/HPC Orchestration
Senior Kubernetes Engineer — GPU/HPC Orchestration

NorthMark Strategies • Dallas (TX)

On-site
USD 120,000 - 160,000
Company-Paid Lunch Stipend
100% Employer-Paid Medical
401(k) matching
+1
Senior Kubernetes Engineer
Senior Kubernetes Engineer

NorthMark Strategies LLC • Dallas (TX)

On-site
USD 150,000 - 210,000
Company-Paid Lunch Stipend
100% Employer-Paid Medical
401(k) Company Match
+1
Senior Kubernetes Engineer for AI/ML GPU Compute Platform
Senior Kubernetes Engineer for AI/ML GPU Compute Platform

GTN Technical Staffing • Dallas (TX)

Hybrid
USD 150,000 - 210,000
Performance bonus
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Senior Software Engineer, Cloud-Native Stack – CSP Engagements

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 288,000