Senior Kubernetes Engineer

NorthMark Compute & Cloud

Dallas (TX)

On-site

USD 130,000 - 180,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Company-Paid lunch stipend
Medical, dental, vision benefits
401(k) matching up to 6%
Paid parental leave

Job summary

NorthMark Compute & Cloud in Dallas seeks a Senior Kubernetes Engineer to design and operate GPU-accelerated container platforms at scale for AI/ML, HPC and LLM training. You will enable high-performance workloads across hybrid or on-prem environments with deep NVIDIA and Kubernetes expertise.

You will implement GPU scheduling, device plugins, MIG, and custom operators, while ensuring security, observability, and CI/CD through GitOps workflows and IaC tooling.

Qualifications

  • Extensive Kubernetes production experience with GPU workloads.
  • Experience with NVIDIA GPU Operator and device plugins.
  • Proficiency in Go or Python for operator development.
  • Knowledge of CRDs, RBAC, and scheduler extensions.

Responsibilities

  • Architect Kubernetes clusters optimized for GPU workloads in on-prem or hybrid environments.
  • Develop and maintain custom operators and controllers for automation.
  • Integrate NVIDIA device plugins and MIG into scheduling layer.
  • Tune scheduling with kube-scheduler plugins, Slurm, and Volcano.
  • Collaborate with HPC/ML/DevOps teams for multi-tenant, high-throughput clusters.
  • Drive observability with Prometheus, Grafana, OpenTelemetry.
  • Enforce security using RBAC, OPA, or Gatekeeper.
  • Maintain CI/CD pipelines for Kubernetes using GitOps (ArgoCD/FluxCD).
  • Contribute to IaC using Terraform, Helm, and Kustomize.

Skills

Kubernetes
GPU scheduling
Go or Python

Tools

NVIDIA GPU Operator
NVML
MIG
DCGM
Prometheus
Grafana
OpenTelemetry
ArgoCD
FluxCD
Terraform
Helm
Kustomize

Job description

The Company

NorthMark Compute & Cloud (NMC²) is backed by dedicated leadership and investment, with a clear mission as it operates at the bleeding edge of technology. Its goal is to scale and enhance the high-performance computing (HPC) and cloud infrastructure that supports its clients' research, production, and delivery, enabling breakthroughs that shape the industries of tomorrow. Its engineers build critical infrastructure to eliminate friction in scientific research, simulations, analysis, and decision-making, accelerating discovery and driving faster innovation.

The Position

We are seeking a highly skilled Senior Kubernetes Engineer to join our NMC2 office in Dallas.

In this role, you will design, implement, and optimise GPU-accelerated container platforms at scale, enabling high-performance workloads (AI/ML, HPC, LLM training) across hybrid or on-prem environments.

You will have deep expertise with both NVIDIA and Kubernetes ecosystems, including GPU scheduling, device plugins and custom operators.

Responsibilities
  • Architecting and operating Kubernetes clusters optimised for GPU workloads, leveraging NVIDIA GPU Operator, Network Operator and DCGM
  • Developing, deploying and maintaining custom Kubernetes operators and controllers to automate infrastructure services
  • Integrating NVIDIA device plugins, Multi-Instance GPU (MIG) and GPU sharing features into the scheduling layer
  • Optimising GPU utilisation and job placement through scheduler extensions, such as kube-scheduler plugins, Slurm and Volcano
  • Collaborating with HPC, ML and DevOps teams to ensure multi-tenant, high-throughput cluster performance
  • Driving observability and telemetry integrations using Prometheus, Grafana, DCGM Exporter and OpenTelemetry
  • Implementing secure multi-user and multi-namespace GPU isolation, with RBAC and policy enforcement, such as OPA or Gatekeeper
  • Maintaining CI/CD pipelines for Kubernetes infrastructure using GitOps, ArgoCD and FluxCD
  • Contributing to infrastructure-as-code, using Terraform, Helm, and Kustomize
  • Participating in performance tuning, incident response and production readiness reviews
Requirements
  • Extensive experience with Kubernetes in production-grade environments and working with NVIDIA and Kubernetes, including GPU Operator, device plugin, NVML, MIG and DCGM
  • Proficiency in Go or Python for operator development and Kubernetes controller logic
  • Deep understanding of Kubernetes internals, including CRDs, RBAC, custom controllers and scheduler extensions
  • Experience with GPU-intensive workloads, for example for LLMs, training pipelines and scientific computing
  • Hands-on experience with Helm, Kustomize and GitOps workflows
  • Familiarity with CNI plugins, especially NVIDIA CNI and Multus
  • Experience with monitoring GPU metrics and cluster health using Prometheus and DCGM Exporter

It is impossible to list every requirement for, or responsibility of, any position. Similarly, we cannot identify all the skills a position may require since job responsibilities and the Company’s needs may change over time. Therefore, the above job description is not comprehensive or exhaustive. The Company reserves the right to adjust, add to or eliminate any aspect of the above description. The Company also retains the right to require all employees to undertake additional or different job responsibilities when necessary to meet business needs.

Must be legally authorized to work in the United States without the need for employer sponsorship, now or at any time in the future.

Benefits & Perks
  • Company-Paid Lunch Stipend: Lunch is provided via GrubHub
  • Company-Paid Benefits: 100% Employer-Paid Medical in our High Deductible Health Plan, Dental and Vision benefits for employees and their families, 16 weeks of Paid Parental Leave, Employee Assistance Program, Life insurance, Short-Term Disability and Long-Term Disability
  • 401(k): Company will match 100% of your contributions up to 6%
  • Optional Employee-Paid Benefits: Medical insurance in our PPO plan and a variety of other benefits such as Health Savings Accounts (with Company Contribution!), Flexible Spending Accounts, Supplemental Life Insurance, Wellhub and more.
  • Time Off: 25 days of Paid Time Off plus 12 company holidays
EQUAL OPPORTUNITY EMPLOYER

NORTHMARK STRATEGIES LLC IS AN EQUAL EMPLOYMENT OPPORTUNITY EMPLOYER. THE COMPANY'S POLICY IS NOT TO DISCRIMINATE AGAINST ANY APPLICANT OR EMPLOYEE BASED ON RACE, COLOR, RELIGION, NATIONAL ORIGIN, GENDER, AGE, SEXUAL ORIENTATION, GENDER IDENTITY OR EXPRESSION, MARITAL STATUS, MENTAL OR PHYSICAL DISABILITY, AND GENETIC INFORMATION, OR ANY OTHER BASIS PROTECTED BY APPLICABLE LAW. THE FIRM ALSO PROHIBITS HARASSMENT OF APPLICANTS OR EMPLOYEES BASED ON ANY OF THESE PROTECTED CATEGORIES.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Kubernetes Engineer
Senior Kubernetes Engineer

NorthMark Strategies • United States

On-site
USD 120,000 - 160,000
Company-Paid Lunch Stipend
100% Employer-Paid Medical
401(k) matching
+1
Senior Kubernetes Engineer
Senior Kubernetes Engineer

NorthMark Strategies LLC • Dallas (TX)

On-site
USD 150,000 - 210,000
Company-Paid Lunch Stipend
100% Employer-Paid Medical
401(k) Company Match
+1
Senior Kubernetes Engineer
Senior Kubernetes Engineer

NMC2 • Dallas (TX)

On-site
USD 120,000 - 160,000
Senior HPC Hardware Engineer
Senior HPC Hardware Engineer

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 140,000 - 220,000
Lunch stipend
Medical benefits - employer-paid
Dental & Vision benefits
+6
Senior HPC Hardware Engineer
Senior HPC Hardware Engineer

NorthMark Strategies • United States

On-site
USD 140,000 - 190,000
Lunch stipend
Medical benefits (employer-paid)
Parental leave 16 weeks
+2
Software Engineer, Fleet Automation
Software Engineer, Fleet Automation

NorthMark Strategies • Town of Texas (WI)

On-site
USD 120,000 - 180,000
Lunch stipend
Employer-paid health & dental & vision
Parental leave 16 weeks
+4
HPC Orchestration Architect
HPC Orchestration Architect

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 180,000 - 260,000
Lunch stipend
Employer-paid medical benefits
Parental leave (16 weeks)
+2
Senior HPC Hardware Engineer
Senior HPC Hardware Engineer

NorthMark Compute and Cloud LLC • United States

On-site
USD 150,000 - 230,000
Lunch stipend
Medical benefits
Dental & Vision
+7
Senior Kubernetes Engineer
Senior Kubernetes Engineer

GTN Technical Staffing • Dallas (TX)

On-site
USD 140,000 - 200,000
Senior HPC Hardware Engineer
Senior HPC Hardware Engineer

NorthMark Strategies LLC • Dallas (TX), Northern (KY)

Hybrid
USD 140,000 - 190,000
Lunch stipend
Company-paid medical benefits
Dental and Vision for employees and 가족
+6