Backend Engineer - Kubernetes & GPU Cluster Automation

TensorWave Inc.

United States

Remote

USD 180,000 - 240,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Stock Options
Medical, Dental, Vision insurance for
Company Health Savings Account
Short/Long Term Disability Insurance
Life and Voluntary Insurance
Pet & Legal Insurance
Flexible PTO
Paid Holidays
Parental Leave
401(k)

Job summary

TensorWave Inc. is seeking a Software Engineer (Back-end) to join our platform team.

You will own end-to-end automation of provisioning, configuring, and operating large-scale GPU clusters across bare metal, Kubernetes, and Slurm environments, building tooling and pipelines for hundreds of GPU nodes. This is a hands-on role working closely with cross-functional partners to support business objectives while upholding our standards for excellence, collaboration, and impact.

Qualifications

  • 5+ years in infrastructure or platform engineering.
  • 3+ years writing production Go.
  • Deep understanding of Kubernetes internals, including informers and work queues, controller-runtime and client-go, CRDs, custom controllers, and operators.
  • Experience building Kubernetes Operators.
  • Experience building gRPC and REST APIs in Go at production scale.

Responsibilities

  • Build and maintain fully automated pipelines for provisioning bare metal GPU clusters from zero to production.
  • Automate Slurm and Kubernetes cluster lifecycle—bootstrapping, upgrades, node provisioning, and decommissioning at scale.
  • Develop and maintain infrastructure for GPU node configuration, including drivers and firmware.
  • Own cluster validation pipelines, automating health checks and GPU burn-in tests.
  • Build day-2 operations automation, including node remediation, rolling upgrades, and automated drain/cordon workflows.
  • Write and maintain runbooks and documentation to enable reliable, repeatable operations.
  • Own the full observability stack for automation services, provisioning pipelines, and cluster health systems.

Skills

Go
Kubernetes internals
Prometheus
Grafana
OpenTelemetry
Loki
gRPC
REST APIs
Bare metal provisioning
DevOps automation

Tools

ArgoCD
GitHub Actions
Argo Workflows
Terraform
Ansible

Job description

TensorWave Inc. is seeking a Software Engineer (Back-end) to join our platform team.

You will own end-to-end automation of provisioning, configuring, and operating large-scale GPU clusters across bare metal, Kubernetes, and Slurm environments, building tooling and pipelines for hundreds of GPU nodes. This is a hands-on role working closely with cross-functional partners to support business objectives while upholding our standards for excellence, collaboration, and impact.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Backend Platform Engineer – GPU Cluster Automation
Backend Platform Engineer – GPU Cluster Automation

TensorWave • Las Vegas (NV)

On-site
USD 140,000 - 210,000
Stock Options
Excellent health insurance
401(k)
+6
Software Engineer (Back-end)
Software Engineer (Back-end)

TensorWave • Las Vegas (NV)

On-site
USD 140,000 - 210,000
Stock Options
Excellent health insurance
401(k)
+6
Senior Manager, GPU Cluster Engineering & Deployment
Senior Manager, GPU Cluster Engineering & Deployment

TensorWave Inc. • United States

Remote
USD 180,000 - 240,000
Stock Options
100% paid Medical, Dental, and Vision
401(k)
+2
Kubernetes Platform Architect – Staff Engineer
Kubernetes Platform Architect – Staff Engineer

TensorWave Inc. • United States

Remote
USD 180,000 - 240,000
Stock Options
Medical, Dental, and Vision insurance
Health Savings Account Contributions
+9
Senior Software Engineer: GPU Infra, Kubernetes & GitOps
Senior Software Engineer: GPU Infra, Kubernetes & GitOps

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 224,000 - 357,000
Senior Kubernetes & GPU Infra Engineer for AI-scale Compute
Senior Kubernetes & GPU Infra Engineer for AI-scale Compute

Kindredventures • United States

On-site
USD 140,000 - 190,000
Kubernetes-Native GPU AI Platform Engineer
Kubernetes-Native GPU AI Platform Engineer

GTN Technical Staffing • Town of Texas (WI), Fort Worth (TX)

Hybrid
USD 150,000 - 210,000
Relocation assistance
Hybrid work arrangement
Remote work flexibility
GPU Compute Infrastructure Engineer
GPU Compute Infrastructure Engineer

Linuxcareers • San Francisco (CA), Northern (KY)

Hybrid
USD 210,000 - 270,000
Platform Engineer - GPU Infra & Kubernetes
Platform Engineer - GPU Infra & Kubernetes

Together AI • San Francisco (CA)

On-site
USD 160,000 - 280,000
Equity
Health insurance
Competitive benefits
Level 2 Technical Support Engineer - GPU/AI Infra
Level 2 Technical Support Engineer - GPU/AI Infra

AI Chopping Block • Las Vegas (NV), Northern (KY)

Hybrid
USD 90,000 - 130,000
Stock Options
Medical Insurance (Employee)
Dental Insurance (Employee)
+13