Senior Manager, GPU Cluster Engineering & Deployment

TensorWave Inc.

United States

Remote

USD 180,000 - 240,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Stock Options
100% paid Medical, Dental, and Vision
401(k)
Flexible PTO
Parental Leave

Job summary

TensorWave in the United States is seeking a Senior Manager, Cluster Engineering & Deployment to own and run the deployment machine for GPU clusters. You’ll manage network bring-up, cabling validation, and production gate readiness across sites, partnering with Data Center Integration teams and vendors.

You will drive velocity by optimizing bring-up time, define spares and tooling, and lead multiple concurrent builds with schedule accountability, ensuring every cluster meets predefined

Qualifications

  • 10+ years across network deployment, cluster/HPC bring-up, or large-scale infrastructure delivery.
  • Hands-on fabric bring-up experience at scale (hundreds of switches / thousands of links per deployment).
  • Strong operational rigor: building and enforcing playbooks, gates, metrics, and blameless defect loops.
  • Team leadership with schedule accountability across multiple concurrent builds or sites.

Responsibilities

  • Own the cluster deployment playbook and drive its evolution: staged bring-up, automated config push, link/optics validation, cabling verification against L1 port maps, and fault triage during deployment windows.
  • Lead deployment engineering across concurrent cluster builds, through team leads and on-site engineers; coordinate daily with Data Center Integration field teams and cabling vendors.
  • Drive deployment velocity engineering: cut bring-up time per cluster through tooling, pre-staging, and defect-source elimination, and set the targets the team is measured against.
  • Own defect feedback loops to Network Engineering (design), Layer One (cabling quality), and vendors (hardware/optics RMA patterns), holding those partners accountable to resolution.
  • Define spares, test equipment, and deployment tooling requirements per site, and standardize them across sites.

Skills

Network deployment
Cluster bring-up
Large-scale infra
Team leadership
Fabric bring-up
Operational rigor

Tools

Python
Ansible

Job description

TensorWave in the United States is seeking a Senior Manager, Cluster Engineering & Deployment to own and run the deployment machine for GPU clusters. You’ll manage network bring-up, cabling validation, and production gate readiness across sites, partnering with Data Center Integration teams and vendors.

You will drive velocity by optimizing bring-up time, define spares and tooling, and lead multiple concurrent builds with schedule accountability, ensuring every cluster meets predefined

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Manager, Cluster Deployment & HPC Automation
Senior Manager, Cluster Deployment & HPC Automation

TensorWave • Las Vegas (NV)

On-site
USD 180,000 - 240,000
Stock Options
Medical Insurance
Dental Insurance
+7
Senior Manager, Cluster Engineering & Deployment
Senior Manager, Cluster Engineering & Deployment

TensorWave • Las Vegas (NV)

On-site
USD 180,000 - 240,000
Stock Options
Medical Insurance
Dental Insurance
+7
Senior Manager, Cluster Engineering & Deployment
Senior Manager, Cluster Engineering & Deployment

TensorWave Inc. • United States

On-site
USD 180,000 - 240,000
Stock Options
100% paid Medical, Dental, and Vision
401(k)
+2
Backend Engineer - Kubernetes & GPU Cluster Automation
Backend Engineer - Kubernetes & GPU Cluster Automation

TensorWave Inc. • United States

Remote
USD 180,000 - 240,000
Stock Options
Medical, Dental, Vision insurance for
Company Health Savings Account
+7
Backend Platform Engineer – GPU Cluster Automation
Backend Platform Engineer – GPU Cluster Automation

TensorWave • Las Vegas (NV)

On-site
USD 140,000 - 210,000
Stock Options
Excellent health insurance
401(k)
+6
Senior Data Center Build & Design Program Lead
Senior Data Center Build & Design Program Lead

TensorWave • United States

Remote
USD 140,000 - 230,000
Stock Options
Health Insurance
401(k)
+3
Senior Data Center Network Engineer - Automation & On-Call
Senior Data Center Network Engineer - Automation & On-Call

TensorWave • Las Vegas (NV)

On-site
USD 120,000 - 180,000
Stock Options
Medical, Dental, and Vision insurance
401(k)
+3
Design Manager — Hyperscale Data Center Projects
Design Manager — Hyperscale Data Center Projects

TensorWave • Las Vegas (NV)

On-site
USD 120,000 - 180,000
Stock Options
Medical, Dental, Vision insurance for
Health Savings Account Contributions
+8
Data Center Manager — 24/7 Operations Leader
Data Center Manager — 24/7 Operations Leader

TensorWave Inc. • Anniston (AL)

On-site
USD 90,000 - 140,000
Stock Options
Medical Insurance (100% for employees)
401(k)
+3
Data Center Technician — AI Compute Infra, GPUs, On‑Site
Data Center Technician — AI Compute Infra, GPUs, On‑Site

TensorWave • Miami (FL)

On-site
USD 55,000 - 75,000
Stock Options
Medical, Dental, Vision insurance
HSA Contributions
+2