Senior TPM: AI Infrastructure & HPC Operations

Nscale

Houston (TX)

On-site

USD 140,000 - 200,000

Full time

29 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Base salary + equity
Career growth
Dynamic startup environment

Job summary

Nscale is seeking a Technical Program Manager to lead AI infrastructure and HPC operations in a high-scale data center environment in Houston. You will drive cross-functional programs spanning hardware, network, software, and vendor partnerships to ensure stability and growth of GPU fleets and InfiniBand networks.

You will implement robust KPI tracking (availability, uptime) and dashboards, mentor teams on incident and change processes, and shape readiness roadmaps for new hardware deployments.

Qualifications

  • 5+ years in a Technical Program Management role driving large-scale infrastructure or software programs.
  • Strong foundational knowledge of data center infrastructure, distributed systems, Linux, and networking concepts.
  • Experience with modern program management methodologies (Agile, Scrum); PMP is a plus; excellent communication and presentation skills.

Responsibilities

  • Lead planning, execution, and delivery of strategic operational programs for AI infrastructure and HPC environments.
  • Define and track KPIs (Availability, Uptime) and build dashboards for real-time leadership visibility.
  • Standardize incident, change, and postmortem processes to reduce toil and MTTR.
  • Coordinate cross-functionally between Hardware, Compute Platform, Network, and Data Center Operations; manage dependencies and risks.
  • Translate capacity planning into delivery roadmaps; ensure new hardware is integrated into the control plane.
  • Identify technical, schedule, and resource risks; communicate impacts to stakeholders.

Skills

Technical Program Management
Data Center Infrastructure
Linux & Networking
Agile / PMP
Stakeholder Communication
SRE / CI–CD

Education

Bachelor's or Master's in CS/Engineering

Tools

CI/CD tooling
Networking tooling
SRE tooling

Job description

Nscale is seeking a Technical Program Manager to lead AI infrastructure and HPC operations in a high-scale data center environment in Houston. You will drive cross-functional programs spanning hardware, network, software, and vendor partnerships to ensure stability and growth of GPU fleets and InfiniBand networks.

You will implement robust KPI tracking (availability, uptime) and dashboards, mentor teams on incident and change processes, and shape readiness roadmaps for new hardware deployments.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infra TPM: Scale & Reliability Leader
AI Infra TPM: Scale & Reliability Leader

Nscale • New York (NY)

On-site
USD 140,000 - 190,000
Equity
Competitive base salary
Dynamic startup environment
AI Infra TPM: Scale, Availability & Ops Leadership
AI Infra TPM: Scale, Availability & Ops Leadership

Greenhouse Software, Inc. • New York (NY), Northern (KY)

Hybrid
USD 150,000 - 230,000
Senior TPM - AI Infrastructure & GPU Deployments
Senior TPM - AI Infrastructure & GPU Deployments

Hamilton Barnes Associates Limited • New York (NY)

On-site
USD 200,000 - 240,000
Bonus
Equity
Principal Technical Program Manager (TPM) - AI Infrastructure Operations
Principal Technical Program Manager (TPM) - AI Infrastructure Operations

Nscale • Houston (TX)

On-site
USD 140,000 - 200,000
Base salary + equity
Career growth
Dynamic startup environment
Principal Technical Program Manager (TPM) - AI Infrastructure Operations
Principal Technical Program Manager (TPM) - AI Infrastructure Operations

Nscale • New York (NY)

On-site
USD 140,000 - 190,000
Equity
Competitive base salary
Dynamic startup environment
Principal Technical Program Manager (TPM) - AI Infrastructure Operations
Principal Technical Program Manager (TPM) - AI Infrastructure Operations

Nscale • Seattle (WA)

On-site
USD 150,000 - 210,000
Equity
Annual reviews
Principal Technical Program Manager (TPM) - AI Infrastructure Operations New Houston; New York; San Francisco; Seattle
Principal Technical Program Manager (TPM) - AI Infrastructure Operations New Houston; New York; San Francisco; Seattle

Greenhouse Software, Inc. • New York (NY), Northern (KY)

Hybrid
USD 150,000 - 230,000
Senior Datacenter AI Systems TPM
Senior Datacenter AI Systems TPM

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 168,000 - 322,000
Equity
Benefits
Staff TPM: AI Infrastructure Deployments & GPU Cloud
Staff TPM: AI Infrastructure Deployments & GPU Cloud

Crusoe • Bellevue (WA)

On-site
USD 200,000 - 240,000
Equity
Paid time off
Health insurance
+2
Remote TPM for AI/GPU Infrastructure Deployments
Remote TPM for AI/GPU Infrastructure Deployments

5C Group • United States

On-site
USD 155,000 - 175,000