AI Infra TPM: Scale & Reliability Leader

Nscale

New York (NY)

On-site

USD 140,000 - 190,000

Full time

30 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
Competitive base salary
Dynamic startup environment

Job summary

Nscale is seeking a Technical Program Manager to lead operations for our AI infrastructure and HPC fleet in New York. You will drive cross-functional programs spanning data center hardware, software/firmware rollouts, and tooling, with a sharp focus on SLA, uptime, and availability.

Responsibilities include KPI-driven leadership, cross-team coordination, and risk mitigation to ensure scalable, reliable GPU/InfiniBand infrastructure for customer workloads.

Qualifications

  • 5+ years leading large-scale infrastructure or software programs.
  • Strong Linux, data center, and networking fundamentals.
  • Proven metrics-driven with uptime/availability targets (SLOs/SLIs).
  • Experience with Agile/Scrum and PMP is preferred.

Responsibilities

  • Own planning, execution, and delivery of strategic operational programs for AI infrastructure and HPC environments.
  • Develop and track KPIs for Availability and Uptime with dashboards and regular leadership updates.
  • Standardize incident, change, and postmortem processes to reduce toil and MTTR.
  • Coordinate cross-functional efforts between Hardware, Compute Platform, Network, Data Center Ops and external vendors.
  • Translate capacity planning into infrastructure delivery roadmaps and ensure new hardware integrates into the control plane.
  • Identify and mitigate technical, schedule, and resource risks and communicate impacts to stakeholders.

Skills

5+ years TPM
Data center infrastructure
Linux
Networking
Agile/Scrum/PMP
Metrics-driven

Education

Bachelor's or Master's degree

Tools

CI/CD tooling
Automation tooling

Job description

Nscale is seeking a Technical Program Manager to lead operations for our AI infrastructure and HPC fleet in New York. You will drive cross-functional programs spanning data center hardware, software/firmware rollouts, and tooling, with a sharp focus on SLA, uptime, and availability.

Responsibilities include KPI-driven leadership, cross-team coordination, and risk mitigation to ensure scalable, reliable GPU/InfiniBand infrastructure for customer workloads.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infra TPM: Scale, Availability & Ops Leadership
AI Infra TPM: Scale, Availability & Ops Leadership

Greenhouse Software, Inc. • New York (NY), Northern (KY)

Hybrid
USD 150,000 - 230,000
Senior TPM: AI Infrastructure & HPC Operations
Senior TPM: AI Infrastructure & HPC Operations

Nscale • Houston (TX)

On-site
USD 140,000 - 200,000
Base salary + equity
Career growth
Dynamic startup environment
AI Infra TPM: Scale GPU & Inference
AI Infra TPM: Scale GPU & Inference

GMI Cloud • Mountain View (CA)

On-site
USD 150,000 - 230,000
Principal Technical Program Manager (TPM) - AI Infrastructure Operations
Principal Technical Program Manager (TPM) - AI Infrastructure Operations

Nscale • Houston (TX)

On-site
USD 140,000 - 200,000
Base salary + equity
Career growth
Dynamic startup environment
Principal Technical Program Manager (TPM) - AI Infrastructure Operations
Principal Technical Program Manager (TPM) - AI Infrastructure Operations

Nscale • Seattle (WA)

On-site
USD 150,000 - 210,000
Equity
Annual reviews
Principal Technical Program Manager (TPM) - AI Infrastructure Operations
Principal Technical Program Manager (TPM) - AI Infrastructure Operations

Nscale • New York (NY)

On-site
USD 140,000 - 190,000
Equity
Competitive base salary
Dynamic startup environment
Senior Infra TPM: Scale, Reliability & AI Ops
Senior Infra TPM: Scale, Reliability & AI Ops

Glean • San Francisco (CA)

Hybrid
USD 198,000 - 235,500
Medical coverage
Vision coverage
Dental coverage
+5
AI Infrastructure TPM: Scale Hardware for AI Systems
AI Infrastructure TPM: Scale Hardware for AI Systems

Meta • Menlo Park (CA)

On-site
USD 168,000 - 234,000
Cloud Infrastructure TPM - AI & Scale Leadership
Cloud Infrastructure TPM - AI & Scale Leadership

AI Chopping Block, Inc. • Sunnyvale (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Senior TPM - AI Infrastructure & GPU Deployments
Senior TPM - AI Infrastructure & GPU Deployments

Hamilton Barnes Associates Limited • New York (NY)

On-site
USD 200,000 - 240,000
Bonus
Equity