Principal Technical Program Manager (TPM) - AI Infrastructure Operations

Nscale

Houston (TX)

On-site

USD 140,000 - 200,000

Full time

32 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Base salary + equity
Career growth
Dynamic startup environment

Job summary

Nscale is seeking a Technical Program Manager to lead AI infrastructure and HPC operations in a high-scale data center environment in Houston. You will drive cross-functional programs spanning hardware, network, software, and vendor partnerships to ensure stability and growth of GPU fleets and InfiniBand networks.

You will implement robust KPI tracking (availability, uptime) and dashboards, mentor teams on incident and change processes, and shape readiness roadmaps for new hardware deployments.

Qualifications

  • 5+ years in a Technical Program Management role driving large-scale infrastructure or software programs.
  • Strong foundational knowledge of data center infrastructure, distributed systems, Linux, and networking concepts.
  • Experience with modern program management methodologies (Agile, Scrum); PMP is a plus; excellent communication and presentation skills.

Responsibilities

  • Lead planning, execution, and delivery of strategic operational programs for AI infrastructure and HPC environments.
  • Define and track KPIs (Availability, Uptime) and build dashboards for real-time leadership visibility.
  • Standardize incident, change, and postmortem processes to reduce toil and MTTR.
  • Coordinate cross-functionally between Hardware, Compute Platform, Network, and Data Center Operations; manage dependencies and risks.
  • Translate capacity planning into delivery roadmaps; ensure new hardware is integrated into the control plane.
  • Identify technical, schedule, and resource risks; communicate impacts to stakeholders.

Skills

Technical Program Management
Data Center Infrastructure
Linux & Networking
Agile / PMP
Stakeholder Communication
SRE / CI–CD

Education

Bachelor's or Master's in CS/Engineering

Tools

CI/CD tooling
Networking tooling
SRE tooling

Job description

As a Technical Program Manager (TPM) for AI Infrastructure Operations, you will be the operational backbone of our high-scale, high-performance AI and High-Performance Computing (HPC) environment. You will be responsible for driving complex, cross-functional programs that ensure the stability, availability, and growth of our cutting-edge GPU fleet and Infiniband network fabrics. This role requires a blend of deep technical understanding, rigorous program management, and a relentless focus on delivering against key operational metrics (SLAs, Uptime, Availability). You will bridge the gap between engineering execution and strategic business goals, directly impacting our ability to serve customer workloads at scale.

Key Responsibilities
  • Program Leadership: Own the planning, execution, and delivery of strategic operational programs, including new data center AI infrastructure build-outs, large-scale fleet software/firmware rollouts, and the implementation of new operational tooling (in partnership with SRE).
  • Metrics and Reporting: Establish, track, and drive accountability against critical infrastructure KPIs, specifically focusing on Availability (Target 97.5%) and Uptime (Target 99%). Develop clear dashboards and communication rhythms to provide leadership with real-time visibility into operational health, program status, and risk.
  • Process Engineering: Analyze and optimize operational workflows across Fleet Operations, Network Operations, and SRE teams. Drive the standardization of incident management, change management, and postmortem processes to reduce toil and improve Mean Time to Recovery (MTTR).
  • Cross-Functional Coordination: Serve as the primary liaison between engineering teams (Hardware, Compute Platform, Network), Data Center Operations, and external vendors (GPU, Network hardware). Proactively identify and resolve dependencies, risks, and roadblocks.
  • Capacity and Readiness: Partner with Data Science/Operation Programs to translate capacity planning models into actionable infrastructure delivery and readiness roadmaps. Ensure that new hardware (GPUs, NICs, switches) is successfully integrated into the operational control plane and meets go-live criteria.
  • Risk Management: Proactively identify technical, schedule, and resource risks related to AI infrastructure scaling and stability. Develop mitigation strategies and communicate impacts clearly to stakeholders.
Required Qualifications
  • Experience: 5+ years of experience in a Technical Program Management role, successfully driving large-scale, complex infrastructure or software engineering programs.
  • Technical Domain Knowledge: Strong foundational understanding of data center infrastructure, distributed systems, Linux, and networking concepts.
  • Program Management Rigor: Proven expertise in modern program management methodologies (Agile, Scrum, PMP certification preferred). Exceptional organizational, communication, and presentation skills.
  • Metrics-Driven Approach: Demonstrable experience in defining, tracking, and improving system performance based on operational metrics (e.g., Uptime, Availability, MTTR, SLOs/SLIs).
  • Execution in Ambiguity: Ability to thrive in a fast-paced, high-growth environment, managing multiple priorities and adapting to evolving technical requirements.
Preferred Qualifications
  • Direct experience managing programs related to data center infrastructure build-outs and hardware commissioning processes.
  • Specific domain knowledge of AI/HPC infrastructure, including NVIDIA GPUs, InfiniBand/RDMA networks, and the challenges of tightly-coupled systems.
  • Experience in a hyperscale or public cloud environment supporting 24/7 mission-critical services.
  • Familiarity with SRE principles, automation tooling, and continuous integration/continuous deployment (CI/CD) pipelines for infrastructure.
  • A Bachelor's or Master's degree in a technical field (Computer Science, Engineering, etc.) or equivalent practical experience.
What We Can Offer You
  • Highly competitive package (base + equity) with reviews every 12 months.
  • Join the fastest-growing tech startup, your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI.
  • Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support.

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Technical Program Manager (TPM) - AI Infrastructure Operations
Principal Technical Program Manager (TPM) - AI Infrastructure Operations

Nscale • New York (NY)

On-site
USD 140,000 - 190,000
Equity
Competitive base salary
Dynamic startup environment
Principal Technical Program Manager (TPM) - AI Infrastructure Operations
Principal Technical Program Manager (TPM) - AI Infrastructure Operations

Nscale • Seattle (WA)

On-site
USD 150,000 - 210,000
Equity
Annual reviews
Principal Technical Program Manager (TPM) - AI Infrastructure Operations New Houston; New York; San Francisco; Seattle
Principal Technical Program Manager (TPM) - AI Infrastructure Operations New Houston; New York; San Francisco; Seattle

Greenhouse Software, Inc. • New York (NY), Northern (KY)

Hybrid
USD 150,000 - 230,000
Technical Program Manager, AI Physical Deployment
Technical Program Manager, AI Physical Deployment

Greenhouse Software, Inc. • Seattle (WA)

Remote
USD 190,000 - 236,000
Equity package
Remote-first culture
Annual performance reviews
Principal Technical Program Manager, AI Physical Deployment
Principal Technical Program Manager, AI Physical Deployment

Nscale • Seattle (WA)

Hybrid
USD 250,000 - 303,000
Base + equity
Flexible work arrangements
Career growth plan
+1
Infrastructure Software Engineer, Fleet & Automation New Houston; New York; San Francisco; Seattle
Infrastructure Software Engineer, Fleet & Automation New Houston; New York; San Francisco; Seattle

Nscale • New York (NY)

On-site
USD 140,000 - 210,000
Competitive package
Equity
Growth opportunities
Technical Program Manager
Technical Program Manager

GMI Cloud • Mountain View (CA)

On-site
USD 150,000 - 230,000
Technical Program Manager - Data Center / HPC Infrastructure and Operations
Technical Program Manager - Data Center / HPC Infrastructure and Operations

CIeNET International • Seattle (WA)

On-site
USD 140,000 - 190,000
Medical, Dental, Vision, Life Ins.
401(k) Matching
PTO & Holidays
+2
Principal Infrastructure Engineer, AI Cluster Performance & Validation
Principal Infrastructure Engineer, AI Cluster Performance & Validation

Nscale • New York (NY), San Francisco (CA), Seattle (WA)

On-site
USD 180,000 - 240,000
Technical Program Manager - Data Center / HPC Infrastructure and Operations
Technical Program Manager - Data Center / HPC Infrastructure and Operations

Cienet-International • Seattle (WA)

On-site
USD 140,000 - 180,000
Medical insurance
Dental insurance
Vision insurance
+3