Technical Compute Qualification Manager

Together AI

New York (NY)

On-site

USD 120,000 - 180,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health insurance
Dental and vision insurance
401(k) plan
Pre-tax commuter benefits
Paid time off

Job summary

Together AI is seeking a Technical Program Manager to lead the qualification of compute providers, driving structured evaluations across compute, networking, storage, power, cooling and operations.

You will coordinate engineering partners, review provider specs and test results, and deliver go/no-go recommendations to ensure new capacity meets standards. Strong analytical and cross-functional collaboration skills are essential.

Qualifications

  • 5+ years in technical program management, preferably in compute/data center environments.
  • Experience reading and evaluating hardware specifications and performance claims.
  • Ability to write scripts or queries (Python/SQL) to analyze provider data.

Responsibilities

  • Run end-to-end qualification processes for new compute capacity.
  • Coordinate engineering partners to validate specs and test results.
  • Produce clear go/no-go recommendations for capacity deployments.
  • Maintain auditable records of evaluation outcomes.
  • Perform parallel provider evaluations with timelines and stakeholder alignment.
  • Translate validation findings into actionable leadership decisions.

Skills

Technical program management
Stakeholder management
Data analysis
Reading and questioning specs

Tools

Python
SQL

Job description


  • As Technical Program Manager, Compute Qualification, you will run the process that screens and qualifies prospective compute providers, taking each prospective deployment through a structured evaluation across compute, networking, storage, power, cooling, and operations

  • You will coordinate various engineering partners through validation, review provider specifications and test results, and produce clear go/no-go recommendations on whether new capacity meets our standards

  • It is a high-impact, process-driven role for someone technical enough to know when a spec sheet does not add up, and additional diligence needs to be completed, and organized enough to drive many evaluations to closure in parallel

  • You will deep-dive into critical hardware performance metrics, proactively identifying potential bottlenecks in cluster architecture before they impact our end customers training or inference workloads.

  • Conduct diligence and work with engineering teams to make assessments regarding technical and operational resilience

  • Own and continuously improve the end-to-end qualification process for new compute capacity, from initial provider intake through final go/no-go recommendation

  • Run multiple provider evaluations in parallel, setting timelines, tracking status, and keeping every stakeholder aligned on what is needed and by when

  • Partner with infrastructure engineering, network engineering, data center engineering, and SRE teams to plan and coordinate technical validation, then translate their findings into clear decisions for leadership

  • Review provider technical specifications and questionnaire responses for completeness and accuracy, flagging gaps, inconsistencies, and risks that warrant follow-up

  • Conduct first-pass analysis of provider data yourself: compare specifications across suppliers , sanity-check performance claims, and surface issues before deeper engineering review

  • Maintain the standards, templates, and documentation that define what meets spec across compute, networking, storage, power, cooling, and operational support

  • Build a structured, auditable record of evaluation outcomes that informs sourcing decisions and scales the qualification function as the team grows


Benefits


  • Competitive health insurance plans

  • Dental and vision insurance

  • Pre-tax flexible spending accounts

  • Mental health support and services

  • Income protection & retirement

  • AD&D insurance

  • 401(k) plan

  • STD & LTD insurance

  • Life insurance

  • Monthly commuting stipend + pre-tax bene

  • Flexible time off policy

  • Team-driven celebrations and events

  • Monthly team lunches



  • Willingness to travel to provider and data center sites as needed

  • Proven ability to run multiple complex, cross-functional workstreams to deadline, with strong organization and stakeholder management

  • Working technical fluency across data center infrastructure: server and GPU hardware, high-performance networking (InfiniBand or Ethernet fabrics), storage, and power and cooling fundamentals; enough depth to read a detailed technical specification and know what to question

  • Excellent written and verbal communication; able to turn dense technical detail into clear recommendations for both engineers and executives

  • Hands-on comfort with data: able to write scripts or queries (for example, Python or SQL) to compare, validate, and analyze provider specifications and test results independently

  • 5+ years in technical program or project management, infrastructure program management, or a comparable technical operations role, ideally involving hardware, data center, or large-scale compute environments

  • Experience qualifying, commissioning, or accepting GPU clusters or HPC infrastructure against defined performance and reliability standards

  • Familiarity with AI training and inference infrastructure, including interconnect topologies, cluster bring-up, and acceptance testing

  • Experience in AI/HPC cluster design

  • Background working directly with hardware vendors, colocation providers, or cloud capacity providers

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda Innovation • California (MO)

Hybrid
USD 180,000 - 240,000
Compute Qualification Lead for HPC & AI Infra
Compute Qualification Lead for HPC & AI Infra

Together AI • New York (NY)

On-site
USD 120,000 - 180,000
Health insurance
Dental and vision insurance
401(k) plan
+2
Technical Program Leader - AI Infrastructure
Technical Program Leader - AI Infrastructure

Designworks Talent • Bellevue (KY)

Hybrid
USD 180,000 - 240,000
Technical Program Manager - Data Center / HPC Infrastructure and Operations
Technical Program Manager - Data Center / HPC Infrastructure and Operations

CIeNET Technologies • Seattle (WA)

On-site
USD 140,000 - 180,000
Medical Insurance
Dental Insurance
Vision Insurance
+5
Infrastructure Lab Manager
Infrastructure Lab Manager

Blue Signal Search • Austin (TX)

On-site
USD 140,000 - 190,000
HPC Infrastructure Engineer
HPC Infrastructure Engineer

Arcadia • San Francisco (CA)

On-site
USD 180,000 - 260,000
Data Center Compute Engineer
Data Center Compute Engineer

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive compensation
Equity opportunity
Comprehensive benefits
+2
Remote Compute Qualification Technical Program Manager
Remote Compute Qualification Technical Program Manager

Together AI • San Francisco (CA)

On-site
USD 200,000 - 250,000
Principal Infrastructure Engineer, AI Cluster Performance & Validation
Principal Infrastructure Engineer, AI Cluster Performance & Validation

Nscale • New York (NY), San Francisco (CA), Seattle (WA)

On-site
USD 180,000 - 240,000
Technical Program Manager - Data Center / HPC Infrastructure and Operations
Technical Program Manager - Data Center / HPC Infrastructure and Operations

Cienet-International • Seattle (WA)

On-site
USD 140,000 - 180,000
Medical insurance
Dental insurance
Vision insurance
+3