- The AI/ML TPM team owns delivery and execution across CoreWeave’s AI/ML Platform Services organization. The team partners closely with Product, Engineering, Research, Infrastructure, and Go-to-Market teams to deliver scalable, reliable, and high-performance platforms that support the full AI lifecycl
- AI/ML TPMs drive alignment and execution across highly technical, cross-functional teams to ensure the successful delivery of customer-facing infrastructure and platform capabilities used by researchers, engineers, and enterprise customers
- As a Technical Program Manager, you will lead complex, cross-functional programs across Performance & Benchmarking within our AI/ML Platform Services organization
- This team is responsible for ensuring CoreWeave’s infrastructure is performant, stable, and validated for demanding AI workloads before and as it reaches customers
- The work spans infrastructure verification, benchmarking, observability, and performance readiness across new hardware platforms, clusters, and model workloads
- You will partner with engineering, infrastructure, product, capacity, and go-to-market teams to drive programs that improve workload performance, validate new environments, operationalize benchmarking frameworks, and create visibility into how CoreWeave systems perform across models, hardware generations, and deployment contexts
- Drive end-to-end program execution for performance and benchmarking initiatives spanning infrastructure validation, performance testing, benchmark execution, observability, and launch readiness
- Partner with engineering and infrastructure teams to deliver programs that verify new hardware platforms, clusters, and software environments meet CoreWeave standards for performance and stability
- Lead cross‑functional efforts to operationalize benchmarking frameworks that measure model performance, runtime efficiency, GPU utilization, and workload reliability across environments
- Coordinate dependencies across platform engineering, infrastructure, capacity, product, and go‑to‑market teams to ensure performance findings are translated into roadmap priorities, customer readiness, and external proof points
- Build program mechanisms for release readiness, benchmark planning, risk management, issue escalation, and post‑launch review for performance‑sensitive infrastructure initiatives
- Establish dashboards, operating cadences, and success metrics to improve performance visibility, infrastructure validation coverage, benchmark repeatability, and time‑to‑readiness for new platforms
- Help drive prioritization across performance bottlenecks, test gaps, and benchmark requests by aligning stakeholders on goals, tradeoffs, and measurable outcomes
- Create clarity across ambiguous technical programs by aligning teams around performance goals, validation criteria, and execution milestones
Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience5+ years of technical program management experience in cloud infrastructure, distributed systems, high-performance computing, or AI/ML platformsUnderstanding of benchmarking methodologies, reproducibility, test coverage, and the tradeoffs between performance, stability, utilization, and customer readinessStrong technical fluency in distributed systems, GPU or accelerator-based infrastructure, workload performance measurement, and large-scale infrastructure operationsExperience with AI/ML benchmarking, performance analysis, or infrastructure validation for training and inference workloadsFamiliarity with GPU cluster architecture, workload observability, hardware bring‑up, and performance bottleneck analysisFamiliarity with translating technical performanceExcellent communication skills, with experience influencing engineering, product, and infrastructure stakeholdersExperience leading large‑scale cross‑functional programs involving performance engineering, benchmarking, validation systems, or infrastructure readinessExperience building launch processes, release governance, dependency management, and operational review mechanisms in fast‑scaling environmentsDemonstrated ability to define program metrics and drive measurable outcomes in performance, reliability, scale, or operational maturityWondering if you’re a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams – even if you aren’t a 100% skill or experience match