MTS: GPU Cluster Benchmarking & Performance

S27a

San Francisco, New York (CA, NY)

Hybrid

USD 140,000 - 210,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Generous PTO
Office stipend
Competitive healthcare (medical,Dental
Conference support

Job summary

SemiAnalysis is seeking a Senior MTS to build the next generation of ClusterMAX benchmarks and run evaluations across GPU clusters, focusing on storage IO, bandwidth, and fault-tolerance.

You will collaborate with executives and engineers across neoclouds, hyperscalers, and chip vendors to extend our TCO and goodput methodologies, and to author research that shapes AI infrastructure decisions.

Qualifications

  • Hands-on experience operating GPU clusters: Slurm and/or Kubernetes, InfiniBand/RoCE fabrics, distributed storage.
  • Strong Python and shell scripting; comfort building benchmark harnesses and CI pipelines.
  • Understanding of distributed training/inference failure modes — node failures, checkpointing, blast radius, recovery.
  • Experience similar to our Technical Consultant profile is a plus: due diligence, TCO analysis, and client-facing technical communication.
  • Security mindset: multi-tenant isolation, bare-metal vs virtualized trade-offs.

Responsibilities

  • Leading development of the next generation of ClusterMAX benchmarks: storage IO and bandwidth, NCCL/RCCL collectives, fault-tolerance and goodput measurement, multi-tenant security and isolation testing.
  • Deploying to and evaluating dozens of GPU clusters across hyperscalers and neoclouds (GB300/GB200 NVL72, B300/B200, H200, MI355X, TPUv7)
  • Extending our TCO and goodput methodology into automated, reproducible tests
  • Working with executives and engineers at 200+ neoclouds, hyperscalers, and chip vendors
  • Authoring technical research analyzing benchmark results, reliability, security, and ease of use, with direct authorship recognition

Skills

GPU cluster operations
Slurm
Kubernetes
InfiniBand/RoCE fabrics
Distributed storage
Python
Shell scripting
Benchmark harnesses & CI pipelines
Understanding distributed training/inf
Client-facing technical communication

Tools

CI pipelines tooling
Benchmarking tools

Job description

SemiAnalysis is seeking a Senior MTS to build the next generation of ClusterMAX benchmarks and run evaluations across GPU clusters, focusing on storage IO, bandwidth, and fault-tolerance.

You will collaborate with executives and engineers across neoclouds, hyperscalers, and chip vendors to extend our TCO and goodput methodologies, and to author research that shapes AI infrastructure decisions.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ClusterMAX Benchmarking Consultant
ClusterMAX Benchmarking Consultant

SemiAnalysis • New York (NY)

On-site
USD 120,000 - 190,000
Member of Technical Staff - ClusterMAX
Member of Technical Staff - ClusterMAX

S27a • San Francisco (CA), New York (NY)

Hybrid
USD 140,000 - 210,000
Generous PTO
Office stipend
Competitive healthcare (medical,Dental
+1
Technical Consultant
Technical Consultant

SemiAnalysis • New York (NY)

On-site
USD 120,000 - 190,000
GPU Systems Engineer - Low-Latency HPC & AI Clusters
GPU Systems Engineer - Low-Latency HPC & AI Clusters

Tower Research Capital • New York (NY)

On-site
USD 200,000 - 300,000
Generous paid time off policies
Hybrid working opportunities
Free breakfast, lunch & snacks
+4
Senior GPU Systems Performance Architect
Senior GPU Systems Performance Architect

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
MTS Inference: GPU Kernel & Performance Architect
MTS Inference: GPU Kernel & Performance Architect

Sail Research • San Francisco (CA)

On-site
USD 120,000 - 160,000
Meals provided
Studio Display for every employee
Senior GPU Benchmarking & Optimization Engineer
Senior GPU Benchmarking & Optimization Engineer

Webhosting • United States

On-site
USD 140,000 - 150,000
Health insurance
401(k) plan with matching
Professional development reimbursement
+4
GPU ML Benchmarking Engineer for Next-Gen AI Infra
GPU ML Benchmarking Engineer for Next-Gen AI Infra

Nebius • Amsterdam (VA)

On-site
USD 130,000 - 190,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+2
Senior AI GPU Cluster Architect
Senior AI GPU Cluster Architect

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Product Manager - GPU Cloud & AI Infrastructure
Senior Product Manager - GPU Cloud & AI Infrastructure

MaxIT Consulting - Max Corporate Group • Massachusetts

On-site
USD 140,000 - 210,000
Healthcare benefits
401(k) savings plan
Annual bonus or incentive
+7