Member of Technical Staff – ClusterMAX

S27a

San Francisco, Northern (CA, KY)

Hybrid

USD 120,000 - 180,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Generous PTO
Conference support
Health insurance
Office stipend

Job summary

SemiAnalysis is seeking a hands-on MTS on ClusterMAX to develop and run GPU cluster benchmarks, evaluate storage IO, NCCL/RCCL collectives, and security tests across diverse hyperscaler environments.

You will work with executives and engineers to analyze benchmark results, improve test automation, and publish technical research with attribution. This role blends data, systems, and economic analysis in a fast‑moving AI and semiconductor context.

Qualifications

  • Hands-on experience operating GPU clusters with Slurm or Kubernetes.
  • Strong Python and shell scripting abilities.
  • Understanding distributed training/inference failure modes — node failures, checkpointing, blast radius, recovery.
  • Security mindset: multi-tenant isolation, bare‑metal vs virtualized trade-offs.

Responsibilities

  • Lead development of the next generation of ClusterMAX benchmarks: storage IO and bandwidth, NCCL/RCCL collectives, fault‑tolerance and goodput measurement, multi‑tenant security and isolation testing.
  • Deploying to and evaluating dozens of GPU clusters across hyperscalers and neoclouds.
  • Extending the TCO and goodput methodology into automated, reproducible tests.
  • Collaborating with executives and engineers at 200+ neoclouds, hyperscalers, and chip vendors.
  • Authoring technical research analyzing benchmark results, reliability, security, and ease of use.

Skills

Slurm
Kubernetes
InfiniBand RoCE
Distributed storage
Python
Shell scripting
CI pipelines

Job description

Employment Type: Full-Time
Work Setting: In-office/Remote
Work Location: United States, New York, Mexico, San Francisco, Canada
Work Hours: Office hours
Find out more here: https://semianalysis.com
About SemiAnalysis

SemiAnalysis is an independent research and analysis firm specializing in the Semiconductor and AI industries. Our in-depth coverage spans the entire supply chain, from semiconductor fabrication processes to state-of-the-art AI Models, CUDA kernels, and GPU cloud infrastructure. We are recognized as the leading authority on AI infrastructure, with the highest concentration of industry experts within one team, and a deep-rooted passion for delving into the intricacies.

We’re a global team of over 20 analysts & engineers, each with extensive networks across the semiconductor supply chain and AI ecosystem, publishing industry‑shaping articles while participating in 40+ conferences annually.

Our newsletter reaches more than 200 000 subscribers worldwide, including senior management and C‑suite leaders at the leading semiconductor and AI companies.

Some of industry‑shaping articles are:

  • InferenceMAX: The world first open inference benchmark that continuous benchmarks performance of popular frontier models

  • MI300X vs H100 vs H200 Training: Extensive Benchmarking & Deep Dive into CUDA & ROCm training developer experience & performance

  • Trainium2 Architecture & Networking: Deep Dive into Amazon’s new Trainium2 rack scale system & 3D torus scale up topology

Find out more here: https://semianalysis.com

The Role

ClusterMAX™ is the industry‑standard GPU cloud rating system — 95% market coverage by volume, 84 providers rated, 209 tracked, and 140+ customer interviews behind each release. We recently published our cluster TCO and goodput framework in How Much Do GPU Clusters Really Cost?, showing that goodput expense alone swings 6–21% of total cluster TCO depending on fault‑tolerance approach. We are now testing providers for ClusterMAX 3.0 with expanded benchmarks, security requirements, and analysis.
As an MTS on ClusterMAX, you will build the benchmarks and run the evaluations that determine how the world’s GPU clouds get rated.

What you'll work on
  • Leading development of the next generation of ClusterMAX™ benchmarks: storage IO and bandwidth, NCCL/RCCL collectives, fault‑tolerance and goodput measurement, multi‑tenant security and isolation testing

  • Deploying to and evaluating dozens of GPU clusters across hyperscalers and neoclouds (GB300/GB200 NVL72, B300/B200, H200, MI355X, TPUv7)

  • Extending our TCO and goodput methodology (see the ClusterMAX TCO & Goodput calculator) into automated, reproducible tests

  • Working with executives and engineers at 200+ neoclouds, hyperscalers, and chip vendors

  • Authoring technical research analyzing benchmark results, reliability, security, and ease of use, with direct authorship recognition

What we're looking for
  • Hands‑on experience operating GPU clusters: Slurm and/or Kubernetes, InfiniBand/RoCE fabrics, distributed storage

  • Strong Python and shell scripting; comfort building benchmark harnesses and CI pipelines

  • Understanding of distributed training/inference failure modes — node failures, checkpointing, blast radius, recovery

  • Experience similar to our Technical Consultant profile is a plus: due diligence, TCO analysis, and client‑facing technical communication

  • Security mindset: multi‑tenant isolation, bare‑metal vs virtualized trade‑offs

Why SemiAnalysis?

At SemiAnalysis, you’re placed directly inside the most important conversations shaping the future of AI and semiconductors. You’ll develop a first‑principles understanding of the technical and economic dynamics behind the world’s most consequential technology, with rare end‑to‑end visibility across the entire stack—from silicon and systems to models, software, and real‑world deployment. Our work is read and relied upon by the people allocating capital, building infrastructure, and setting strategy across the industry, giving your research and analysis real, measurable impact.

The environment is built for learning and autonomy. You’ll have the freedom to chase interesting ideas, dig into emerging trends early, and continuously expand your technical and economic toolkit, supported by generous PTO, office stipends, competitive healthcare (medical, dental, vision), and support for conferences and ongoing learning.
SemiAnalysis LLC participates in the E-Verify program to confirm the employment eligibility of all newly hired employees. For more information, visit e-verify.gov.

  • E-Verify Poster: https://www.everify.gov/sites/default/files/everify/posters/EVerifyParticipationPoster.pdf

  • Right to Work Poster: https://www.justice.gov/crt/case-document/file/1133936/dl?inline=

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - ClusterMAX
Member of Technical Staff - ClusterMAX

S27a • San Francisco (CA), New York (NY)

Hybrid
USD 140,000 - 210,000
Generous PTO
Office stipend
Competitive healthcare (medical,Dental
+1
Member of Technical Staff - Tokenomics
Member of Technical Staff - Tokenomics

S27a • San Francisco (CA), New York (NY)

Hybrid
USD 140,000 - 190,000
Generous PTO
Office stipend
Competitive healthcare (medical,Dental
+2
Member of Technical Staff – Tokenomics
Member of Technical Staff – Tokenomics

S27a • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
MTS - Research Analyst, AI Infrastructure & Economics
MTS - Research Analyst, AI Infrastructure & Economics

Socket.dev • San Francisco (CA)

Hybrid
USD 60,000 - 90,000
Senior Research Analyst, AI Infrastructure & Economics
Senior Research Analyst, AI Infrastructure & Economics

S27a • San Francisco (CA), New York (NY)

On-site
USD 120,000 - 190,000
Member of Technical Staff - Inference
Member of Technical Staff - Inference

S27a • San Francisco (CA), New York (NY)

Hybrid
USD 150,000 - 210,000
Research Analyst, AI Infrastructure & Economics
Research Analyst, AI Infrastructure & Economics

SemiAnalysis • San Francisco (CA)

Hybrid
USD 65,000 - 90,000
Research Analyst, AI Infrastructure & Economics
Research Analyst, AI Infrastructure & Economics

S27a • San Francisco (CA), New York (NY)

On-site
USD 60,000 - 90,000
MTS - Senior Research Analyst, AI Infrastructure & Economics
MTS - Senior Research Analyst, AI Infrastructure & Economics

Socket.dev • San Francisco (CA)

Hybrid
USD 120,000 - 190,000
Member of Technical Staff - Frontier System Modelling
Member of Technical Staff - Frontier System Modelling

S27a • San Francisco (CA), New York (NY)

Hybrid
USD 150,000 - 210,000