Senior AI Cluster Engineer – GPU Performance & Pre-Sales

Iframe

San Francisco (CA)

On-site

USD 220,000 - 320,000

Full time

14 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
Competitive base salary band

Job summary

iframe.ai is seeking a Senior Customer Engineer in the US to own three to five reserved-capacity accounts, guiding training and inference performance end-to-end. You will collaborate with a sales engineer and an account manager to manage technical relationships from kickoff through renewal.

Five-plus years in distributed- or ML-systems engineering with 64+ GPU jobs, strong PyTorch and CUDA know-how, and the ability to discuss technical scope with executives are required.

Qualifications

  • Five-plus years of distributed-systems or ML-systems engineering — production experience with 64+ GPU jobs is required.
  • Strong PyTorch / FSDP / Megatron-LM debugging skills.
  • Working knowledge of CUDA and at least one of Triton or CUTLASS.
  • Strong communicator. You'll write four post-mortems a quarter and present at customer all-hands.
  • Comfort with executive-level conversations on technical scope, timelines, and trade-offs.

Responsibilities

  • Own three to five reserved-capacity accounts as their named customer engineer. Most are AI labs or AI-native scale-ups running 256–2048-GPU jobs.
  • Profile distributed training jobs: NCCL collectives, gradient overlap, checkpoint cost, MFU. Make and defend specific recommendations.
  • Tune kernels and configurations alongside the runtime team — your changes go through the same review process and ship in the same release train.
  • Run the technical pre-sales for your accounts' expansions and renewals: pilot design, scope, success criteria, post-mortem.
  • Lead the post-mortem on any P1 affecting your accounts; escalate root-cause work into the runtime, cluster, or platform teams.
  • Carry the customer-engineering on-call rotation alongside runtime and cluster SRE — about one week per six.

Skills

Distributed systems
ML systems engineering
PyTorch
FSDP
Megatron-LM

Tools

CUDA
Triton
CUTLASS
NCCL
wandb

Job description

iframe.ai is seeking a Senior Customer Engineer in the US to own three to five reserved-capacity accounts, guiding training and inference performance end-to-end. You will collaborate with a sales engineer and an account manager to manage technical relationships from kickoff through renewal.

Five-plus years in distributed- or ML-systems engineering with 64+ GPU jobs, strong PyTorch and CUDA know-how, and the ability to discuss technical scope with executives are required.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Performance & Efficiency Engineer - Equity Eligible
Senior AI Performance & Efficiency Engineer - Equity Eligible

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Competitive benefits
AI GPU Cluster Support Engineer
AI GPU Cluster Support Engineer

Together Computer Inc • United States

Hybrid
USD 90,000 - 150,000
Startup equity
Health insurance
Remote work flexibility
Senior AI Infrastructure Lead: GPU Clusters & LLMs
Senior AI Infrastructure Lead: GPU Clusters & LLMs

Cadence Design Systems • San Jose (CA)

On-site
USD 137,000 - 254,000
Senior GPU Cluster Architect for AI Infra at Scale
Senior GPU Cluster Architect for AI Infra at Scale

Partner Company • United States

Remote
USD 184,000 - 318,000
Medical insurance
Dental insurance
Vision insurance
+1
Senior AI Infrastructure Architect — Enterprise GPU Clusters
Senior AI Infrastructure Architect — Enterprise GPU Clusters

NVIDIA • California (MO)

On-site
USD 184,000 - 287,500
Equity
Benefits
Senior AI Infrastructure Architect – GPU Clusters
Senior AI Infrastructure Architect – GPU Clusters

NVIDIA • California (MO)

On-site
USD 184,000 - 356,500
Equity
Benefits
Senior GPU Compute Solutions Architect
Senior GPU Compute Solutions Architect

Computacenter AG & Co. oHG • Northern (KY)

Hybrid
USD 190,000 - 230,000
Senior GPU Compute Architect for AI Data Center
Senior GPU Compute Architect for AI Data Center

Computacenter AG & Co. oHG • United States

Remote
USD 170,000 - 200,000
Lead AI Infra Architect for Enterprise GPU Clusters
Lead AI Infra Architect for Enterprise GPU Clusters

NVIDIA • Town of Texas (WI)

On-site
USD 184,000 - 287,500
Remote GPU Cluster Architect - AI Infrastructure Leader
Remote GPU Cluster Architect - AI Infrastructure Leader

Jobgether SRL • United States

Remote
USD 184,000 - 318,000
Medical, dental, vision insurance
Remote work reimbursement
RSUs may be available
+3