Senior ML Performance Engineer: LLM Benchmarking & GPU

Amadeus Search

San Francisco (CA)

Hybrid

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary
Equity and bonus opportunities
Medical, dental, and vision coverage
Retirement savings plan
Additional wellness benefits

Job summary

A leading AI infrastructure company is seeking a Senior ML Performance Engineer to design a comprehensive performance testing platform for large language models. This role requires a minimum of 7 years in performance engineering and strong experience with GPU programming and ML inference workloads. Candidates should have expertise in Python and C/C++. The position offers competitive compensation, equity, and wellness benefits in a hybrid work environment.

Qualifications

  • 7+ years in performance engineering or benchmarking roles.
  • Strong knowledge of ML inference workloads.
  • Experience building performance testing infrastructure from scratch.

Responsibilities

  • Design and implement a performance testing platform for LLM inference workloads.
  • Define benchmarking methodologies and metrics.
  • Collaborate with engineers to integrate performance testing.

Skills

Performance engineering
Benchmarking
ML inference
Python
C/C++
GPU optimization
Analytical skills

Tools

CUDA
ROCm
PyTorch
TensorFlow
ONNX Runtime

Job description

A leading AI infrastructure company is seeking a Senior ML Performance Engineer to design a comprehensive performance testing platform for large language models. This role requires a minimum of 7 years in performance engineering and strong experience with GPU programming and ML inference workloads. Candidates should have expertise in Python and C/C++. The position offers competitive compensation, equity, and wellness benefits in a hybrid work environment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Performance Engineer — Scalable DL Pipelines & Optimization
ML Performance Engineer — Scalable DL Pipelines & Optimization

Optiver US LLC • New York (NY)

On-site
USD 160,000 - 260,000
Competitive compensation package
Global profit-sharing pool
401(k) match up to 50%
+2
Staff ML Performance Engineer — Scalable Inference & CUDA
Staff ML Performance Engineer — Scalable Inference & CUDA

Modal • New York (NY)

On-site
USD 120,000 - 160,000
Software Engineer - ML Model Performance
Software Engineer - ML Model Performance

Baseten • San Francisco (CA)

On-site
USD 150,000 - 250,000
Staff ML Performance Engineer: Scale Training Throughput
Staff ML Performance Engineer: Scale Training Throughput

Wayve • Sunnyvale (CA)

On-site
USD 130,000 - 160,000
ML Model Performance Engineer - Inference and Acceleration
ML Model Performance Engineer - Inference and Acceleration

Baseten • New York (NY)

On-site
USD 200,000 - 275,000
ML Performance Engineer – Real-Time Inference
ML Performance Engineer – Real-Time Inference

Odyssey • Palo Alto (CA)

On-site
USD 130,000 - 160,000
Senior ML Systems Engineer - Model Inference & Efficiency
Senior ML Systems Engineer - Model Inference & Efficiency

Cohere • New York (NY)

Hybrid
USD 100,000 - 150,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+4
Senior ML Infrastructure Engineer - GPU & Scale
Senior ML Infrastructure Engineer - GPU & Scale

TensorWave • Las Vegas (NV)

On-site
USD 120,000 - 150,000
Competitive Salary
Stock Options
100% paid Medical, Dental, and Vision insurance
+8
Senior ML Performance Engineer
Senior ML Performance Engineer

well-funded deeptech startup • California (MO)

On-site
USD 200,000 - 250,000
Performance Engineer, Large-Scale ML Systems
Performance Engineer, Large-Scale ML Systems

Anthropic • New York (NY)

Hybrid
USD 280,000 - 850,000
Competitive compensation
Optional equity donation matching
Generous vacation and parental leave
+1