Senior ML Performance Engineer

Amadeus Search

San Francisco (CA)

Hybrid

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary
Equity and bonus opportunities
Medical, dental, and vision coverage
Retirement savings plan
Additional wellness benefits

Job summary

A leading AI infrastructure company is seeking a Senior ML Performance Engineer to design a comprehensive performance testing platform for large language models. This role requires a minimum of 7 years in performance engineering and strong experience with GPU programming and ML inference workloads. Candidates should have expertise in Python and C/C++. The position offers competitive compensation, equity, and wellness benefits in a hybrid work environment.

Qualifications

  • 7+ years in performance engineering or benchmarking roles.
  • Strong knowledge of ML inference workloads.
  • Experience building performance testing infrastructure from scratch.

Responsibilities

  • Design and implement a performance testing platform for LLM inference workloads.
  • Define benchmarking methodologies and metrics.
  • Collaborate with engineers to integrate performance testing.

Skills

Performance engineering
Benchmarking
ML inference
Python
C/C++
GPU optimization
Analytical skills

Tools

CUDA
ROCm
PyTorch
TensorFlow
ONNX Runtime

Job description

Position: Senior ML Performance Engineer
Location: SF Bay Area (US) or Toronto (Canada) – Hybrid
Employment Type: Full-Time
Industry: AI Infrastructure / Compiler Systems
Overview

A venture-backed AI infrastructure company is building a high-performance, portable compiler designed to let developers “build once, deploy anywhere.” This includes cloud, edge, and hybrid environments — all optimized for resource efficiency, scalability, and sustainable AI development.

The team is looking for a Senior ML Performance Engineer to architect and lead a Performance Testing Platform from the ground up, measuring and optimizing the performance of large language models (LLMs) before and after compiler optimization on modern GPU architectures.

This role sits at the intersection of ML systems, GPU architecture, and performance engineering, with high visibility into product quality and customer impact.

Key Responsibilities
  • Design and implement a comprehensive performance testing platform for LLM inference workloads across GPU clusters

  • Define benchmarking methodologies, metrics, and test suites (latency, throughput, memory utilization, power consumption, and model accuracy)

  • Establish baseline performance for unoptimized models and validate post-optimization improvements

  • Build automated pipelines for continuous performance validation across compiler releases and model updates

  • Investigate performance bottlenecks using GPU profilers and system-level monitoring

  • Collaborate with compiler engineers, ML engineers, and DevOps to integrate performance testing into development workflows

  • Create dashboards and reporting to track performance trends, regressions, and wins

  • Document best practices for GPU-based ML performance testing

Required Qualifications
  • 7+ years in performance engineering, benchmarking, or systems engineering roles

  • Strong knowledge of ML inference workloads, particularly transformer-based LLMs

  • Hands-on GPU programming and optimization experience (CUDA, ROCm, or similar)

  • Strong programming skills in Python and C/C++

  • Proven experience building performance testing infrastructure or benchmarking platforms from scratch

  • Experience with ML frameworks: PyTorch, TensorFlow, ONNX Runtime, vLLM, TensorRT-LLM

  • Proficiency with profiling and debugging GPU workloads

  • Experience with CI/CD systems and test automation frameworks

  • Strong analytical skills with the ability to design experiments, analyze results, and communicate findings clearly

Nice to Have
  • AMD GPU experience (Mi200/Mi300) and ROCm ecosystem

  • Compiler optimization knowledge

  • Distributed inference and multi-GPU workloads

  • ML model quantization, pruning, and optimization techniques

  • High-performance computing or systems-level optimization

  • Infrastructure-as-code experience: Kubernetes, Docker, Terraform

  • Contributions to open-source ML or systems projects

Personal Attributes
  • Detail-oriented — able to spot subtle regressions

  • Self-driven and accountable

  • Collaborative and team-oriented

  • Passionate about sustainable AI

  • Clear and effective communicator

Compensation & Benefits
  • Competitive salary, dependent on experience and location

  • Equity and bonus opportunities

  • Medical, dental, and vision coverage

  • Retirement savings plan

  • Additional wellness benefits

Why This Role Is Unique
  • Build the infrastructure that validates high-performance ML models

  • Influence core product quality and customer outcomes

  • Work in a highly technical, high-impact environment at the forefront of AI systems

  • Collaborate across a globally distributed team

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer- Model Performance Systems
Software Engineer- Model Performance Systems

Baseten • San Francisco (CA)

On-site
USD 160,000 - 200,000
Competitive compensation
Equity
Medical/dental/vision insurance
+4
Senior ML Performance Engineer
Senior ML Performance Engineer

well-funded deeptech startup • California (MO)

On-site
USD 200,000 - 250,000
Member of Technical Staff - GPU Performance Engineer
Member of Technical Staff - GPU Performance Engineer

Liquid AI • San Francisco (CA)

On-site
USD 120,000 - 180,000
Competitive base salary with equity
100% medical, dental, and vision premiums
401(k) matching up to 4%
+2
AI Compiler and Performance Engineer
AI Compiler and Performance Engineer

ElastixAI INC. • Seattle (WA)

Hybrid
USD 180,000 - 240,000
Competitive compensation
Comprehensive medical, dental, and vision coverage
Flexible Time Off (FTO)
+4
Senior LLM Inference Engineer — Performance & GPU Optimization
Senior LLM Inference Engineer — Performance & GPU Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Software Engineer, Model Performance Systems
Software Engineer, Model Performance Systems

Baseten • New York (NY)

On-site
USD 160,000 - 200,000
Competitive compensation with equity
100% medical, dental, and vision insurance coverage
Generous PTO including Winter Break
+3
Performance Architect - AI Hardware
Performance Architect - AI Hardware

TEEMA • United States

Remote
USD 140,000 - 210,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000
AI Engineer - Model Performance
AI Engineer - Model Performance

Dormont Manufacturing Co • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive compensation and benefits
Supportive environment for innovation and growth
Frontier AI Workloads - Performance and Scalability Engineer
Frontier AI Workloads - Performance and Scalability Engineer

AMD • San Jose (CA)

On-site
USD 150,000 - 200,000