Staff ML Performance Engineer - GPU & Inference

Modal

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Modal is building an infrastructure layer for AI and is seeking strong engineers to optimize ML systems for performance at scale. You will contribute to Modal’s container runtime and open-source projects, pushing language and diffusion models toward higher throughput and lower latency.

The role emphasizes working with Torch, TensorRT, CUDA, and NVIDIA GPU architectures to maximize efficiency, while exploring low-level OS foundations to improve performance and reliability.

Qualifications

  • 5+ years of experience writing high-quality, high-performance code.
  • Experience with Torch, high-level ML frameworks, and inference engines (vLLM or TensorRT).
  • Familiarity with Nvidia GPU architecture and CUDA.
  • Experience with ML performance engineering and boosting GPU performance (e.g., debugging SM occupancy, rewriting algorithms, reducing host overhead).
  • Nice-to-have: familiarity with low-level OS foundations (Linux kernel, file systems, containers).

Responsibilities

  • Build ML systems that perform at scale and low latency.
  • Contribute to open-source projects and Modal’s container runtime.
  • Advance language and diffusion models toward higher throughput.

Skills

5+ years of experience writing high‑on

Tools

Torch
TensorRT
CUDA
Linux kernel
Containers

Job description

Modal is building an infrastructure layer for AI and is seeking strong engineers to optimize ML systems for performance at scale. You will contribute to Modal’s container runtime and open-source projects, pushing language and diffusion models toward higher throughput and lower latency.

The role emphasizes working with Torch, TensorRT, CUDA, and NVIDIA GPU architectures to maximize efficiency, while exploring low-level OS foundations to improve performance and reliability.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Performance Engineer - GPU & Inference
Senior ML Performance Engineer - GPU & Inference

Modal Labs • New York (NY)

On-site
Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Modal Labs • New York (NY)

On-site
USD 150,000 - 190,000
Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Modal • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Modal • New York (NY)

On-site
USD 120,000 - 160,000
Member of Technical Staff, ML Systems
Member of Technical Staff, ML Systems

TensorScale AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff ML Performance Engineer — Scalable Inference & CUDA
Staff ML Performance Engineer — Scalable Inference & CUDA

Modal • New York (NY)

On-site
USD 120,000 - 160,000
ML Inference Engineer San Francisco · Engineering · Full Time →
ML Inference Engineer San Francisco · Engineering · Full Time →

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Visa sponsorship
Relocation support
Generous health, dental, and vision coverage
Principal ML Infrastructure Engineer (Relocation Available)
Principal ML Infrastructure Engineer (Relocation Available)

Franklin Fitch • Dallas (TX)

On-site
USD 100,000 - 140,000
GPU Performance Engineer: Scale ML Inference & Systems
GPU Performance Engineer: Scale ML Inference & Systems

Anthropic • New York (NY)

Hybrid
USD 280,000 - 850,000
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000