Senior AI Performance Engineer - GPUs & Distributed Systems

Fireworks AI

United States

On-site

USD 150,000 - 260,000

Full time

7 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Fireworks AI is seeking a Software Engineer focused on Performance Optimization to push speed and efficiency across our AI infrastructure. You will own optimization at every layer, from GPU kernels to large-scale distributed systems.

Collaborating with research, infrastructure, and systems teams, you will target bottlenecks in LLMs, VLMs, and video models, improve latency and memory use, and drive CUDA, Triton, and PyTorch optimizations for production workloads.

Qualifications

  • Bachelor's degree in CS/CE/EE or equivalent practical experience.
  • 5+ years of experience in performance optimization or HPC.
  • Experience optimizing large models for training and inference (LLMs, VLMs, or video models).
  • Proficiency in CUDA or ROCm and GPU profiling tools (Nsight, nvprof, CUPTI).
  • Familiarity with PyTorch and performance-critical model execution.
  • Experience with distributed systems debugging and multi-GPU optimization.
  • Contributions to open-source ML or HPC infrastructure.
  • Knowledge of compiler stacks or ML compilers (e.g., torch.compile, Triton, XLA).
  • Familiarity with cloud-scale AI infrastructure and orchestration tools (Kubernetes).
  • Background in ML systems engineering or hardware-aware model design.

Responsibilities

  • Analyze and optimize performance across GPU kernels and distributed systems.
  • Profile bottlenecks and implement low-level optimizations using CUDA, Triton, and related tools.
  • Collaborate with ML researchers to co-design hardware-efficient model architectures.
  • Scale inference and training across multi-GPU, multi-node environments.
  • Build and maintain performance benchmarking and monitoring infrastructure.

Skills

Performance optimization
GPU kernel optimization
CUDA
Triton
PyTorch
Distributed systems
Profiling
Load balancing
Hardware-aware optimization
High-performance computing

Education

Bachelor's degree in Computer Science/Engineering
Master's or PhD in Computer Science/Electrical Engineering

Tools

CUDA
ROCm
Nsight
nvprof
CUPTI
PyTorch
Triton
XLA
Kubernetes

Job description

Fireworks AI is seeking a Software Engineer focused on Performance Optimization to push speed and efficiency across our AI infrastructure. You will own optimization at every layer, from GPU kernels to large-scale distributed systems.

Collaborating with research, infrastructure, and systems teams, you will target bottlenecks in LLMs, VLMs, and video models, improve latency and memory use, and drive CUDA, Triton, and PyTorch optimizations for production workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Performance Optimization Engineer for AI Infrastructure
Performance Optimization Engineer for AI Infrastructure

Fireworks AI • San Mateo (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff, Performance Optimization
Member of Technical Staff, Performance Optimization

Fireworks AI • San Mateo (CA)

On-site
USD 180,000 - 260,000
AI Training Infrastructure Engineer - Scale & Performance
AI Training Infrastructure Engineer - Scale & Performance

SupportFinity™ • San Francisco (CA)

On-site
USD 175,000 - 220,000
Meaningful equity
Competitive salary
Comprehensive benefits package
Research Engineer, GPU Performance
Research Engineer, GPU Performance

Harnham • California (MO)

On-site
USD 120,000 - 160,000
Senior AI Training Performance Engineer (GPU & Scale)
Senior AI Training Performance Engineer (GPU & Scale)

figure.ai • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Member of Technical Staff (Performance Optimization)
Member of Technical Staff (Performance Optimization)

Fireworks AI • United States

On-site
USD 150,000 - 260,000
LLM Infrastructure Engineer — Performance & Scale
LLM Infrastructure Engineer — Performance & Scale

Fireworks AI • United States

On-site
USD 150,000 - 210,000
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior AI Systems Performance Engineer
Senior AI Systems Performance Engineer

NVIDIA • Austin (TX)

On-site
USD 272,000 - 432,000
Equity
Benefits
Senior Performance Engineer: AI Workload Optimization
Senior Performance Engineer: AI Workload Optimization

NVIDIA • Redmond (WA)

On-site
USD 224,000 - 432,000
Equity
Benefits