GPU Performance Engineer: Accelerate AI Video Pipelines

HeyGen

Los Angeles, Palo Alto, San Francisco (CA, CA, CA)

On-site

USD 130,000 - 170,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Competitive salary
Dynamic work environment
Growth opportunities
Collaborative culture
Access to latest tools

Job summary

HeyGen is seeking a Software Engineer focused on GPU performance to accelerate AI video applications and related workloads. You will work across model execution, inference infrastructure, and profiling to reduce latency and cost while increasing throughput.

Ideal candidates will have a strong background in GPU hardware, profiling, and Python-based ML frameworks, with experience in delivering production-grade performance improvements at scale.

Qualifications

  • Experience optimizing GPU-based AI workloads or high-performance computing systems.
  • Proficiency in Python and PyTorch or a similar ML framework.
  • Strong curiosity about GPU hardware, including memory bandwidth, cache behavior, tensor cores, and CPU-GPU data movement.
  • Experience using profiling tools to connect hardware behavior to application-level bottlenecks and validate improvements.
  • Ability to turn performance experiments into reliable production changes and communicate tradeoffs clearly.

Responsibilities

  • Use profiling tools to investigate GPU utilization, kernel execution, memory bandwidth, and CPU–GPU data movement.
  • Identify bottlenecks across model execution, preprocessing, and inference serving, then measure the impact of each optimization.
  • Improve performance through batching, scheduling, memory management, and better GPU utilization.
  • Develop or integrate high-performance GPU kernels when existing implementations limit performance.
  • Build benchmarks and automated checks that catch performance regressions across representative video workloads.
  • Collaborate with AI researchers and infrastructure engineers to bring optimizations into production.
  • Measure the effect of changes on latency, throughput, cost, and output quality.

Skills

GPU optimization
Python
PyTorch
Profiling
Latency reduction
Performance tuning

Education

B.S. in CS/EE

Tools

CUDA
Triton
C++

Job description

HeyGen is seeking a Software Engineer focused on GPU performance to accelerate AI video applications and related workloads. You will work across model execution, inference infrastructure, and profiling to reduce latency and cost while increasing throughput.

Ideal candidates will have a strong background in GPU hardware, profiling, and Python-based ML frameworks, with experience in delivering production-grade performance improvements at scale.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer, GPU Performance
Software Engineer, GPU Performance

HeyGen • Los Angeles (CA), Palo Alto (CA), San Francisco (CA)

On-site
USD 130,000 - 170,000
Competitive salary
Dynamic work environment
Growth opportunities
+2
GPU Performance Engineer: Scale ML Inference & Systems
GPU Performance Engineer: Scale ML Inference & Systems

Anthropic • New York (NY)

Hybrid
USD 280,000 - 850,000
Research Engineer, GPU Performance
Research Engineer, GPU Performance

Harnham • California (MO)

On-site
USD 120,000 - 160,000
Senior AI Training Performance Engineer (GPU & Scale)
Senior AI Training Performance Engineer (GPU & Scale)

figure.ai • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Performance Engineer: AI/GPU Benchmarking & Optimization
Performance Engineer: AI/GPU Benchmarking & Optimization

Thomas To • Santa Clara (CA)

On-site
USD 136,000 - 213,000
Equity
Benefits
AI DevTech Engineer: Accelerate Generative AI on GPUs
AI DevTech Engineer: Accelerate Generative AI on GPUs

NVIDIA • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Equity
Benefits
Backend Infrastructure Engineer - Scalable ML & Data Systems
Backend Infrastructure Engineer - Scalable ML & Data Systems

HeyGen • San Francisco (CA)

On-site
USD 180,000 - 240,000
401k plan
Health benefits
Generous PTO
+2
AI DevTech Engineer: Equity & GPU-Accelerated AI
AI DevTech Engineer: Equity & GPU-Accelerated AI

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Kernel Engineer: GPU Performance & Inference
Kernel Engineer: GPU Performance & Inference

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 240,000
Performance Engineer
Performance Engineer

Tangier Partners • Boston (MA)

On-site
USD 150,000 - 210,000