Software Engineer, GPU Performance

HeyGen

Los Angeles, Palo Alto, San Francisco (CA, CA, CA)

On-site

USD 130,000 - 170,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Competitive salary
Dynamic work environment
Growth opportunities
Collaborative culture
Access to latest tools

Job summary

HeyGen is seeking a Software Engineer focused on GPU performance to accelerate AI video applications and related workloads. You will work across model execution, inference infrastructure, and profiling to reduce latency and cost while increasing throughput.

Ideal candidates will have a strong background in GPU hardware, profiling, and Python-based ML frameworks, with experience in delivering production-grade performance improvements at scale.

Qualifications

  • Experience optimizing GPU-based AI workloads or high-performance computing systems.
  • Proficiency in Python and PyTorch or a similar ML framework.
  • Strong curiosity about GPU hardware, including memory bandwidth, cache behavior, tensor cores, and CPU-GPU data movement.
  • Experience using profiling tools to connect hardware behavior to application-level bottlenecks and validate improvements.
  • Ability to turn performance experiments into reliable production changes and communicate tradeoffs clearly.

Responsibilities

  • Use profiling tools to investigate GPU utilization, kernel execution, memory bandwidth, and CPU–GPU data movement.
  • Identify bottlenecks across model execution, preprocessing, and inference serving, then measure the impact of each optimization.
  • Improve performance through batching, scheduling, memory management, and better GPU utilization.
  • Develop or integrate high-performance GPU kernels when existing implementations limit performance.
  • Build benchmarks and automated checks that catch performance regressions across representative video workloads.
  • Collaborate with AI researchers and infrastructure engineers to bring optimizations into production.
  • Measure the effect of changes on latency, throughput, cost, and output quality.

Skills

GPU optimization
Python
PyTorch
Profiling
Latency reduction
Performance tuning

Education

B.S. in CS/EE

Tools

CUDA
Triton
C++

Job description

About HeyGen

At HeyGen, our mission is to make visual storytelling accessible to all. Over the last decade, visual content has become the preferred method of information creation, consumption, and retention. But the ability to create such content, in particular videos, continues to be costly and challenging to scale. Our ambition is to build technology that equips more people with the power to reach, captivate, and inspire audiences.
Learn more at www.heygen.com . Visit our Mission and Culture doc here .

Position Summary

HeyGen is building AI applications including Avatar IV, Photo Avatar, Interactive Avatar, and Video Translation. We’re looking for a Software Engineer focused on GPU performance to make the systems behind these experiences faster and more efficient.

You will work across model execution and inference infrastructure, using profiling and measurement to improve latency, throughput, and GPU cost. This role is a fit for an engineer who enjoys understanding how software uses the hardware beneath it.

Key Responsibilities
  • Use NVIDIA Nsight Systems, Nsight Compute, and PyTorch Profiler to investigate GPU utilization, kernel execution, memory bandwidth, and CPU–GPU data movement.
  • Identify bottlenecks across model execution, preprocessing, and inference serving, then measure the impact of each optimization.
  • Improve performance through batching, scheduling, memory management, and better GPU utilization.
  • Develop or integrate high-performance GPU kernels when existing implementations limit performance.
  • Build benchmarks and automated checks that catch performance regressions across representative video workloads.
  • Collaborate with AI researchers and infrastructure engineers to bring optimizations into production.
  • Measure the effect of changes on latency, throughput, cost, and output quality.
Qualifications
  • Experience optimizing GPU-based AI workloads or high-performance computing systems.
  • Proficiency in Python and experience with PyTorch or a similar machine learning framework.
  • Strong curiosity about GPU hardware, including memory bandwidth, cache behavior, tensor cores, and data movement between CPU and GPU.
  • Experience using profiling tools to connect hardware behavior to application-level bottlenecks and validate improvements.
  • Ability to turn performance experiments into reliable production changes and communicate tradeoffs clearly.
Preferred Qualifications
  • Experience with CUDA, Triton, or C++ GPU programming.
  • Experience optimizing video, image, audio, diffusion, or Transformer models.
  • Familiarity with multi-GPU inference, GPU interconnects, quantization, or large-scale model serving.
  • Experience building performance benchmarks or regression testing infrastructure.
  • Prior experience in a fast-paced technology environment.
What HeyGen Offers
  • Competitive salary and benefits package.
  • Dynamic and inclusive work environment.
  • Opportunities for professional growth and advancement.
  • Collaborative culture that values innovation and creativity.
  • Access to the latest technologies and tools.

HeyGen is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Join us at HeyGen

Help us make AI video faster and more accessible. We’d love to hear from you.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Performance Engineer: Accelerate AI Video Pipelines
GPU Performance Engineer: Accelerate AI Video Pipelines

HeyGen • Los Angeles (CA), Palo Alto (CA), San Francisco (CA)

On-site
USD 130,000 - 170,000
Competitive salary
Dynamic work environment
Growth opportunities
+2
Software Engineer - GPU Kernel
Software Engineer - GPU Kernel

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
Backend Engineer - Infrastructure
Backend Engineer - Infrastructure

HeyGen • San Francisco (CA)

On-site
USD 180,000 - 240,000
401k plan
Health benefits
Generous PTO
+2
Applied AI Engineer, Video Agent
Applied AI Engineer, Video Agent

HeyGen, Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000
Equity in a high-growth AI company
401K Match
Health, dental, and vision benefits
+1
Applied AI Engineer, Video Agent
Applied AI Engineer, Video Agent

Engg • San Francisco (CA), Palo Alto (CA), Los Angeles (CA)

On-site
USD 180,000 - 240,000
Equity
401K Match
Health benefits
+3
Applied AI Engineer, Video Agent
Applied AI Engineer, Video Agent

HeyGen • San Francisco (CA)

On-site
USD 180,000 - 240,000
Equity in a high-growth AI company
401K Match
Health, dental, and vision benefits
+2
GPU Performance Profiling Engineer
GPU Performance Profiling Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Research Engineer
Research Engineer

HeyGen • Los Angeles (CA)

On-site
USD 90,000 - 130,000
Competitive salary and benefits
Opportunities for advancement
Innovative and inclusive work environment
+1
Performance Engineer, GPU
Performance Engineer, GPU

Anthropic • New York (NY)

On-site
USD 280,000 - 850,000
Senior High Performance AI Engineer
Senior High Performance AI Engineer

NVIDIA • California (MO)

On-site
USD 184,000 - 287,500
Equity options
Generous benefits package