GPU Performance Engineer: CUDA, Triton & Inference

Brahma Consulting Group

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

35 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Brahma is conducting this search for a client and is building a team focused on accelerating generative video and image workloads. This early engineering hire reports to the CEO and joins a founding group with deep expertise in distributed systems and ML research.

You’ll optimize GPU performance, implement CUDA/Triton kernels, and build scalable inference and training engines across multiple GPUs. Ideal candidates bring CUDA, Triton, and GPU profiling experience and a passion for deep technical

Qualifications

  • 1+ years working on deep learning inference or training systems, or distributed systems.
  • Hands-on experience with CUDA, Triton, PyTorch internals, or GPU profiling.
  • Strong CS fundamentals and a drive to go deep on hard technical problems.

Responsibilities

  • Optimize GPU performance for training and inference on image and video workloads.
  • Profile and remove bottlenecks at kernel, memory, system, and cluster level.
  • Write CUDA and Triton kernels that ship to production.
  • Build distributed inference and training engines for diffusion models across GPUs and nodes.
  • Own communication performance: NCCL, RDMA over InfiniBand or RoCE, and disaggregated serving.
  • Build benchmarking and regression harnesses so performance gains hold in production.

Skills

CUDA
Triton
PyTorch internals
GPU profiling
Distributed systems

Tools

Nsight
RDMA
NCCL

Job description

Brahma is conducting this search for a client and is building a team focused on accelerating generative video and image workloads. This early engineering hire reports to the CEO and joins a founding group with deep expertise in distributed systems and ML research.

You’ll optimize GPU performance, implement CUDA/Triton kernels, and build scalable inference and training engines across multiple GPUs. Ideal candidates bring CUDA, Triton, and GPU profiling experience and a passion for deep technical

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Machine Learning Engineer, GPU Performance
Machine Learning Engineer, GPU Performance

Brahma Consulting Group • San Francisco (CA)

On-site
USD 150,000 - 210,000
GPU Transformer Performance Engineer (Triton/CUDA)
GPU Transformer Performance Engineer (Triton/CUDA)

Luma AI • United States

Remote
USD 180,000 - 280,000
GPU Systems Research Intern (CUDA/Triton)
GPU Systems Research Intern (CUDA/Triton)

Togetherai • San Francisco (CA)

On-site
USD 80,000 - 96,000
Housing stipend
Competitive compensation
Research Engineer, GPU Performance
Research Engineer, GPU Performance

Harnham • California (MO)

On-site
USD 120,000 - 160,000
Remote GPU Kernel Engineer: CUDA/Triton Optimization
Remote GPU Kernel Engineer: CUDA/Triton Optimization

anyone-ai • United States

On-site
USD 138,000 - 248,000
Inference Runtime Performance Engineer — GPU Kernels
Inference Runtime Performance Engineer — GPU Kernels

iFrame Corporation • San Francisco (CA)

Remote
USD 220,000 - 360,000
Senior GPU AI Inference Systems Engineer
Senior GPU AI Inference Systems Engineer

NVIDIA • California (MO)

On-site
USD 196,000 - 288,000
Equity
Comprehensive benefits
Senior System Software Engineer, Dynamo-Triton Inference
Senior System Software Engineer, Dynamo-Triton Inference

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 224,000 - 357,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
Senior AI Performance Engineer - GPU & Triton Optimizations
Senior AI Performance Engineer - GPU & Triton Optimizations

lumalabs-ai • San Francisco (CA)

On-site
USD 237,000 - 395,000