Machine Learning Engineer, GPU Performance

Brahma Consulting Group

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

28 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Brahma is conducting this search for a client and is building a team focused on accelerating generative video and image workloads. This early engineering hire reports to the CEO and joins a founding group with deep expertise in distributed systems and ML research.

You’ll optimize GPU performance, implement CUDA/Triton kernels, and build scalable inference and training engines across multiple GPUs. Ideal candidates bring CUDA, Triton, and GPU profiling experience and a passion for deep technical

Qualifications

  • 1+ years working on deep learning inference or training systems, or distributed systems.
  • Hands-on experience with CUDA, Triton, PyTorch internals, or GPU profiling.
  • Strong CS fundamentals and a drive to go deep on hard technical problems.

Responsibilities

  • Optimize GPU performance for training and inference on image and video workloads.
  • Profile and remove bottlenecks at kernel, memory, system, and cluster level.
  • Write CUDA and Triton kernels that ship to production.
  • Build distributed inference and training engines for diffusion models across GPUs and nodes.
  • Own communication performance: NCCL, RDMA over InfiniBand or RoCE, and disaggregated serving.
  • Build benchmarking and regression harnesses so performance gains hold in production.

Skills

CUDA
Triton
PyTorch internals
GPU profiling
Distributed systems

Tools

Nsight
RDMA
NCCL

Job description

Brahma is conducting this search on behalf of a client.

We're a small, well-funded founding team rebuilding the training and inference stack for generative video and image models. Today's stack was built for language models. We're co-designing across GPU kernels, distributed systems, and the models themselves to make these workloads dramatically faster and cheaper. Our inference engine is live with customers in generative media and robotics.

This is an early engineering hire. You'll report directly to the CEO and work alongside a founding team with deep expertise in distributed systems, kernel optimization, cloud infrastructure, and ML research.

What you'll do
  • Optimize GPU performance for training and inference on image and video generation workloads
  • Profile and remove bottlenecks at the kernel, memory, system, and cluster level using Nsight and related tools
  • Write CUDA and Triton kernels that ship to production
  • Build distributed inference and training engines for diffusion models across multiple GPUs and nodes
  • Own communication performance: NCCL, RDMA over InfiniBand or RoCE, and disaggregated serving
  • Build benchmarking and regression harnesses so performance gains hold in production
What we're looking for
  • 1+ years working on deep learning inference or training systems, or distributed systems
  • Hands-on experience with CUDA, Triton, PyTorch internals, or GPU profiling
  • Strong CS fundamentals and a drive to go deep on hard technical problems
  • You'd rather make a model ten times faster than train one
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Performance Engineer: CUDA, Triton & Inference
GPU Performance Engineer: CUDA, Triton & Inference

Brahma Consulting Group • San Francisco (CA)

On-site
USD 150,000 - 210,000
Research Engineer, GPU Performance
Research Engineer, GPU Performance

Harnham • California (MO)

On-site
USD 120,000 - 160,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • San Francisco (CA)

On-site
USD 140,000 - 210,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • New York (NY)

On-site
USD 220,000 - 485,000
GPU Optimization Engineer
GPU Optimization Engineer

Vincere • San Francisco (CA)

On-site
USD 230,000 - 300,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • Palo Alto (CA)

On-site
USD 220,000 - 485,000
GPU Performance Engineer: Scale ML Inference & Systems
GPU Performance Engineer: Scale ML Inference & Systems

Anthropic • New York (NY)

Hybrid
USD 280,000 - 850,000