AI Performance Engineer: CUDA & Multi-GPU Optimization

Brillfy Technology Inc

United States

On-site

USD 150,000 - 210,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Brillfy Technology Inc. is seeking a hands-on GPU/AI systems engineer focused on CUDA optimization, multi-GPU performance, and system-wide profiling.

The role requires deep expertise in GPU architecture and performance tuning, with experience across NVIDIA tools and distributed AI stacks. The ideal candidate will profile AI/ML workloads, identify bottlenecks, and implement latency and throughput enhancements across compute, memory, and networking layers, collaborating with AI, DevOps, and

Qualifications

  • Strong hands-on CUDA programming and GPU performance optimization.
  • Deep understanding of GPU architecture and memory hierarchy.
  • Experience with Nsight and CUDA profiling tools.

Responsibilities

  • Profile and optimize AI/ML workloads across multi-GPU and multi-node systems.
  • Identify bottlenecks across compute, memory, networking, and orchestration layers.
  • Optimize CUDA kernels (memory coalescing, shared memory usage, occupancy tuning).
  • Improve inference performance using TensorRT, Triton, DeepStream, NeMo.
  • Analyze and improve latency, throughput, GPU utilization, and memory efficiency.
  • Work on distributed AI systems using Apache Ray, NCCL, Kubernetes GPU scheduling.
  • Build benchmarking frameworks and performance monitoring systems.
  • Collaborate with AI, DevOps, and Infrastructure teams for system-wide optimization.

Skills

CUDA programming
GPU architecture
Performance profiling
NVIDIA ecosystem (TensorRT, Triton, Ne

Tools

Nsight tools

Job description

Brillfy Technology Inc. is seeking a hands-on GPU/AI systems engineer focused on CUDA optimization, multi-GPU performance, and system-wide profiling.

The role requires deep expertise in GPU architecture and performance tuning, with experience across NVIDIA tools and distributed AI stacks. The ideal candidate will profile AI/ML workloads, identify bottlenecks, and implement latency and throughput enhancements across compute, memory, and networking layers, collaborating with AI, DevOps, and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Performance Engineer
Senior AI Performance Engineer

Brillfy Technology Inc • United States

On-site
USD 150,000 - 210,000
AI Performance Software Engineer – GPU & DL Optimizations
AI Performance Software Engineer – GPU & DL Optimizations

AMD • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Comprehensive benefits package
Inclusive culture
Career advancement opportunities
AI & HPC GPU Compute Performance Engineer
AI & HPC GPU Compute Performance Engineer

engineeringjobs.net, Inc. • San Jose (CA)

On-site
USD 150,000 - 190,000
Medical Insurance
Dental Insurance
Vision Insurance
+16
Performance Engineer: AI/GPU Benchmarking & Optimization
Performance Engineer: AI/GPU Benchmarking & Optimization

Thomas To • Santa Clara (CA)

On-site
USD 136,000 - 213,000
Equity
Benefits
Senior Performance Engineer: AI Workload Optimization
Senior Performance Engineer: AI Workload Optimization

NVIDIA • Redmond (WA)

On-site
USD 224,000 - 431,250
Equity
Benefits
Senior AI Training Performance Engineer (GPU & Scale)
Senior AI Training Performance Engineer (GPU & Scale)

figure.ai • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Senior AI/HPC GPU Engineer - Performance & Systems Expert
Senior AI/HPC GPU Engineer - Performance & Systems Expert

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
AI DevTech Engineer: Equity & GPU-Accelerated AI
AI DevTech Engineer: Equity & GPU-Accelerated AI

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
AI Compute Performance Engineer
AI Compute Performance Engineer

Cisco Systems, Inc. • Milpitas (CA)

On-site
USD 155,000 - 223,000
Medical, dental, vision insurance
401(k) with Cisco matching
Paid parental leave
+2
AI Compute Performance Engineer
AI Compute Performance Engineer

020 Cisco Systems, Inc. • Milpitas (CA)

On-site
USD 155,000 - 223,000