Senior AI Performance Engineer

Brillfy Technology Inc

United States

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Brillfy Technology Inc. is seeking a hands-on GPU/AI systems engineer focused on CUDA optimization, multi-GPU performance, and system-wide profiling.

The role requires deep expertise in GPU architecture and performance tuning, with experience across NVIDIA tools and distributed AI stacks. The ideal candidate will profile AI/ML workloads, identify bottlenecks, and implement latency and throughput enhancements across compute, memory, and networking layers, collaborating with AI, DevOps, and

Qualifications

  • Strong hands-on CUDA programming and GPU performance optimization.
  • Deep understanding of GPU architecture and memory hierarchy.
  • Experience with Nsight and CUDA profiling tools.

Responsibilities

  • Profile and optimize AI/ML workloads across multi-GPU and multi-node systems.
  • Identify bottlenecks across compute, memory, networking, and orchestration layers.
  • Optimize CUDA kernels (memory coalescing, shared memory usage, occupancy tuning).
  • Improve inference performance using TensorRT, Triton, DeepStream, NeMo.
  • Analyze and improve latency, throughput, GPU utilization, and memory efficiency.
  • Work on distributed AI systems using Apache Ray, NCCL, Kubernetes GPU scheduling.
  • Build benchmarking frameworks and performance monitoring systems.
  • Collaborate with AI, DevOps, and Infrastructure teams for system-wide optimization.

Skills

CUDA programming
GPU architecture
Performance profiling
NVIDIA ecosystem (TensorRT, Triton, Ne

Tools

Nsight tools

Job description

Job Description

This is a hands-on engineering role, requiring deep expertise in CUDA, GPU architecture, and performance profiling.


Key Responsibilities


  • Profile and optimize AI/ML workloads across multi-GPU and multi-node systems

  • Identify bottlenecks across compute, memory, networking, and orchestration layers

  • Optimize CUDA kernels (memory coalescing, shared memory usage, occupancy tuning)

  • Improve inference performance using TensorRT, Triton, DeepStream, NeMo

  • Analyze and improve latency, throughput, GPU utilization, and memory efficiency

  • Work on distributed AI systems using Apache Ray, NCCL, Kubernetes GPU scheduling

  • Build benchmarking frameworks and performance monitoring systems

  • Collaborate with AI, DevOps, and Infrastructure teams for system-wide optimization


Required Skills


  • Strong hands-on CUDA programming and GPU performance optimization

  • Deep understanding of GPU architecture and memory hierarchy

  • Experience with Nsight, CUDA profiling tools, performance benchmarking

  • Hands-on experience with NVIDIA ecosystem (Triton, TensorRT, NeMo, DeepStream)

  • Experience with distributed AI systems (multi-GPU, multi-node, NCCL, Ray)

  • Experience working with AI models such as YOLO, GPT, LLaMA, Transformers

  • Strong understanding of AI system performance metrics (latency, throughput, utilization)


Preferred


  • Experience working at NVIDIA or similar GPU/AI infrastructure companies

  • Experience with real-time video / Vision AI systems

  • Experience with large-scale production AI deployments


Interview Process (Mandatory)


  • Candidates will receive a technical handout 1 day before interview

  • 90-minute deep-dive demo discussion (NOT theoretical)

  • Candidate must explain:

  • Bottleneck identification approach

  • GPU optimization strategies

  • System-level performance improvements

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Performance and Efficiency Engineer
Senior AI Performance and Efficiency Engineer

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Competitive benefits
Senior AI Training Performance Architect
Senior AI Training Performance Architect

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Research Engineer, GPU Performance
Research Engineer, GPU Performance

Harnham • California (MO)

On-site
USD 120,000 - 160,000
Senior Software Engineer, AI Performance Analysis
Senior Software Engineer, AI Performance Analysis

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Solutions Architect, AI Performance Engineering
Senior Solutions Architect, AI Performance Engineering

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits package
Senior AI Performance and Efficiency Engineer
Senior AI Performance and Efficiency Engineer

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Competitive salaries
Comprehensive benefits package
Equity eligibility
Senior AI Training Performance Architect
Senior AI Training Performance Architect

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Data Center Performance Engineer - Benchmarking and Optimization
Senior Data Center Performance Engineer - Benchmarking and Optimization

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 287,500
Equity
Benefits
Senior System Software Engineer - AI Performance and Efficiency Tools
Senior System Software Engineer - AI Performance and Efficiency Tools

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Senior Performance Engineer - DGX Cloud
Senior Performance Engineer - DGX Cloud

NVIDIA • Austin (TX)

On-site
USD 272,000 - 432,000
Equity
Benefits