GPU ML Infra Intern: Speed Up Training & Profiling

Plus 2

Santa Clara (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Free lunch
Snacks & drinks
401(k) plan

Job summary

PlusAI in Silicon Valley seeks a highly skilled systems/performance engineer to optimize BEV model training pipelines and GPU kernels. You will profile, implement custom kernels with CUDA or Triton, and push performance through LLM-assisted code generation and profiling workflows.

Key focus areas include memory management, PyTorch custom ops, and GPU utilization, with opportunities to influence autonomous vehicle perception workloads and high-performance computing practices.

Responsibilities

  • Identify Training Bottlenecks: Profile BEV model training pipelines to pinpoint computational and memory bottlenecks.
  • Develop Custom Kernels: Design and implement high-performance kernels using CUDA, Triton, or C++ to accelerate training.
  • Leverage LLMs for Optimization: Explore and integrate LLMs to generate high-performance code and optimize kernel logic.
  • Automate Profiling Workflows: Build systems to automate performance profiling with NVIDIA Nsight and PyTorch Profiler.
  • Iterative Performance Tuning: Analyze profiling data to maximize GPU utilization and reduce training times.
  • Systems Programming: Strong proficiency in C++ with memory management and parallel processing principles.
  • Deep Learning Frameworks: Hands-on PyTorch experience with custom operations and autograd.
  • Performance-Oriented Mindset: Problem-solving with a focus on low-level optimization.
  • GPU Programming Experience: Experience writing and optimizing custom GPU kernels with CUDA or Triton.
  • Hardware Profiling Tools: Familiarity with profiling tools like Nsight and PyTorch Profiler.
  • LLM for Code Generation: Experience with LLMs for code writing or refactoring.
  • Autonomous Vehicle Perception: Understanding BEV models and 3D perception.

Job description

PlusAI in Silicon Valley seeks a highly skilled systems/performance engineer to optimize BEV model training pipelines and GPU kernels. You will profile, implement custom kernels with CUDA or Triton, and push performance through LLM-assisted code generation and profiling workflows.

Key focus areas include memory management, PyTorch custom ops, and GPU utilization, with opportunities to influence autonomous vehicle perception workloads and high-performance computing practices.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
ML Performance Engineer: Scale GPU-Driven Training
ML Performance Engineer: Scale GPU-Driven Training

Decisive Point • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
ML Systems Engineer: Optimizing Training & GPU Kernels
ML Systems Engineer: Optimizing Training & GPU Kernels

Jobtailor • Massachusetts

On-site
USD 120,000 - 180,000
GPU Transformer Performance Engineer (Triton/CUDA)
GPU Transformer Performance Engineer (Triton/CUDA)

Luma AI • United States

Remote
USD 180,000 - 280,000
Machine Learning Infrastructure Engineer Intern
Machine Learning Infrastructure Engineer Intern

Plus 2 • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Free lunch
Snacks & drinks
401(k) plan
Remote ML Performance Engineer: AI Throughput & GPU Optimizer
Remote ML Performance Engineer: AI Throughput & GPU Optimizer

Bright Vision Technologies • Redmond (WA)

On-site
USD 100,000 - 150,000
Member of Technical Staff, ML Systems
Member of Technical Staff, ML Systems

TensorScale AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lead GPU ML Training Performance Engineer
Lead GPU ML Training Performance Engineer

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 120,000 - 160,000
Research Engineer, GPU Performance
Research Engineer, GPU Performance

Harnham • California (MO)

On-site
USD 120,000 - 160,000
Senior Inference Performance Engineer — Equity & Hybrid
Senior Inference Performance Engineer — Equity & Hybrid

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000