GPU Systems Engineer — Distributed Training & Inference

TensorScale AI

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

24 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

TensorScale AI in San Francisco seeks a hardware‑aware software engineer to optimize GPU systems for training and inference across image, video, and world‑model workloads. You will push performance from kernel code to distributed engines, profiling bottlenecks and implementing practical improvements that scale.

Responsibilities include CUDA / Triton optimizations, designing efficient distributed inference and training pipelines, and owning communication performance across GPUs and nodes with

Qualifications

  • 3+ years in deep learning inference/training systems, distributed systems, or high-performance computing.
  • Proficiency in CUDA, with hands-on GPU profiling (Nsight).
  • Strong systems instincts across GPU architecture, parallel programming, and compute kernels.
  • Experience with distributed multi-GPU debugging and optimization.
  • Familiarity with PyTorch and performance-critical model execution.
  • Comfort working onsite with the team in the Bay Area.

Responsibilities

  • Optimize system and GPU performance of training and inference for image, video, and world-model workloads.
  • Profile and remove bottlenecks at kernel, memory, system, and cluster levels (Nsight and related tooling).
  • Implement low-level optimizations in CUDA / Triton.
  • Design and improve distributed inference / training engines for diffusion models.
  • Own communication performance across GPUs and nodes: NCCL, RDMA / InfiniBand / RoCE, and disaggregated serving.
  • Build benchmarking and regression harnesses so performance gains stick in production.
  • Hardware-aware kernel, runtime, and model co-design.

Skills

CUDA proficiency
Nsight profiling
Distributed systems
PyTorch
Multi-GPU debugging
Performance optimization

Tools

Nsight
NCCL
InfiniBand
RoCE
TensorRT

Job description

TensorScale AI in San Francisco seeks a hardware‑aware software engineer to optimize GPU systems for training and inference across image, video, and world‑model workloads. You will push performance from kernel code to distributed engines, profiling bottlenecks and implementing practical improvements that scale.

Responsibilities include CUDA / Triton optimizations, designing efficient distributed inference and training pipelines, and owning communication performance across GPUs and nodes with

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Systems Engineer - Scale AI Inference (On-site SF/LA)
GPU Systems Engineer - Scale AI Inference (On-site SF/LA)

Vast.ai Inc. • San Francisco (CA)

On-site
USD 140,000 - 210,000
Health insurance
Dental
Vision
+5
Senior System Software Engineer — GPU AI Inference (Triton)
Senior System Software Engineer — GPU AI Inference (Triton)

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior ML Training Systems Engineer - Distributed CUDA
Senior ML Training Systems Engineer - Distributed CUDA

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior AI Inference Systems Engineer | GPU Kernels & Runtime
Senior AI Inference Systems Engineer | GPU Kernels & Runtime

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Senior AI Training Performance Engineer (GPU & Scale)
Senior AI Training Performance Engineer (GPU & Scale)

figure.ai • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Senior Training Infra Engineer - 800+ GPU Scale
Senior Training Infra Engineer - 800+ GPU Scale

Figureai • San Jose (CA)

On-site
USD 150,000 - 350,000
GPU Systems Engineer for AI Training Clusters
GPU Systems Engineer for AI Training Clusters

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Member of Technical Staff, ML Systems
Member of Technical Staff, ML Systems

TensorScale AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior GPU Systems Engineer: Scale AI Performance
Senior GPU Systems Engineer: Scale AI Performance

NVIDIA AI • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits