Get more replies from employers
Send a job-specific resume in minutes.
TensorScale AI in San Francisco seeks a hardware‑aware software engineer to optimize GPU systems for training and inference across image, video, and world‑model workloads. You will push performance from kernel code to distributed engines, profiling bottlenecks and implementing practical improvements that scale.
Responsibilities include CUDA / Triton optimizations, designing efficient distributed inference and training pipelines, and owning communication performance across GPUs and nodes with
TensorScale AI in San Francisco seeks a hardware‑aware software engineer to optimize GPU systems for training and inference across image, video, and world‑model workloads. You will push performance from kernel code to distributed engines, profiling bottlenecks and implementing practical improvements that scale.
Responsibilities include CUDA / Triton optimizations, designing efficient distributed inference and training pipelines, and owning communication performance across GPUs and nodes with