GPU Systems Engineer - Scale AI Inference (On-site SF/LA)

Vast.ai Inc.

San Francisco (CA)

On-site

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Dental
Vision
Life insurance
401(k) match
Equity
Onsite meals
Startup culture

Job summary

Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance.

This role is based on-site in San Francisco or Los Angeles, with a strong emphasis on HPC techniques and scalable workloads. You will collaborate with technical leadership, evaluate architectures, and contribute to improving GPU

Qualifications

  • Advanced C++ skills with modern standards (C++17/20).
  • Experience with at least one parallel framework (CUDA, HIP, SYCL, OpenCL, OpenACC).
  • Strong background in systems optimization and HPC performance tooling.

Responsibilities

  • Design and optimize GPU kernels and tensor libraries.
  • Translate HPC techniques into scalable AI inference solutions.
  • Evaluate emerging architectures and resource management approaches.
  • Collaborate with leadership to improve GPU infrastructure efficiency.

Skills

CUDA/C++
GPGPU
Python
Linux
Advanced C++ (C++17/20)
Parallel frameworks (CUDA, HIP, SYCL, 

Job description

Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance.

This role is based on-site in San Francisco or Los Angeles, with a strong emphasis on HPC techniques and scalable workloads. You will collaborate with technical leadership, evaluate architectures, and contribute to improving GPU

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Inference GPU Systems Engineer
AI Inference GPU Systems Engineer

Vast.ai Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Comprehensive health, dental, vision, and life insurance
401(k) with company match
Early-stage equity
+2
GPU Systems Engineer – HPC / Parallel Computing New
GPU Systems Engineer – HPC / Parallel Computing New

Vast.ai Inc. • San Francisco (CA)

On-site
USD 140,000 - 210,000
Health insurance
Dental
Vision
+5
GPU Systems Engineer — Distributed Training & Inference
GPU Systems Engineer — Distributed Training & Inference

TensorScale AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
GPU Systems Research Engineer for AI Inference
GPU Systems Research Engineer for AI Inference

Vast.ai • Los Angeles (CA)

On-site
USD 160,000 - 320,000
Comprehensive health, dental, vision, and life insurance
401(k) with company match
Meaningful early-stage equity
+2
GPU Infrastructure Engineer for AI & HPC
GPU Infrastructure Engineer for AI & HPC

Vast.ai Inc. • Los Angeles (CA), Northern (KY)

Hybrid
USD 90,000 - 160,000
Comprehensive health, dental, vision,
401(k) with company match
Meaningful equity
+2
Senior AI Inference Systems Engineer | GPU Kernels & Runtime
Senior AI Inference Systems Engineer | GPU Kernels & Runtime

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Systems/GPU Research Engineer
Systems/GPU Research Engineer

Vast.ai • Los Angeles (CA)

On-site
USD 160,000 - 320,000
Comprehensive health, dental, vision, and life insurance
401(k) with company match
Meaningful early-stage equity
+2
On-site Linux Systems Ops Engineer – GPU & AI Infra
On-site Linux Systems Ops Engineer – GPU & AI Infra

Vast.ai • Los Angeles (CA)

On-site
USD 90,000 - 160,000
Health insurance
401(k) with company match
Meaningful equity
+2
Senior ML Systems Engineer — Inference & Scale
Senior ML Systems Engineer — Inference & Scale

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000
Staff Engineer, GPU AI Inference & RL Infrastructure
Staff Engineer, GPU AI Inference & RL Infrastructure

B Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive medical, dental, and vision insurance
Fully paid parental leave
+2