GPU Infrastructure Engineer for AI & HPC

Vast.ai Inc.

Los Angeles, Northern (CA, KY)

Hybrid

USD 90,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Comprehensive health, dental, vision,
401(k) with company match
Meaningful equity
Onsite meals
Fast-paced startup culture

Job summary

Vast.ai is hiring for a full-time, on-site role in Westwood, Los Angeles to diagnose and resolve Linux and GPU infrastructure issues across NVIDIA drivers, CUDA, Docker and KVM. You’ll own escalations end-to-end, build runbooks, and collaborate with engineering and host support teams.

You’ll work autonomously in Ubuntu environments, automate with Python and Bash, and guide suppliers on hardware, BIOS, firmware, and network configuration to optimize performance.

Qualifications

  • Strong Linux systems operations experience with Ubuntu, RHEL/CentOS, or Debian, including networking, storage, services, and permissions
  • Proficiency with Docker, including container debugging, Docker Compose, image management, and storage/troubleshooting
  • Experience with virtualization platforms such as Proxmox VE, VMware, or similar hypervisors, including VM provisioning and troubleshooting
  • Strong networking fundamentals, including VLANs, DNS, DHCP, NAT, VPNs, firewall rules, and L2/L3 troubleshooting
  • Hands-on experience with NVIDIA GPU drivers, CUDA, and GPU workload troubleshooting
  • Python and Bash scripting for automation and diagnostics
  • Strong written English communication that is clear, professional, and technically precise
  • Experience providing technical support in a customer-facing or internal help desk environment
  • Ability to prioritize across a concurrent queue of escalated tickets, triaging by severity and customer impact, balancing reactive resolution with proactive documentation and tooling

Responsibilities

  • Diagnose and resolve issues across NVIDIA CUDA/GPU drivers, Docker, and KVM virtualization environments
  • Investigate GPU utilization, container resource constraints, thermal throttling, driver conflicts, and disk I/O bottlenecks
  • Assist clients and infrastructure suppliers working with TensorFlow, PyTorch, and other GPU-accelerated workloads
  • Troubleshoot network-layer issues, including VLAN, DNS, DHCP, VPN, NAT, firewall rules, and connectivity failures on host machines
  • Handle escalated support tickets involving GPU workload failures, container issues, networking problems, account infrastructure, and host-side configuration
  • Provide managed support for supplier onboarding and ongoing machine management, including installation, configuration, and post setup troubleshooting
  • Advise suppliers on hardware setup, driver configuration, BIOS and firmware settings, and network configuration for optimal performance
  • Provide coverage for L1 support overflow during peak periods or incidents
  • Write and maintain internal runbooks, escalation guides, and knowledge base articles to reduce repeat escalations
  • Build diagnostic and automation tooling in Python and Bash to reduce manual triage overhead
  • Collaborate with the engineering and support teams to flag and document systemic or recurring platform issues

Skills

Linux troubleshooting
NVIDIA drivers
Docker
KVM virtualization
Python scripting
Bash scripting
Networking basics
TensorFlow/PyTorch
GPU workloads
Customer support

Tools

Proxmox VE
VMware
Docker Compose

Job description

Vast.ai is hiring for a full-time, on-site role in Westwood, Los Angeles to diagnose and resolve Linux and GPU infrastructure issues across NVIDIA drivers, CUDA, Docker and KVM. You’ll own escalations end-to-end, build runbooks, and collaborate with engineering and host support teams.

You’ll work autonomously in Ubuntu environments, automate with Python and Bash, and guide suppliers on hardware, BIOS, firmware, and network configuration to optimize performance.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

On-site Linux Systems Ops Engineer – GPU & AI Infra
On-site Linux Systems Ops Engineer – GPU & AI Infra

Vast.ai • Los Angeles (CA)

On-site
USD 90,000 - 160,000
Health insurance
401(k) with company match
Meaningful equity
+2
AI, HPC & GPU Infrastructure Support Engineer New
AI, HPC & GPU Infrastructure Support Engineer New

Vast.ai Inc. • Los Angeles (CA), Northern (KY)

Hybrid
USD 90,000 - 160,000
Comprehensive health, dental, vision,
401(k) with company match
Meaningful equity
+2
GPU Systems Engineer - Scale AI Inference (On-site SF/LA)
GPU Systems Engineer - Scale AI Inference (On-site SF/LA)

Vast.ai Inc. • San Francisco (CA)

On-site
USD 140,000 - 210,000
Health insurance
Dental
Vision
+5
GPU Systems Engineer – HPC / Parallel Computing New
GPU Systems Engineer – HPC / Parallel Computing New

Vast.ai Inc. • San Francisco (CA)

On-site
USD 140,000 - 210,000
Health insurance
Dental
Vision
+5
Systems Software Engineer - GPU Cloud & Security
Systems Software Engineer - GPU Cloud & Security

Vast.ai • Los Angeles (CA)

On-site
USD 120,000 - 180,000
Comprehensive health, dental, vision,‑
401(k) with company match
Meaningful equity
+2
Senior Full-Stack Engineer — AI Infra & GPU Cloud (Equity)
Senior Full-Stack Engineer — AI Infra & GPU Cloud (Equity)

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Systems Software Engineer - GPU Cloud & Linux Performance
Systems Software Engineer - GPU Cloud & Linux Performance

Vast.ai • San Francisco (CA)

On-site
USD 140,000 - 200,000
AI Inference GPU Systems Engineer
AI Inference GPU Systems Engineer

Vast.ai Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Comprehensive health, dental, vision, and life insurance
401(k) with company match
Early-stage equity
+2
Staff Compute Infra Engineer - GPU & AI Systems
Staff Compute Infra Engineer - GPU & AI Systems

xAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
System Engineer
System Engineer

Acceler8 Talent • Fremont (CA)

On-site
USD 135,000 - 165,000
Comprehensive benefits