GPU Infra Engineer: Linux, NVIDIA, & Automation

Vast.ai

Los Angeles (CA)

On-site

USD 90,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401(k) with company match
Equity
Onsite meals
Startup culture

Job summary

Vast.ai is seeking a Linux/GPU infrastructure engineer to troubleshoot complex issues across NVIDIA drivers, CUDA, Docker, and KVM. The role is based in our Westwood, Los Angeles office with on-site schedules.

You will diagnose, reproduce, and resolve escalations, while collaborating with engineering to build runbooks and tooling. Strong Linux, GPU, Python, and English communication are essential for success.

Qualifications

  • Strong Linux systems operations across Ubuntu, RHEL/CentOS, Debian, including networking, storage, services, and permissions.
  • Proficiency with Docker, including debugging, Compose, image management, and storage troubleshooting.
  • Experience with virtualization platforms (Proxmox, VMware or similar) and VM provisioning.

Responsibilities

  • Diagnose and resolve issues across NVIDIA CUDA/GPU drivers, Docker, and KVM virtualization environments.
  • Investigate GPU utilization, container resource constraints, thermal throttling, driver conflicts, and disk I/O bottlenecks.
  • Assist clients and infrastructure suppliers working with TensorFlow, PyTorch, and other GPU-accelerated workloads.
  • Troubleshoot network-layer issues including VLAN, DNS, DHCP, VPN, NAT, firewall rules, and connectivity on host machines.
  • Handle escalated support tickets involving GPU workload failures, container issues, networking problems, and host-side configuration.
  • Provide managed support for supplier onboarding and ongoing machine management, including installation and post-setup troubleshooting.
  • Advise suppliers on hardware setup, driver configuration, BIOS and firmware settings, and network configuration for optimal performance.
  • Provide coverage for L1 support overflow during peak periods or incidents.
  • Write and maintain internal runbooks, escalation guides, and knowledge base articles to reduce repeat escalations.
  • Build diagnostic and automation tooling in Python and Bash to reduce manual triage overhead.
  • Collaborate with engineering and support teams to flag and document systemic or recurring platform issues.

Skills

Linux
NVIDIA CUDA
Python
Bash scripting
Networking

Tools

Docker
Proxmox VE
VMware
KVM

Job description

Vast.ai is seeking a Linux/GPU infrastructure engineer to troubleshoot complex issues across NVIDIA drivers, CUDA, Docker, and KVM. The role is based in our Westwood, Los Angeles office with on-site schedules.

You will diagnose, reproduce, and resolve escalations, while collaborating with engineering to build runbooks and tooling. Strong Linux, GPU, Python, and English communication are essential for success.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Infrastructure Engineer for AI & HPC
GPU Infrastructure Engineer for AI & HPC

Vast.ai Inc. • Los Angeles (CA), Northern (KY)

Hybrid
USD 90,000 - 160,000
Comprehensive health, dental, vision,
401(k) with company match
Meaningful equity
+2
On-site Linux Systems Ops Engineer – GPU & AI Infra
On-site Linux Systems Ops Engineer – GPU & AI Infra

Vast.ai • Los Angeles (CA)

On-site
USD 90,000 - 160,000
Health insurance
401(k) with company match
Meaningful equity
+2
GPU Infra Production Engineer: Automation & Validation
GPU Infra Production Engineer: Automation & Validation

Socket.dev • United States

Remote
USD 60,000 - 80,000
Health insurance
401k matching
Professional development
+6
Systems Software Engineer - GPU Cloud & Security
Systems Software Engineer - GPU Cloud & Security

Vast.ai • Los Angeles (CA)

On-site
USD 120,000 - 180,000
Comprehensive health, dental, vision,‑
401(k) with company match
Meaningful equity
+2
GPU Infra Production Engineer — Automation & Validation
GPU Infra Production Engineer — Automation & Validation

Vultr • United States

On-site
USD 60,000 - 80,000
Premium health insurance
401(k) with company match
Professional development reimbursement
+6
GPU Infra Production Engineer - Automation & Validation
GPU Infra Production Engineer - Automation & Validation

Webhosting • Northern (KY)

Hybrid
USD 60,000 - 80,000
Insurance
401(k) matching
Professional development reimbursement
+5
Remote Infra Operations Engineer - GPU Cloud
Remote Infra Operations Engineer - GPU Cloud

Nscale • Barstow (TX)

Remote
USD 80,000 - 110,000
Competitive pay with equity
Flexible workplace
Dynamic progression plan
Lead AI Infra Architect—GPU-Scale Infrastructure
Lead AI Infra Architect—GPU-Scale Infrastructure

NVIDIA • Austin (TX)

On-site
USD 184,000 - 288,000
Senior Full-Stack Engineer — AI Infra & GPU Cloud (Equity)
Senior Full-Stack Engineer — AI Infra & GPU Cloud (Equity)

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Security Engineer (Onsite LA) - GPU Cloud & DevSecOps
Security Engineer (Onsite LA) - GPU Cloud & DevSecOps

Vast.ai • Los Angeles (CA)

On-site
USD 140,000 - 210,000
Health benefits
401(k) with company match
Early-stage equity
+2