On-site Linux Systems Ops Engineer – GPU & AI Infra

Vast.ai

Los Angeles (CA)

On-site

USD 90,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401(k) with company match
Meaningful equity
Onsite meals
Snacks and collaboration with founders

Job summary

Vast.ai in Los Angeles seeks an experienced systems operations support engineer to tackle escalated infrastructure issues across hardware, BIOS, networking, Ubuntu, Docker, NVIDIA CUDA, and GPUs. You will own end-to-end escalations, build runbooks, and create automation to prevent recurrences.

You will work with L1 support, engineering, and host teams on systemic issues, communicating clearly with technical and non-technical audiences. This is a full-time, on-site role in Westwood, LA.

Qualifications

  • Strong Linux systems operations experience with Ubuntu or similar, including networking and storage.
  • Proficiency with Docker, debugging containers, and related tooling.
  • Experience with virtualization platforms (Proxmox, VMware) and GPU workloads.

Responsibilities

  • Handle escalated support tickets including GPU workload failures, container issues, networking problems, and host configuration.
  • Provide managed support for supplier onboarding and ongoing machine management, including installation and post-setup troubleshooting.
  • Assist clients and infrastructure suppliers with TensorFlow, PyTorch, and other GPU workloads.
  • Provide coverage for L1 support during peak periods or incidents.
  • Diagnose and resolve issues across Docker, CUDA drivers, and KVM virtualization environments.
  • Troubleshoot network-layer issues (VLAN, DNS, DHCP, VPN, NAT, firewall).
  • Investigate performance issues related to GPU utilization and container resources.
  • Advise suppliers on installation best practices (hardware setup, driver config, BIOS/firmware, networks).
  • Write and maintain runbooks, escalation guides, and knowledge base articles.
  • Build diagnostic and automation tooling in Python and Bash to reduce manual triage.
  • Collaborate with engineering and support to flag systemic issues.

Skills

Linux systems
Ubuntu
Docker
Networking
GPU/CUDA
Python scripting
Bash scripting
Virtualization
Troubleshooting
Clear written comms

Tools

Docker
Proxmox VE
VMware
NVIDIA CUDA drivers

Job description

Vast.ai in Los Angeles seeks an experienced systems operations support engineer to tackle escalated infrastructure issues across hardware, BIOS, networking, Ubuntu, Docker, NVIDIA CUDA, and GPUs. You will own end-to-end escalations, build runbooks, and create automation to prevent recurrences.

You will work with L1 support, engineering, and host teams on systemic issues, communicating clearly with technical and non-technical audiences. This is a full-time, on-site role in Westwood, LA.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Infrastructure Engineer for AI & HPC
GPU Infrastructure Engineer for AI & HPC

Vast.ai Inc. • Los Angeles (CA), Northern (KY)

Hybrid
USD 90,000 - 160,000
Comprehensive health, dental, vision,
401(k) with company match
Meaningful equity
+2
GPU Infra Engineer: Linux, NVIDIA, & Automation
GPU Infra Engineer: Linux, NVIDIA, & Automation

Vast.ai • Los Angeles (CA)

On-site
USD 90,000 - 160,000
Health insurance
401(k) with company match
Equity
+2
Security Engineer (Onsite LA) - GPU Cloud & DevSecOps
Security Engineer (Onsite LA) - GPU Cloud & DevSecOps

Vast.ai • Los Angeles (CA)

On-site
USD 140,000 - 210,000
Health benefits
401(k) with company match
Early-stage equity
+2
Security Engineer - GPU Cloud, DevSecOps & Equity
Security Engineer - GPU Cloud, DevSecOps & Equity

Vast.ai Inc. • Los Angeles (CA)

On-site
USD 150,000 - 210,000
Comprehensive health, dental, vision,
401(k) with company match
Meaningful early-stage equity
+1
Systems Software Engineer - GPU Cloud & Security
Systems Software Engineer - GPU Cloud & Security

Vast.ai • Los Angeles (CA)

On-site
USD 120,000 - 180,000
Comprehensive health, dental, vision,‑
401(k) with company match
Meaningful equity
+2
Systems Operations Support Engineer — Linux
Systems Operations Support Engineer — Linux

Vast.ai • Los Angeles (CA)

On-site
USD 90,000 - 160,000
Health insurance
401(k) with company match
Meaningful equity
+2
GPU Systems Engineer - Scale AI Inference (On-site SF/LA)
GPU Systems Engineer - Scale AI Inference (On-site SF/LA)

Vast.ai Inc. • San Francisco (CA)

On-site
USD 140,000 - 210,000
Health insurance
Dental
Vision
+5
GPU HPC Engineer for AI Inference - SF/LA On-site
GPU HPC Engineer for AI Inference - SF/LA On-site

Vast.ai • Los Angeles (CA)

On-site
USD 120,000 - 180,000
Health, dental, vision, and life保险
401(k) with company match
Meaningful early-stage equity
+2
HPC & GPU Infrastructure Support Engineer
HPC & GPU Infrastructure Support Engineer

Vast.ai • Los Angeles (CA)

On-site
USD 90,000 - 160,000
Health insurance
401(k) with company match
Equity
+2
AI, HPC & GPU Infrastructure Support Engineer New
AI, HPC & GPU Infrastructure Support Engineer New

Vast.ai Inc. • Los Angeles (CA), Northern (KY)

Hybrid
USD 90,000 - 160,000
Comprehensive health, dental, vision,
401(k) with company match
Meaningful equity
+2