On-site Linux Systems Ops Engineer – GPU & AI Infra

Vast.ai

Los Angeles (CA)

On-site

USD 90,000 - 160,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) with company match
Meaningful equity
Onsite meals
Snacks and collaboration with founders

Job summary

Vast.ai in Los Angeles seeks an experienced systems operations support engineer to tackle escalated infrastructure issues across hardware, BIOS, networking, Ubuntu, Docker, NVIDIA CUDA, and GPUs. You will own end-to-end escalations, build runbooks, and create automation to prevent recurrences.

You will work with L1 support, engineering, and host teams on systemic issues, communicating clearly with technical and non-technical audiences. This is a full-time, on-site role in Westwood, LA.

Qualifications

  • Strong Linux systems operations experience with Ubuntu or similar, including networking and storage.
  • Proficiency with Docker, debugging containers, and related tooling.
  • Experience with virtualization platforms (Proxmox, VMware) and GPU workloads.

Responsibilities

  • Handle escalated support tickets including GPU workload failures, container issues, networking problems, and host configuration.
  • Provide managed support for supplier onboarding and ongoing machine management, including installation and post-setup troubleshooting.
  • Assist clients and infrastructure suppliers with TensorFlow, PyTorch, and other GPU workloads.
  • Provide coverage for L1 support during peak periods or incidents.
  • Diagnose and resolve issues across Docker, CUDA drivers, and KVM virtualization environments.
  • Troubleshoot network-layer issues (VLAN, DNS, DHCP, VPN, NAT, firewall).
  • Investigate performance issues related to GPU utilization and container resources.
  • Advise suppliers on installation best practices (hardware setup, driver config, BIOS/firmware, networks).
  • Write and maintain runbooks, escalation guides, and knowledge base articles.
  • Build diagnostic and automation tooling in Python and Bash to reduce manual triage.
  • Collaborate with engineering and support to flag systemic issues.

Skills

Linux systems
Ubuntu
Docker
Networking
GPU/CUDA
Python scripting
Bash scripting
Virtualization
Troubleshooting
Clear written comms

Tools

Docker
Proxmox VE
VMware
NVIDIA CUDA drivers

Job description

Vast.ai in Los Angeles seeks an experienced systems operations support engineer to tackle escalated infrastructure issues across hardware, BIOS, networking, Ubuntu, Docker, NVIDIA CUDA, and GPUs. You will own end-to-end escalations, build runbooks, and create automation to prevent recurrences.

You will work with L1 support, engineering, and host teams on systemic issues, communicating clearly with technical and non-technical audiences. This is a full-time, on-site role in Westwood, LA.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Security Engineer (Onsite LA) - GPU Cloud & DevSecOps
Security Engineer (Onsite LA) - GPU Cloud & DevSecOps

Vast.ai • Los Angeles (CA)

On-site
USD 140,000 - 210,000
Health benefits
401(k) with company match
Early-stage equity
+2
Systems Software Engineer - GPU Cloud & Security
Systems Software Engineer - GPU Cloud & Security

Vast.ai • Los Angeles (CA)

On-site
USD 120,000 - 180,000
Comprehensive health, dental, vision,‑
401(k) with company match
Meaningful equity
+2
Systems Operations Support Engineer — Linux
Systems Operations Support Engineer — Linux

Vast.ai • Los Angeles (CA)

On-site
USD 90,000 - 160,000
Health insurance
401(k) with company match
Meaningful equity
+2
On-Site Infra Engineer: Kubernetes, GPUs & AI
On-Site Infra Engineer: Kubernetes, GPUs & AI

Nebula • El Segundo (CA)

On-site
USD 100,000 - 250,000
Equity participation
Health, dental, vision
401(k) with matching
+2
Senior Systems & Infra Automation Engineer
Senior Systems & Infra Automation Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior AI-Driven Cloud & Linux Infra Engineer
Senior AI-Driven Cloud & Linux Infra Engineer

NVIDIA • California (MO)

On-site
USD 168,000 - 270,000
Customer Support Engineer, AI/GPU Infra
Customer Support Engineer, AI/GPU Infra

SF Compute • San Francisco (CA)

On-site
USD 120,000 - 165,000
Generous equity
Visa sponsorship
401(k) matching
+5
Staff Compute Infra Engineer - GPU & AI Systems
Staff Compute Infra Engineer - GPU & AI Systems

xAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Comprehensive medical, vision, and dental coverage
Access to a 401(k) retirement plan
+3
Customer Support Engineer – GPU/AI Infra
Customer Support Engineer – GPU/AI Infra

The San Francisco Compute Company • San Francisco (CA)

On-site
USD 90,000 - 120,000
Generous equity grant
Visa sponsorships
Retirement matching
+5
Senior GPU Infra Engineer — Remote
Senior GPU Infra Engineer — Remote

Nscale • Seattle (WA)

On-site
USD 120,000 - 170,000
Remote-first culture
Equity plan
Flexible workplace