Systems Operations Support Engineer — Linux

Vast.ai

Los Angeles (CA)

On-site

USD 90,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401(k) with company match
Meaningful equity
Onsite meals
Snacks and collaboration with founders

Job summary

Vast.ai in Los Angeles seeks an experienced systems operations support engineer to tackle escalated infrastructure issues across hardware, BIOS, networking, Ubuntu, Docker, NVIDIA CUDA, and GPUs. You will own end-to-end escalations, build runbooks, and create automation to prevent recurrences.

You will work with L1 support, engineering, and host teams on systemic issues, communicating clearly with technical and non-technical audiences. This is a full-time, on-site role in Westwood, LA.

Qualifications

  • Strong Linux systems operations experience with Ubuntu or similar, including networking and storage.
  • Proficiency with Docker, debugging containers, and related tooling.
  • Experience with virtualization platforms (Proxmox, VMware) and GPU workloads.

Responsibilities

  • Handle escalated support tickets including GPU workload failures, container issues, networking problems, and host configuration.
  • Provide managed support for supplier onboarding and ongoing machine management, including installation and post-setup troubleshooting.
  • Assist clients and infrastructure suppliers with TensorFlow, PyTorch, and other GPU workloads.
  • Provide coverage for L1 support during peak periods or incidents.
  • Diagnose and resolve issues across Docker, CUDA drivers, and KVM virtualization environments.
  • Troubleshoot network-layer issues (VLAN, DNS, DHCP, VPN, NAT, firewall).
  • Investigate performance issues related to GPU utilization and container resources.
  • Advise suppliers on installation best practices (hardware setup, driver config, BIOS/firmware, networks).
  • Write and maintain runbooks, escalation guides, and knowledge base articles.
  • Build diagnostic and automation tooling in Python and Bash to reduce manual triage.
  • Collaborate with engineering and support to flag systemic issues.

Skills

Linux systems
Ubuntu
Docker
Networking
GPU/CUDA
Python scripting
Bash scripting
Virtualization
Troubleshooting
Clear written comms

Tools

Docker
Proxmox VE
VMware
NVIDIA CUDA drivers

Job description

About Us

Vast.ai's cloud powers AI projects and businesses all over the world. We are democratizing and decentralizing AI computing — reshaping our future for the benefit of humanity. Our mission is to organize, optimize, and orient the world's computation.

We value elegance, ownership, integrity, and continuous learning. You'll have the opportunity to dive into state-of-the-art AI systems while collaborating with a globally distributed team.

About the Role

This is a systems operations support role focused on deep-diving into escalated infrastructure issues that go beyond frontline triage. You’ll be the engineering resource our L1 support team relies on when tickets become complex, investigating and resolving issues across the full infrastructure stack—including hardware, BIOS and firmware, networking, Ubuntu, Docker, NVIDIA CUDA and GPUs, and KVM-based virtual machines.

You’ll own complex escalations end-to-end: gathering evidence, reproducing issues, identifying the root cause, proposing solutions, and working with the appropriate teams to bring each issue to resolution. The best engineers in this role don’t just resolve individual tickets—they identify recurring patterns, improve operational tooling, and build runbooks that prevent future incidents. You’ll collaborate directly with the engineering and host support teams on systemic infrastructure issues.

Strong Linux systems knowledge, technical depth, and support experience are the primary requirements. You should be comfortable working autonomously in Ubuntu environments, troubleshooting hardware, networking, containers, virtual machines, and GPU workloads, and clearly communicating your findings and proposed solutions to both technical and non-technical audiences.

Vast.ai users or hosts strongly preferred.

Location and Schedule

This is a full-time position based in our Westwood, Los Angeles office.

Available schedules:

  • Monday–Friday: Fully on-site

  • Sunday–Thursday: Four days on-site and one day working from home

Key Responsibilities

  • Handle escalated support tickets involving GPU workload failures, container issues, networking problems, account infrastructure, and host-side configuration

  • Provide managed support for supplier onboarding and ongoing machine management, acting as a technical resource through installation, configuration, and post-setup troubleshooting

  • Assist clients and infrastructure suppliers working with TensorFlow, PyTorch, and other GPU-accelerated workloads

  • Provide coverage for L1 support overflow during peak periods or incidents

  • Diagnose and resolve issues across Docker, NVIDIA CUDA/GPU drivers, and KVM virtualization environments

  • Troubleshoot network-layer issues, including VLAN, DNS, DHCP, VPN, NAT, firewall rules, and connectivity failures on host machines

  • Investigate performance issues involving GPU utilization, container resource constraints, thermal throttling, driver conflicts, and disk I/O bottlenecks

  • Advise suppliers on installation best practices, including hardware setup, driver configuration, BIOS/firmware settings, and network configuration for optimal performance

  • Write and maintain internal runbooks, escalation guides, and knowledge base articles to reduce repeat escalations

  • Build diagnostic and automation tooling in Python and Bash to reduce manual triage overhead

  • Collaborate with the engineering and support teams to flag and document systemic or recurring platform issues

You Are

  • Experienced with Linux, especially Ubuntu, and comfortable troubleshooting from the command line

  • Someone who enjoys debugging difficult problems and fixing broken systems

  • Methodical and focused on finding root causes, not just temporary fixes

  • Able to manage complex tickets independently

  • A clear written communicator with an interest in AI infrastructure and GPU computing

Must-Haves

  • Strong Linux systems operations experience with Ubuntu, RHEL/CentOS, or Debian, including networking, storage, services, and permissions

  • Proficiency with Docker, including container debugging, Docker Compose, image management, cgroup limits, and Docker storage and filesystem troubleshooting

  • Experience with virtualization platforms such as Proxmox VE, VMware, or similar hypervisors, including VM provisioning and troubleshooting

  • Strong networking fundamentals, including VLANs, DNS, DHCP, NAT, VPNs, firewall rules, and L2/L3 troubleshooting

  • Hands‑on experience with NVIDIA GPU drivers, CUDA, and GPU workload troubleshooting

  • Python and Bash scripting skills for automation and diagnostic tooling

  • Strong written English communication that is clear, professional, and technically precise

  • Experience providing technical support in a customer‑facing or internal help desk environment

  • Ability to prioritize across a concurrent queue of escalated tickets, triaging by severity and customer impact, balancing reactive resolution against proactive documentation and tooling work, and making clear judgment calls on when to escape versus own resolution end-to-end

Nice-to-Haves

  • Familiarity with AI/ML frameworks (TensorFlow, PyTorch) and running GPU‑accelerated containers

  • Monitoring and observability experience (Prometheus, Grafana)

  • Relevant certifications: RHCSA, CompTIA Linux+, or similar

  • Knowledge of the Vast.ai platform as a client or infrastructure supplier

Interview Process (~1 week)

After you submit your application, our technical team will review your experience and qualifications. Selected candidates will proceed through the following stages:

  • 15 minutes — Initial Screening (Virtual): A brief conversation about your background, availability, and interest in the role

  • 45 minutes — Experience Interview (Virtual): An introduction to Vast.ai and a deeper discussion of your technical and support experience

  • 2 hours — Meet and Greet and Technical Assessment (On-site): Meet the team and complete an LLM‑assisted Linux systems operations assessment

Annual Salary Range

$90,000 — $160,000 + equity + benefits

Vast.ai is hiring across all experience levels with compensation commensurate with background, experience and potential.

Benefits
  • Comprehensive health, dental, vision, and life insurance

  • 401(k) with company match

  • Meaningful early‑stage equity

  • Onsite meals, snacks, and close collaboration with founders/tech leaders

  • Ambitious, fast‑paced startup culture where initiative is rewarded

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI, HPC & GPU Infrastructure Support Engineer New
AI, HPC & GPU Infrastructure Support Engineer New

Vast.ai Inc. • Los Angeles (CA), Northern (KY)

Hybrid
USD 90,000 - 160,000
Comprehensive health, dental, vision,
401(k) with company match
Meaningful equity
+2
Technical Support Engineer II (Linux)
Technical Support Engineer II (Linux)

Vast.ai Inc. • Los Angeles (CA)

On-site
USD 90,000 - 150,000
Health insurance
401(k) with company match
Meaningful equity
+2
HPC & GPU Infrastructure Support Engineer
HPC & GPU Infrastructure Support Engineer

Vast.ai • Los Angeles (CA)

On-site
USD 90,000 - 160,000
Health insurance
401(k) with company match
Equity
+2
C++ Software Engineer
C++ Software Engineer

Vast.ai • Los Angeles (CA)

On-site
USD 120,000 - 180,000
Comprehensive health, dental, vision,‑
401(k) with company match
Meaningful equity
+2
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Vast.ai • San Francisco (CA)

On-site
USD 180,000 - 230,000
Health, dental, vision, life insurance
401(k) with company match
Equity
+2
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Vast.ai Inc. • San Francisco (CA)

On-site
USD 180,000 - 230,000
Health benefits
401(k) with company match
Equity
+2
Systems/GPU Research Engineer
Systems/GPU Research Engineer

Vast.ai Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Comprehensive health, dental, vision, and life insurance
401(k) with company match
Early-stage equity
+2
Technical Product Manager, Infrastructure New
Technical Product Manager, Infrastructure New

Vast.ai Inc. • Los Angeles (CA)

On-site
USD 170,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+4
Systems/GPU Research Engineer
Systems/GPU Research Engineer

Vast.ai • Los Angeles (CA)

On-site
USD 160,000 - 320,000
Comprehensive health, dental, vision, and life insurance
401(k) with company match
Meaningful early-stage equity
+2
Developer Relations Engineer
Developer Relations Engineer

Vast.ai • Los Angeles (CA)

On-site
USD 160,000 - 200,000
On-site in SF/LA
Health, dental, vision insurance
Travel/conference budget