Technical Support Engineer II (Linux)

Vast.ai Inc.

Los Angeles (CA)

On-site

USD 90,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401(k) with company match
Meaningful equity
Onsite meals
Startup culture

Job summary

Vast.ai in Los Angeles is seeking a senior technical support engineer specializing in escalated infrastructure issues. You will diagnose and resolve issues across Ubuntu, Docker, NVIDIA CUDA/GPU, and virtualization (KVM), while building runbooks and tooling to reduce repeats.

Friendly for autonomous work and collaboration with engineering. This full-time, on-site role at our Westwood office requires strong Linux, Docker, GPU, and networking expertise, with on-call responsibilities and onboarding

Qualifications

  • Solid Linux SysOps experience across Ubuntu Server, RHEL/CentOS, Debian.
  • Proficiency with Docker: container debugging, image management, cgroup limits.
  • Experience with virtualization: Proxmox VE, VMware; VM provisioning and troubleshooting.
  • Networking fundamentals: VLAN, DNS, DHCP, NAT, VPN, firewall rules.
  • Hands-on NVIDIA GPU drivers, CUDA, and GPU workload troubleshooting.
  • Scripting in Python and Bash for automation.
  • Strong English written communication: clear and professional.
  • Customer-facing support or internal helpdesk experience.
  • Ability to prioritize concurrent escalated tickets and decide on escalation.

Responsibilities

  • Handle escalated support tickets including GPU workload failures and container issues.
  • Diagnose and resolve Docker, CUDA, and virtualization issues (KVM).
  • Troubleshoot networking on host machines: VLAN, DNS, DHCP, VPN, NAT, firewall.
  • Investigate GPU utilization, container resources, thermal throttling, driver conflicts.
  • Advise hosts on installation, driver config, BIOS settings for performance.
  • Provide managed support for onboarding and ongoing machine management.
  • Write and maintain runbooks, escalation guides, and knowledge base articles.
  • Build diagnostic and automation tooling in Python and Bash.
  • Collaborate with engineering to flag systemic issues and document them.
  • Assist clients and hosts using AI frameworks (TensorFlow, PyTorch).
  • Provide coverage for L1 support during peak periods per on-call rotation.

Skills

Linux SysOps
Docker
Networking
NVIDIA GPU
CUDA
Python
Bash
English communication
On-call readiness
Troubleshooting depth
Documentation writing

Tools

Docker
Proxmox VE
VMware
NVIDIA drivers
Python
Bash

Job description

About Us

Vast.ai's cloud powers AI projects and businesses all over the world. We are democratizing and decentralizing AI computing — reshaping our future for the benefit of humanity. Our mission is to organize, optimize, and orient the world's computation.

We value elegance, ownership, integrity, and continuous learning. You'll have the opportunity to dive into state-of-the-art AI systems while collaborating with a globally distributed team.

About the Role

This is a technical support role focused on escalated infrastructure issues that go beyond frontline triage. You'll be the engineering resource our L1 support team leans on when tickets get complex: diagnosing and resolving issues across the full stack — hardware/BIOS/firmware, networking, Ubuntu, Docker, NVIDIA CUDA/GPU, and virtualization (KVM).

You'll handle higher-complexity issues, own escalation resolution end-to-end, and contribute to internal documentation and runbooks. The best engineers in this role don't just resolve tickets — they build the tooling and runbooks that eliminate recurring ones. You'll collaborate directly with the engineering team and host support team on systemic issues.

Strong technical depth and support experience are the primary requirements. You should be comfortable working autonomously across Ubuntu environments, diagnosing container and GPU issues, and communicating findings clearly to both technical and non-technical audiences.

Vast.ai users or hosts strongly preferred.

This role is full-time and onsite in our office in Westwood (LA)
Schedule: Sunday - Thursday.

Key Responsibilities
  • Handle escalated support tickets, including GPU workload failures, container issues, networking problems, account infrastructure, and host-side configuration

  • Diagnose and resolve issues across Docker, NVIDIA CUDA/GPU drivers, and virtualization environments (KVM)

  • Troubleshoot network-layer issues: VLAN, DNS, DHCP, VPN, NAT, firewall rules, and connectivity failures on host machines

  • Investigate performance issues on GPU utilization, container resource constraints, thermal throttling, driver conflicts, disk I/O bottlenecks

  • Advise suppliers (hosts) on installation best practices — hardware setup, driver configuration, BIOS/firmware settings, and network configuration for optimal performance

  • Provide managed support for supplier onboarding and ongoing machine management, acting as a technical resource through installation, configuration, and post-setup troubleshooting

  • Write and maintain internal runbooks, escalation guides, and knowledge base articles to reduce repeat escalations

  • Build diagnostic and automation tooling in Python and Bash to reduce manual triage overhead

  • Collaborate with the engineering team and infrastructure support team to flag and document systemic or recurring platform issues

  • Assist clients and infrastructure suppliers working with AI frameworks (TensorFlow, PyTorch) and GPU-accelerated workloads

  • Provide coverage for L1 support team overflow during peak periods or incidents, per a defined on-call rotation

You Are
  • Fluent in Linux — you navigate systems, read logs, and solve problems from the command line without hesitation

  • Methodical and thorough: you gather data, dig into root causes, and don't settle for surface-level fixes

  • A self-starter who can manage a queue of complex tickets with minimal supervision

  • Adaptable to a defined on-call rotation which may include weekend coverage

  • A clear written communicator: able to explain technical findings and write useful internal documentation

  • Genuinely curious about AI infrastructure, GPU computing, and distributed systems

Must-Haves
  • Solid Linux SysOps experience: Ubuntu Server, RHEL/CentOS, Debian; comfortable with systems, networking, storage, and permissions

  • Proficiency with Docker: container debugging, Docker Compose, image management, cgroup resource limits, Docker storage/filesystem management

  • Experience with virtualization: Proxmox VE, VMware, or similar hypervisors; provisioning and troubleshooting VMs

  • Networking fundamentals: VLAN, DNS, DHCP, NAT, VPN, firewall rules, and general L2/L3 troubleshooting

  • Hands-on experience with NVIDIA GPU drivers, CUDA, and GPU workload troubleshooting (essential)

  • Scripting in Python and Bash for automation and diagnostic tooling

  • Strong English written communication: clear, professional, and technically precise

  • Experience providing technical support in a customer-facing or internal helpdesk context

  • Ability to prioritize across a concurrent queue of escalated tickets, triaging by severity and customer impact, balancing reactive resolution against proactive documentation and tooling work, and making clear judgment calls on when to escalate versus own resolution end-to-end

Nice-to-Haves
  • Familiarity with AI/ML frameworks (TensorFlow, PyTorch) and running GPU-accelerated containers

  • Monitoring and observability experience (Prometheus, Grafana)

  • Relevant certifications: RHCSA, CompTIA Linux+, or similar

  • Knowledge of the Vast.ai platform as a client or infrastructure supplier

Annual Salary Range

$90,000 – $150,000 + equity + benefits

Vast.ai is hiring across all experience levels with compensation commensurate with background, experience and potential.

Benefits
  • Comprehensive health, dental, vision, and life insurance

  • 401(k) with company match

  • Meaningful early-stage equity

  • Onsite meals, snacks, and close collaboration with founders/tech leaders

  • Ambitious, fast-paced startup culture where initiative is rewarded

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Systems Operations Support Engineer — Linux
Systems Operations Support Engineer — Linux

Vast.ai • Los Angeles (CA)

On-site
USD 90,000 - 160,000
Health insurance
401(k) with company match
Meaningful equity
+2
AI, HPC & GPU Infrastructure Support Engineer New
AI, HPC & GPU Infrastructure Support Engineer New

Vast.ai Inc. • Los Angeles (CA), Northern (KY)

Hybrid
USD 90,000 - 160,000
Comprehensive health, dental, vision,
401(k) with company match
Meaningful equity
+2
HPC & GPU Infrastructure Support Engineer
HPC & GPU Infrastructure Support Engineer

Vast.ai • Los Angeles (CA)

On-site
USD 90,000 - 160,000
Health insurance
401(k) with company match
Equity
+2
C++ Software Engineer
C++ Software Engineer

Vast.ai • Los Angeles (CA)

On-site
USD 120,000 - 180,000
Comprehensive health, dental, vision,‑
401(k) with company match
Meaningful equity
+2
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Vast.ai • San Francisco (CA)

On-site
USD 180,000 - 230,000
Health, dental, vision, life insurance
401(k) with company match
Equity
+2
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Vast.ai Inc. • San Francisco (CA)

On-site
USD 180,000 - 230,000
Health benefits
401(k) with company match
Equity
+2
Technical Product Manager, Infrastructure New
Technical Product Manager, Infrastructure New

Vast.ai Inc. • Los Angeles (CA)

On-site
USD 170,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+4
Developer Relations Engineer
Developer Relations Engineer

Vast.ai • Los Angeles (CA)

On-site
USD 160,000 - 200,000
On-site in SF/LA
Health, dental, vision insurance
Travel/conference budget
Systems/GPU Research Engineer
Systems/GPU Research Engineer

Vast.ai Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Comprehensive health, dental, vision, and life insurance
401(k) with company match
Early-stage equity
+2
Developer Relations Engineer
Developer Relations Engineer

Vast.ai • San Francisco (CA)

On-site
USD 160,000 - 200,000
Health insurance
401(k) with match
Onsite meals
+1