Senior Linux Systems Engineer: Virtualization for AI

Jobtailor

California (MO)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Qualifications

  • 5+ years building production systems software, platform infrastructure, virtualization, or Linux-based distributed systems.
  • Experience with virtualization technologies such as KVM, QEMU, VFIO, virtio, Kata Containers, KubeVirt, Firecracker, or gVisor.
  • Experience building or operating Kubernetes platforms and container runtimes at scale.
  • Strong systems programming skills in Go, Rust, C/C++, or a combination thereof.
  • Comfortable debugging production failures spanning hardware, OS, container runtimes, virtualization, and distributed infrastructure.
  • Experience profiling and optimizing system performance using perf, eBPF, ftrace, bpftrace, flame graphs, or crash analysis.

Responsibilities

  • Design secure execution environments using containerd, runc, gVisor, Kata Containers, KubeVirt, and KVM/QEMU.
  • Build GPU-aware runtime infrastructure supporting VFIO, Kata, NVIDIA GPU Operator, and PCIe passthrough for multi-tenant AI workloads.
  • Improve Linux kernel and hypervisor performance through scheduling, memory management, I/O, NUMA locality, and virtualization primitives.
  • Debug complex interactions across Linux, KVM, GPU drivers, firmware, and Kubernetes when workloads don’t behave as expected.
  • Develop observability and debugging tooling using eBPF, perf, tracepoints, and kernel tracing infrastructure.
  • Improve container and VM startup performance, resource isolation, and runtime efficiency for latency-sensitive AI workloads.
  • Extend virtualization infrastructure supporting virtio devices, IOMMU, SR-IOV, mediated devices, nested virtualization, and hardware passthrough.
  • Profile production systems and build performance analysis tooling to identify bottlenecks across kernels, hypervisors, container runtimes, storage, networking, and GPUs.
  • Collaborate with security, platform, networking, and GPU infrastructure teams to define the next generation of runtime isolation and workload execution.

Skills

Linux systems knowledge
High performance systems
Debugging complex infra
Strong programming

Tools

KVM/QEMU
Kata Containers
KubeVirt
gVisor
VFIO
virtio
NVIDIA GPU Operator
Kubernetes platform

Job description

  • Design secure execution environments using containerd, runc, gVisor, Kata Containers, KubeVirt, and KVM/QEMU.
  • Build GPU‑aware runtime infrastructure supporting VFIO, Kata, NVIDIA GPU Operator, and PCIe passthrough for multi‑tenant AI workloads.
  • Improve Linux kernel and hypervisor performance through optimization of scheduling, memory management, I/O, NUMA locality, and virtualization primitives.
  • Debug complex interactions across Linux, KVM, GPU drivers, firmware, and Kubernetes when workloads don't behave as expected.
  • Develop observability and debugging tooling using eBPF, perf, tracepoints, and kernel tracing infrastructure.
  • Improve container and VM startup performance, resource isolation, and runtime efficiency for latency‑sensitive AI inference and training workloads.
  • Extend virtualization infrastructure supporting virtio devices, IOMMU, SR‑IOV, mediated devices, nested virtualization, and hardware passthrough.
  • Profile production systems and build performance analysis tooling to identify bottlenecks across kernels, hypervisors, container runtimes, storage, networking, and GPUs.
  • Collaborate with security, platform, networking, and GPU infrastructure teams to define the next generation of runtime isolation and workload execution.
Requirements
  • 5+ years building production systems software, platform infrastructure, virtualization, or Linux‑based distributed systems.
  • Strong Linux systems knowledge, including namespaces, cgroups, scheduling, memory management, filesystems, networking, and process lifecycle.
  • Experience with virtualization technologies such as KVM, QEMU, VFIO, virtio, Kata Containers, KubeVirt, Firecracker, or gVisor.
  • Experience building or operating Kubernetes platforms and container runtimes at scale.
  • Strong systems programming skills in Go, Rust, C/C++, or a combination thereof.
  • Comfortable debugging production failures that span hardware, operating systems, container runtimes, virtualization, and distributed infrastructure.
  • Experience profiling and optimizing system performance using tools such as perf, eBPF, ftrace, bpftrace, flame graphs, or crash analysis.
Core Competencies

Demonstrates expertise in building and optimizing production systems software and virtualization infrastructure, with a strong focus on Linux systems, container runtimes, and performance analysis. Proficient in debugging complex interactions across hardware and software layers to ensure efficient and secure execution environments.

Highest‑signal resume keywords
  • Linux Systems Knowledge
  • Virtualization Technologies (KVM, QEMU, Kata Containers)
  • Systems Programming (Go, Rust, C/C++)
  • Performance Profiling and Optimization
  • Kubernetes Platform Experience
ATS Optimization Keywords
Hard Skills
  • Linux Kernel Optimization
  • Containerd
  • GVisor
  • KubeVirt
  • GPU‑Aware Runtime Infrastructure
  • EBPF
  • VFIO
  • PCIe Passthrough
  • Debugging Complex Interactions
  • Performance Analysis Tooling
Soft Skills
  • Collaboration
  • Problem‑Solving
Industry Keywords
  • Distributed Systems
  • Virtualization Primitives
  • Resource Isolation
  • Latency‑Sensitive Workloads
  • Multi‑Tenant AI Workloads
Tools & Technologies
  • NVIDIA GPU Operator
  • KVM/QEMU
  • Flame Graphs
  • Crash Analysis
  • Ftrace
  • Bpftrace
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Systems Engineer, Virtualization
Senior Systems Engineer, Virtualization

Jobtailor • California (MO)

On-site
USD 150,000 - 210,000
Virtualization & Orchestration Engineer
Virtualization & Orchestration Engineer

Jobtailor • Bellevue (WA)

On-site
USD 150,000 - 210,000
Lead Software Engineer – Cloud & Software Defined Storage Solutions, C++
Lead Software Engineer – Cloud & Software Defined Storage Solutions, C++

Jobtailor • California (MO)

On-site
USD 150,000 - 190,000
Member of Technical Staff – AI Cloud Infrastructure
Member of Technical Staff – AI Cloud Infrastructure

Jobtailor • California (MO)

On-site
USD 150,000 - 190,000
Principal Product Manager – Hardware Accelerator Virtualization
Principal Product Manager – Hardware Accelerator Virtualization

Jobtailor • Boston (MA)

On-site
USD 150,000 - 230,000
Linux System Developer
Linux System Developer

Jobtailor • Fall River (MA)

On-site
USD 120,000 - 150,000
System Software Engineer – AI
System Software Engineer – AI

Jobtailor • California (MO)

On-site
USD 120,000 - 170,000
Software Engineer – AI & Cloud Engineering
Software Engineer – AI & Cloud Engineering

Jobtailor • Massachusetts

On-site
USD 110,000 - 170,000
Software Engineer – Deployment
Software Engineer – Deployment

Jobtailor • Arlington (VA)

On-site
USD 140,000 - 210,000
Senior Software Engineering Manager – DPU, AI Cloud Infrastructure
Senior Software Engineering Manager – DPU, AI Cloud Infrastructure

Jobtailor • Milpitas (CA)

On-site
USD 180,000 - 240,000