- Design secure execution environments using containerd, runc, gVisor, Kata Containers, KubeVirt, and KVM/QEMU.
- Build GPU‑aware runtime infrastructure supporting VFIO, Kata, NVIDIA GPU Operator, and PCIe passthrough for multi‑tenant AI workloads.
- Improve Linux kernel and hypervisor performance through optimization of scheduling, memory management, I/O, NUMA locality, and virtualization primitives.
- Debug complex interactions across Linux, KVM, GPU drivers, firmware, and Kubernetes when workloads don't behave as expected.
- Develop observability and debugging tooling using eBPF, perf, tracepoints, and kernel tracing infrastructure.
- Improve container and VM startup performance, resource isolation, and runtime efficiency for latency‑sensitive AI inference and training workloads.
- Extend virtualization infrastructure supporting virtio devices, IOMMU, SR‑IOV, mediated devices, nested virtualization, and hardware passthrough.
- Profile production systems and build performance analysis tooling to identify bottlenecks across kernels, hypervisors, container runtimes, storage, networking, and GPUs.
- Collaborate with security, platform, networking, and GPU infrastructure teams to define the next generation of runtime isolation and workload execution.
Requirements
- 5+ years building production systems software, platform infrastructure, virtualization, or Linux‑based distributed systems.
- Strong Linux systems knowledge, including namespaces, cgroups, scheduling, memory management, filesystems, networking, and process lifecycle.
- Experience with virtualization technologies such as KVM, QEMU, VFIO, virtio, Kata Containers, KubeVirt, Firecracker, or gVisor.
- Experience building or operating Kubernetes platforms and container runtimes at scale.
- Strong systems programming skills in Go, Rust, C/C++, or a combination thereof.
- Comfortable debugging production failures that span hardware, operating systems, container runtimes, virtualization, and distributed infrastructure.
- Experience profiling and optimizing system performance using tools such as perf, eBPF, ftrace, bpftrace, flame graphs, or crash analysis.
Core Competencies
Demonstrates expertise in building and optimizing production systems software and virtualization infrastructure, with a strong focus on Linux systems, container runtimes, and performance analysis. Proficient in debugging complex interactions across hardware and software layers to ensure efficient and secure execution environments.
Highest‑signal resume keywords
- Linux Systems Knowledge
- Virtualization Technologies (KVM, QEMU, Kata Containers)
- Systems Programming (Go, Rust, C/C++)
- Performance Profiling and Optimization
- Kubernetes Platform Experience
ATS Optimization Keywords
Hard Skills
- Linux Kernel Optimization
- Containerd
- GVisor
- KubeVirt
- GPU‑Aware Runtime Infrastructure
- EBPF
- VFIO
- PCIe Passthrough
- Debugging Complex Interactions
- Performance Analysis Tooling
Soft Skills
- Collaboration
- Problem‑Solving
Industry Keywords
- Distributed Systems
- Virtualization Primitives
- Resource Isolation
- Latency‑Sensitive Workloads
- Multi‑Tenant AI Workloads
Tools & Technologies
- NVIDIA GPU Operator
- KVM/QEMU
- Flame Graphs
- Crash Analysis
- Ftrace
- Bpftrace