Software Engineer, High Performance Computing

Eventual

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

Just now
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

In-person 4 days/week SF office
Competitive compensation with startup‑
Equity
Catered lunches and dinners
Commuter benefit
Team-building events and poker nights
Health, vision, and dental coverage
Flexible PTO
Latest Apple equipment
401(k) plan with match

Job summary

Eventual in San Francisco is seeking a Systems Engineer for the Dataloading team to craft the layer that turns multi-petabyte video corpora into dict[str, Tensor] on GPUs at line rate. You will work with top labs and labs with Vera Rubin on the horizon.

We will uplevel you on NVL72, CUDA, and SLURM, with a focus on memory, NVMe, network, and CPU interactions to keep GPUs fed and efficient.

Qualifications

  • Experience designing high-throughput data loaders and pipelines.
  • Strong systems-level understanding of memory, IO, and CPU/NVMe interactions.
  • Proficiency in Rust, C++, or C and performance tuning for GPU workloads.

Responsibilities

  • Design and build the video-native dataloader delivering tensors to the GPU.
  • Profile and optimize the data path from object store to device RAM, minimizing copies.
  • Saturate current hardware (B200, GB200, NVL72) on real training jobs.
  • Own performance benchmarks against baselines and track regressions in PRs.
  • Collaborate with partner labs to deploy the loader in training stacks.
  • Coordinate with Storage and Visual Understanding teams on data flow and ingestion.

Skills

Rust
C++
C
io_uring
NUMA
Performance tuning

Tools

CUDA
SLURM
Kubernetes

Job description

About Eventual

Every breakthrough Physical AI system — humanoid robots, autonomous vehicles, video generation models — is trained on petabytes of video, lidar, radar, and sensor data. But today's data platforms (Databricks, Snowflake) were built for spreadsheet-like analytics, not the multimodal corpora that power AI. Robotics and video‑AI teams now lose 20-40% of their training time to dataloading alone. GPU bandwidth has grown 2-3× per generation. Storage and pipelines haven't. The gap widens every year.

Eventual was founded in 2022 to close it. Our open‑source engine, Daft, is the distributed data engine purpose‑built for multimodal AI — already running 2 PB/day at Amazon, 60‑100 PB at another FAANG company, and in production at Mobileye, TogetherAI, and CloudKitchens. We are building a video‑native index on top of our engine for Physical AI that streams curated datasets to GPUs at line rate. Saturates B200s today. Aimed at NVL72 and Vera Rubin tomorrow.

We're building this in partnership with the top PhysicalAI labs and public AI infrastructure companies today. We have raised $30M from Felicis, CRV, Microsoft M12, Citi, Essence, Y Combinator, Caffeinated Capital, Array.vc, and angels from the co‑founders of Databricks and Perplexity. We've assembled a world‑class team from AWS, Render, Pinecone and Tesla. We have spent our careers powering the last generation of PhysicalAI in self‑driving, and are excited to now do this for the next.

Join our small (but powerful!) team working together 4 days/week in our SF Mission district office.

Your Role

As a Systems Engineer on the Dataloading team, you'll build the layer that turns multi‑petabyte video corpora into dict[str, Tensor] already on the GPU at line rate. We work with the top labs training Physical AI on the newest generation hardware — H100, B200, GB200, NVL72, with Vera Rubin on the horizon — on billions of dollars worth of compute, in collaboration with partners that are the largest public AI companies on Earth. Our job is to keep those GPUs fed: rank‑aware sampling, NVMe caching, video and sensor co‑loading, random access into clips, decode pipelining. Streaming alone can already saturate a B200; the hard part is enabling the complex sampling patterns researchers actually need without giving up a single percentage point of MFU.

This is a systems engineering role for someone who feels physical pain when a system is slow. You won't need GPU experience on day one — we'll uplevel you on NVL72, CUDA, and SLURM. We will need you to bring real expertise on what happens between NVMe, network, memory, and CPU, and a deep instinct for where bytes go.

Key Responsibilities
  • Design and build the video‑native dataloader: rank‑aware, NVMe‑cached, random‑access into clips, returns tensors directly to the GPU.
  • Profile and optimize the full data path from object store → NVMe → page cache → host RAM → device RAM. Eliminate every avoidable copy and stall.
  • Saturate the latest hardware (B200, GB200, NVL72) on real customer training jobs. Push toward Vera Rubin bandwidth requirements.
  • Own performance benchmarks against customer baselines (custom DataLoaders, DALI, decord, LeRobot) and against our own historical numbers — regressions get caught at PR time.
  • Partner with researchers at our partner labs to land the loader in their training stack and measure MFU end‑to‑end.
  • Work cross‑team with Storage Infrastructure on the index/format boundary and with Visual Understanding on the model‑output ingestion path.
What We Look For
  • Obsession with systems‑level performance. You can recite Jeff Dean's "numbers every programmer should know" in your sleep. You eat flamegraphs for breakfast.
  • Strong opinions on io_uring — love it or hate it, you've earned the opinion.
  • Live and breathe Rust, C++, or C. You reach for them when it matters and you know why.
  • Strong familiarity with operating systems — page cache, scheduling, syscalls, NUMA, memory hierarchies.
  • A sense for where bytes actually go: NVMe vs. memory vs. network vs. PCIe vs. NVLink, and the throughput and latency budgets of each.
Nice to have
  • Experience working with GPUs is a plus, but you don't need it on day one.
  • Experience working with SLURM, Kubernetes for GPU workloads, or other HPC schedulers.
  • Hands‑on CUDA experience.
  • Deep expertise on memory and caching subsystems — page cache tuning, hugepages, NUMA pinning, GPU-Direct Storage.
  • Worked on video decode pipelines (PyAV, decord, NVDEC) or PyTorch DataLoader internals.
  • Contributed to open‑source systems projects in Rust/C++.
Perks & Benefits
  • In‑person, tight‑knit team — 4 days/week in our SF Mission office.
  • Competitive comp and meaningful startup equity.
  • Catered lunches and dinners for SF employees.
  • Commuter benefit.
  • Team‑building events and poker nights.
  • Health, vision, and dental coverage.
  • Flexible PTO.
  • Latest Apple equipment.
  • 401(k) plan with match.

If slow systems evoke emotional pain for you and you want to spend the next few years making the most expensive GPU clusters on the planet earn their keep, we'd love to talk.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, High Performance Computing
Software Engineer, High Performance Computing

Eventual • United States

On-site
USD 100,000 - 130,000
Catered lunches and dinners
Flexible PTO
Health, vision, and dental coverage
+1
Dataloading Systems Engineer — GPU-Scale Data Pipelines
Dataloading Systems Engineer — GPU-Scale Data Pipelines

Eventual • United States

On-site
Software Engineer, Product
Software Engineer, Product

Alumni Ventures • San Francisco (CA)

On-site
USD 150,000 - 250,000
Competitive compensation and startup equity
Catered lunches and dinners
Commuter benefit
+3
Software Engineer, Large-Scale Data Query Systems
Software Engineer, Large-Scale Data Query Systems

Doist • San Francisco (CA)

On-site
USD 180,000 - 240,000
In-person 4 days/week in office
Competitive comp and startup equity
Catered lunches and dinners for SF
+6
Research Engineer, Multimodal Data
Research Engineer, Multimodal Data

Eventual • San Francisco (CA)

On-site
USD 120,000 - 150,000
Catered lunches and dinners
Commuter benefit
Health, vision, and dental coverage
+2
Software Engineer, Large-Scale Data Query Systems
Software Engineer, Large-Scale Data Query Systems

Eventual • San Francisco (CA)

On-site
USD 150,000 - 250,000
Competitive pay
Startup equity
Catered meals and dinners
+6
Research Engineer, Multimodal Data
Research Engineer, Multimodal Data

Eventual • United States

On-site
USD 150,000 - 210,000
In-person, tight-knit SF team
Competitive compensation
Startup equity
+7
Member of Technical Staff, Specialized Focus
Member of Technical Staff, Specialized Focus

Mixpeek • San Francisco (CA)

On-site
USD 180,000 - 280,000
In-person tight knit team with 4x a-?e
Competitive comp and startup equity
Catered lunches and dinners for SF
+6
Software Engineer, Systems
Software Engineer, Systems

Eventual • California (MO)

On-site
USD 100,000 - 150,000
Catered lunches and dinners for SF employees
Commuter benefit
Team building events & poker nights
+4
Systems/GPU Research Engineer
Systems/GPU Research Engineer

Vast.ai Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Comprehensive health, dental, vision, and life insurance
401(k) with company match
Early-stage equity
+2