Senior HPC Systems Engineer – Throughput & Reliability

San Diego Stealth Startup

San Diego (CA)

On-site

USD 140,000 - 210,000

Full time

22 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

San Diego Stealth Startup seeks an experienced systems engineer to build and improve a fast, observable, and reliable data compute path across ingestion, processing, storage, and delivery. You will work with algorithms, platforms, and infrastructure engineers to ensure throughput and latency constraints are met in production.

Early focus includes benchmarks, stateful multi-threaded pipelines, and production-ready ML model productization for performance and reliability.

Qualifications

  • PhD (CS/Life Sciences or related) with 3+ years of relevant experience or Master’s with 6+ years or Bachelor’s with 8+ years.
  • Contributed to complex production software with state machines, concurrency, and I/O bottlenecks.
  • Shipped production-quality software in C++, Rust, CUDA, C, or C#; familiar with profiling, tracing, debugging, testing, builds, and CI.
  • Able to reason about throughput, latency, memory, storage, network behavior, and error recovery.
  • Experience productizing ML models (neural nets, tree-based, or unsupervised) for production reliability.
  • Works well in a flat, technical team and communicates tradeoffs clearly.

Responsibilities

  • Build and improve a high-throughput compute stack that is fast, observable, recoverable, and operable in production.
  • Establish reproducible hardware benchmarks for accelerated compute, memory transfers, storage throughput, and network streaming.
  • Productize ML models so they meet production requirements for performance, reliability, observability, and quality.
  • Define backpressure, checkpointing, retry, and recovery behavior for disk pressure and network outages.
  • Ensure production interfaces and tests remain durable against future algorithm changes.
  • Collaborate with algorithms, platform, and infrastructure engineers to improve reliability of data and compute paths.

Skills

C++
Rust
CUDA
C
C#
Profiling
Tracing
Debugging
CI

Education

PhD in Computer Science or related field
Master’s degree with 6+ years of experience
Bachelor’s degree with 8+ years of experience

Tools

CI/CD tooling
Profiling tools
Debugging tools

Job description

San Diego Stealth Startup seeks an experienced systems engineer to build and improve a fast, observable, and reliable data compute path across ingestion, processing, storage, and delivery. You will work with algorithms, platforms, and infrastructure engineers to ensure throughput and latency constraints are met in production.

Early focus includes benchmarks, stateful multi-threaded pipelines, and production-ready ML model productization for performance and reliability.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff HPC Software Engineer
Staff HPC Software Engineer

San Diego Stealth Startup • San Diego (CA)

On-site
USD 140,000 - 210,000
Senior Bioinformatics Engineer - Scalable Data Pipelines
Senior Bioinformatics Engineer - Scalable Data Pipelines

San Diego Stealth Startup • San Diego (CA)

On-site
USD 202,000 - 215,000
Senior Software Engineer - HPC Distributed Systems
Senior Software Engineer - HPC Distributed Systems

Adtechtalent • Washington

On-site
USD 140,000 - 230,000
Healthcare
Retirement plans
Disability coverage
+7
Senior Backend Performance Engineer — GPU-Accelerated Data Pipelines
Senior Backend Performance Engineer — GPU-Accelerated Data Pipelines

Three Pillars Recruiting • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Inference Systems Engineer for High-Performance AI
Senior Inference Systems Engineer for High-Performance AI

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Senior Infra & DevOps Engineer — Hardware/Robotics
Senior Infra & DevOps Engineer — Hardware/Robotics

building3 • San Francisco (CA)

On-site
USD 140,000 - 200,000
Senior Data Engineer, Real-Time Cost & Analytics Platform
Senior Data Engineer, Real-Time Cost & Analytics Platform

fal • San Francisco (CA)

On-site
USD 140,000 - 210,000
Interesting work
Learning & growth
Relocation assistance
+1
Lead HPC Systems & Performance Engineer
Lead HPC Systems & Performance Engineer

Stanford Black Limited • Dallas (TX)

On-site
USD 140,000 - 210,000
Senior HPC Engineer — Platform Mastery & Customer Impact
Senior HPC Engineer — Platform Mastery & Customer Impact

Quiet Capital • United States

On-site
USD 100,000 - 130,000
Dallas-Based Lead HPC Systems & Performance Engineer
Dallas-Based Lead HPC Systems & Performance Engineer

Stanford Black Limited • Dallas (TX)

On-site
USD 140,000 - 210,000