Staff HPC Software Engineer

San Diego Stealth Startup

San Diego (CA)

On-site

USD 140,000 - 210,000

Full time

17 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

San Diego Stealth Startup seeks an experienced systems engineer to build and improve a fast, observable, and reliable data compute path across ingestion, processing, storage, and delivery. You will work with algorithms, platforms, and infrastructure engineers to ensure throughput and latency constraints are met in production.

Early focus includes benchmarks, stateful multi-threaded pipelines, and production-ready ML model productization for performance and reliability.

Qualifications

  • PhD (CS/Life Sciences or related) with 3+ years of relevant experience or Master’s with 6+ years or Bachelor’s with 8+ years.
  • Contributed to complex production software with state machines, concurrency, and I/O bottlenecks.
  • Shipped production-quality software in C++, Rust, CUDA, C, or C#; familiar with profiling, tracing, debugging, testing, builds, and CI.
  • Able to reason about throughput, latency, memory, storage, network behavior, and error recovery.
  • Experience productizing ML models (neural nets, tree-based, or unsupervised) for production reliability.
  • Works well in a flat, technical team and communicates tradeoffs clearly.

Responsibilities

  • Build and improve a high-throughput compute stack that is fast, observable, recoverable, and operable in production.
  • Establish reproducible hardware benchmarks for accelerated compute, memory transfers, storage throughput, and network streaming.
  • Productize ML models so they meet production requirements for performance, reliability, observability, and quality.
  • Define backpressure, checkpointing, retry, and recovery behavior for disk pressure and network outages.
  • Ensure production interfaces and tests remain durable against future algorithm changes.
  • Collaborate with algorithms, platform, and infrastructure engineers to improve reliability of data and compute paths.

Skills

C++
Rust
CUDA
C
C#
Profiling
Tracing
Debugging
CI

Education

PhD in Computer Science or related field
Master’s degree with 6+ years of experience
Bachelor’s degree with 8+ years of experience

Tools

CI/CD tooling
Profiling tools
Debugging tools

Job description

We are looking for an experienced systems engineer to build and improve a measured, reliablecomputearchitecture. The work spans high-throughput data ingestion, processing, storage, and delivery.

Build and improve a high-throughputcomputestack,so it is fast, observable, recoverable, and practical tooperatein production environments.

This person will work with algorithms,platforms, and infrastructure engineers. They will not be expected to own every algorithm or infrastructure service. Their core responsibility is making the data and compute path reliable under real throughput, storage, network, and latency constraints.

Early work

In the first three to six months, this person should help:

  • Establish reproducible hardware benchmarks for acceleratedcompute, CPU workloads, memory transfers, storage throughput, and network streaming.
  • Build or harden stateful, multi-threaded pipelines that move data from ingestion through compute and output.
  • Productize machine learning models, including neural networks, tree-based models, and unsupervised models, so they meet production requirements for performance, reliability, observability, and quality.
  • Define backpressure, checkpointing, retry, and recovery behavior for disk pressure, slow consumers, and network outages.
  • Compare alternative processing designs using wall-clock time, memory, storage, and quality measurements.
  • Make the production interfaces and performance tests durable enough that later algorithm changes do not quietly break throughput or recovery behavior.

Required experience

  • This role requires a PhD in Computer Science, Life Sciences, or a related discipline with 3+ years of relevant experience; a master's degree with 6+ years of relevant experience; or abachelor'sdegree with 8+ years of relevant experience.
  • Has contributed to a complex production software system with state machines, concurrency, and realcomputeor I/O bottlenecks. They do not need to have been the technicallead butmust understand how these systems fail and how to debug them.
  • Has shipped production-quality software in at least one of C++, Rust, CUDA, C, or C#. Comfortable with the normal engineering tools: profiling, tracing, debugging, testing, code review, builds, and CI.
  • Can reason concretely about throughput, latency, buffering, memory, storage, network behavior, scheduling, contention, and failure recovery.
  • Uses measurements to guide performance work: canidentifya bottleneck, make a targeted change, quantify the gain, and add a regression guard.
  • Has experience productizing machine learning models, including neural networks, tree-based models, or unsupervised models. Can make these models reliable, measurable, and efficient in a production system. This is not a model-research role.
  • Works well in a flat, highly technical team: canstatetradeoffs clearly, contribute outside a narrow specialty, learn from others, and strengthen areas where the team is currently thin.

Strongly preferred

  • GPU computing, CUDA profiling, or heterogeneous CPU/GPU pipelines.
  • High-throughput storage, networking, streaming, or low-latency systems.
  • Performance-sensitive scientific computing or another data-intensive system where delayed or failed processing has operational consequences.
  • Experience with resilient data pipelines: bounded queues, backpressure, checkpoint/restart, idempotent outputs, and operational telemetry.
  • This is not primarily anMLOps, cloud-platform, data-science, or model-training position.
  • This role complements algorithm and scientific development; it does not unilaterally set scientific requirements, quality criteria, or cloud/platform ownership.
  • The near-term focus is the high-throughputcomputepath and its interfaces, not a general rewrite of company infrastructure.

We are an equal opportunity employer. We thrive on diversity and collaboration.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Systems Engineer – Throughput & Reliability
Senior HPC Systems Engineer – Throughput & Reliability

San Diego Stealth Startup • San Diego (CA)

On-site
USD 140,000 - 210,000
HPC AI Systems Administrator
HPC AI Systems Administrator

MRE Consulting • Houston (TX)

On-site
USD 95,000 - 140,000
Lead HPC Systems & Performance Engineer
Lead HPC Systems & Performance Engineer

Stanford Black Limited • Dallas (TX)

On-site
USD 140,000 - 210,000
High-Performance Computing Talent Community
High-Performance Computing Talent Community

CGG Services SAS • Houston (TX)

On-site
USD 90,000 - 130,000
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000
Senior HPC Applications Engineer
Senior HPC Applications Engineer

Parallel Works • Chicago (IL)

On-site
USD 140,000 - 190,000
Medical coverage
Vision coverage
Dental coverage
+2
Senior HPC Applications Engineer
Senior HPC Applications Engineer

Parallel Works, Inc. • Chicago (IL)

On-site
USD 140,000 - 210,000
Medical, vision, and dental coverage
401(k) with company match
Paid vacation and sick time
+1
Senior Software Engineer - Backend Performance -MarTech/AdTech
Senior Software Engineer - Backend Performance -MarTech/AdTech

Three Pillars Recruiting • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Backend Software Engineer
Senior Backend Software Engineer

Carex Consulting Group • Boston (MA)

On-site
USD 120,000 - 180,000
Staff High-Performance Software Engineer
Staff High-Performance Software Engineer

AtlasBase • South San Francisco (CA)

On-site
USD 150,000 - 230,000