Senior Staff Machine Learning Software Engineer

San Diego Stealth Startup

San Diego (CA)

On-site

USD 202,000 - 215,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

San Diego Stealth Startup in San Diego, CA is building a product where learned models and compute-heavy inference must run inside a tight local runtime budget. This role owns the path from a working prototype to production inference that is measured, packaged, tested, and ready for repeated use in the field.

You will work closely with the people developing the underlying algorithms, but your ownership is different: production readiness, performance, reliability, and the engineering boundary

Qualifications

  • Shipped constrained inference within a target latency, memory, or compute budget.
  • Produced production-grade Rust or modern C++ code with attention to correctness and performance.
  • Experience with CUDA or accelerator-aware execution and memory movement.
  • Demonstrated ownership and leadership on complex projects.

Responsibilities

  • Turning research prototypes into production inference components with explicit latency, throughput, memory, and accuracy budgets.
  • Optimizing the execution path: tensor layout, host/device transfers, batching strategy, kernel launch overhead, mixed precision, quantization, memory reuse.
  • Writing or tuning Rust, C++, and CUDA where needed and validating improvements with profiler output and tests.
  • Building inference-adjacent evaluation machinery: calibration checks, confidence behavior, regression detection tied to product metrics.
  • Maintaining the deployment contract: model artifacts, runtime integration, versioning, reproducibility, and performance gates.

Skills

Rust
C++
CUDA
Performance optimization
Profiling
Distributed systems

Education

PhD (6+ yrs) or MS (10+ yrs) or BS/BA (12+ yrs) in life sciences or tech

Tools

Rust
C++
CUDA

Job description

Position Overview

We are building a product where learned models and compute-heavy inference components have to run inside a tight local runtime budget. Research code is only the starting point. This role owns the path from a working prototype to production inference that is measured, packaged, tested, and ready for repeated use in the field.

Location: San Diego, CA

Job Type: Full-Time

Salary Range: $202,000 – 215,000

Position Overview

We are building a product where learned models and compute-heavy inference components have to run inside a tight local runtime budget. Research code is only the starting point. This role owns the path from a working prototype to production inference that is measured, packaged, tested, and ready for repeated use in the field.

You will work closely with the people developing the underlying algorithms, but your ownership is different: production readiness, performance, reliability, and the engineering boundary between exploratory model work and shipped execution. The strongest fit is someone who can explain the bottleneck they found, the number they moved, the tradeoff they accepted, and the test that kept the fix from regressing.

If your best work is making inference faster, smaller, more predictable, and easier to ship, this role is likely a good match.

Responsibilities

  • Turning research prototypes into production inference components with explicit latency, throughput, memory, and accuracy budgets
  • Optimizing the execution path: tensor layout, host/device transfers, batching strategy, kernel launch overhead, mixed precision, quantization, and memory reuse
  • Writing or tuning Rust, C++, and CUDA where framework-level optimization is not enough, then validating the improvement with profiler output and release-facing tests
  • Building inference-adjacent evaluation machinery: calibration checks, confidence behavior, regression detection, dataset slices, and failure-mode reporting tied to product metrics
  • Maintaining the deployment contract: model artifacts, runtime integration, versioning, reproducibility, and performance gates that block unsafe changes
  • Algorithm research and novel model design live on a separate track. You will collaborate with that team, translate prototypes into production constraints, and surface shipping risks early when a design needs to change.

Qualifications

  • PhD (6+ years), MS (10+ years) or BS/BA (12+ years) of experience in life sciences or technology.
  • Must have demonstrated leadership or ownership with 2 of the 5 areas referenced below successfully:
  • Shipped constrained inference. You have personally moved a model or learned component from prototype to deployed runtime with a real latency, throughput, memory, or power budget. You can name the target, the bottleneck, and the change that closed the gap.
  • Rust/C++ at shipping depth. You have written production code in Rust or modern C++ where correctness, latency, memory layout, and ownership boundaries mattered. You can reason about the runtime behavior of the code you ship, not just its API surface.
  • CUDA and accelerator-aware execution. You are comfortable below Python: custom CUDA extensions or kernels, host/device memory movement, launch overhead, profiler traces, and the practical tradeoffs between framework convenience and a purpose-built implementation.
  • Performance-native judgment. You reason in wall-clock time, memory movement, launch overhead, bandwidth, numerical precision, and error budgets without needing those constraints added late in review.
  • Production engineering discipline. You define typed interfaces, deterministic behavior, reproducible artifacts, meaningful tests, and clean handoffs with upstream research code.

Strongly Preferred

  • Rust at shipping depth, especially FFI boundaries, pyo3 / maturin, async runtimes, or performance-sensitive service code
  • Inference on constrained local hardware, embedded systems, edge devices, or budget-bound accelerator deployments
  • Quantization, mixed precision, model compression, or kernel fusion that shipped beyond a benchmark notebook
  • Calibration or confidence estimation used on production outputs, with monitoring or regression checks attached
  • Public or shareable evidence of engineering quality: code, technical writing, postmortems, talks, or a concrete shipped system you can discuss
  • Comfort using AI-assisted development tools while still owning correctness, tests, and review quality

Nice to Have

  • Real-time or near-real-time signal-processing systems
  • Products that combine learned models with deterministic numerical code
  • Rust- or C++-based inference or numerical pipelines, including custom FFI to CUDA, cuDNN, TensorRT, or similar accelerator libraries

We are an equal opportunity employer. We thrive on diversity and collaboration.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
INFERENCE ENGINEER
INFERENCE ENGINEER

MakerMaker • San Francisco (CA)

On-site
USD 180,000 - 240,000
Distributed Systems Engineer, Real-Time Inference at Scale
Distributed Systems Engineer, Real-Time Inference at Scale

adaption • San Francisco (CA)

On-site
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

B Capital • United States

On-site
USD 180,000 - 230,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

United States Digital Space LLC • San Francisco (CA)

On-site
USD 180,000 - 260,000
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Bala Cynwyd (PA)

On-site
USD 120,000 - 160,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
INFERENCE ENGINEER
INFERENCE ENGINEER

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000