Founding Engineer, ML Inference

Reactor

San Francisco (CA)

On-site

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary
Early equity
Health, dental, and vision coverage
Relocation support

Job summary

A media technology company in San Francisco is seeking a Founding Engineer specializing in ML Inference. This highly technical role requires expertise in the ML infrastructure stack and aims to optimize generative media performance. The ideal candidate will drive innovations in real-time model performance, design in-house inference runtimes, and optimize models through advanced techniques. Competitive salary and relocation support are offered, along with generous health coverage.

Qualifications

  • Strong foundation in systems programming with a track record of identifying and resolving bottlenecks.
  • Deep expertise in ML infrastructure stack, including PyTorch, TensorRT, and more.
  • Knowledge of model compilation and quantization techniques.

Responsibilities

  • Drive performance for real-time model performance for diffusion models.
  • Design a high-performance in-house inference runtime.
  • Optimize models for inference through quantization and pruning.
  • Benchmark model performance to identify bottlenecks.

Skills

Systems programming
PyTorch
TensorRT
Model compilation
Quantization
NVIDIA GPU hardware knowledge
Transformer architectures

Job description

We're building a future where anyone can create interactive media applications that delight, educate, and simulate. Building a new kind of platform for real-time generative media, enabling developers to go from idea to immersive, dynamic experience in seconds. Join a small, focused team of YC and unicorn founders and senior engineers with deep expertise in 3D, generative video, developer platforms, and creative tool, aspiring to continuously push the boundaries of what's possible.

About the Role

We're looking for a Founding Engineer, ML Inference with deep expertise in high-performance ML engineering. This is a highly technical, high-impact role focused on squeezing every drop of performance from generative media models.

You'll work across the model-serving stack, designing novel inference frameworks, optimizing inference performance, and shaping the competitive edge in ultra-low-latency, high-throughput environments.

What You'll Do
  • Drive our frontier position on real-time model performance for diffusion models
  • Design and implement a high-performance in-house inference runtime
  • Implement optimizations using torch.compile, custom CUDA kernels, and specialized inference frameworks
  • Optimize neural network models for inference through quantization, pruning, and architectural modifications while maintaining accuracy
  • Profile and benchmark model performance to identify computational bottlenecks
  • Collaborate directly with model partner teams to directly integrate their models into our platform
Required Skills
  • Strong foundation in systems programming, with a track record of identifying and resolving bottlenecks
  • Deep expertise in the ML infrastructure stack: PyTorch, TensorRT, TransformerEngine, Nsight, ONNX Runtime
  • Model compilation, quantization (INT8/FP16), and advanced serving architectures
  • Working knowledge of GPU hardware (NVIDIA) and the ability to dive deep into the stack as needed
  • Strong understanding of transformer architectures and modern ML model optimization techniques
Logistics

We are based in-person in San Francisco. We believe the best ideas and work come from being together.

• Competitive San Francisco salary and meaningful early equity.

• We sponsor visas. We are committed to working through the process together for the right candidates. If you're currently outside the US, we're also committed to helping you relocate to the US throughout this process.

• We offer generous health, dental, and vision coverage, and relocation support as needed.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Inference Engineer San Francisco · Engineering · Full Time →
ML Inference Engineer San Francisco · Engineering · Full Time →

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Visa sponsorship
Relocation support
Generous health, dental, and vision coverage
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Staff / Principal Machine Learning Engineer, Serving
Staff / Principal Machine Learning Engineer, Serving

Inworld AI • Mountain View (CA)

On-site
USD 270,000 - 500,000
Relocation assistance
Equity options
Comprehensive benefits package
Staff Software Engineer, ML Performance & Systems
Staff Software Engineer, ML Performance & Systems

fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Competitive salary and equity
Visa sponsorship
Health, dental, and vision insurance
+1
Founding ML Inference Engineer — Ultra-Low Latency AI
Founding ML Inference Engineer — Ultra-Low Latency AI

Reactor • San Francisco (CA)

On-site
USD 60,000 - 80,000
Senior Inference Engineer
Senior Inference Engineer

DeepRec.ai • Palo Alto (CA)

Hybrid
USD 180,000 - 240,000
Competitive salary
Equity
Health benefits
+3
Research Engineer, Infrastructure, Inference
Research Engineer, Infrastructure, Inference

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Staff Technical Lead for Inference & ML Performance
Staff Technical Lead for Inference & ML Performance

fal • San Francisco (CA)

On-site
USD 150,000 - 200,000
Machine Learning Researcher
Machine Learning Researcher

SOLANA FOUNDATION • San Francisco (CA)

Hybrid
USD 250,000 - 350,000
Equity in a high-growth startup
Comprehensive benefits