ML Inference Engineer San Francisco · Engineering · Full Time →

Reactor

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Visa sponsorship
Relocation support
Generous health, dental, and vision coverage

Job summary

Reactor is looking for an experienced ML Inference Engineer with deep expertise in high-performance ML engineering. This role focuses on optimizing the performance of generative media models, contributing to Reactor's competitive edge.

The ideal candidate will drive model performance for diffusion models, collaborate with partner teams, and implement optimizations. A Bachelor's degree in a related field is required, along with strong knowledge of GPU hardware and modern ML optimization techniques. The position offers competitive salary and meaningful equity.

Qualifications

  • Strong foundation in systems programming, with a track record of identifying and resolving bottlenecks.
  • Working knowledge of GPU hardware (NVIDIA).
  • Strong understanding of transformer architectures and modern ML optimization techniques.

Responsibilities

  • Drive frontier position on model performance for diffusion models.
  • Design and implement a in-house inference runtime.
  • Profile and benchmark model performance to identify computational bottlenecks.

Skills

Deep expertise in PyTorch
TensorRT
Advanced serving architectures
Model compilation
Quantization

Education

Bachelor's degree in Computer Science, Electrical Engineering, or related

Tools

Nsight
ONNX Runtime

Job description

We're looking for an ML Inference Engineer with deep expertise in high-performance ML engineering. This is a highly technical, high-impact role focused on squeezing every drop of performance from generative media models.

You'll work across the inference stack, designing novel frameworks, optimizing inference performance, and shaping Reactor's competitive edge in ultra-low-latency, high-throughput environments.

What You'll Do
  • Drive our frontier position on model performance for diffusion models
  • Design and implement a high-performance in-house inference runtime
  • Implement optimizations using torch.compile, custom CUDA kernels, and specialized inference frameworks
  • Optimize neural network models through quantization, pruning, and architectural modifications
  • Profile and benchmark model performance to identify computational bottlenecks
  • Collaborate directly with model partner teams to integrate their models into our platform
What We're Looking For
  • Bachelor's degree in Computer Science, Electrical Engineering, or a related technical field (or equivalent practical experience)
  • Strong foundation in systems programming, with a track record of identifying and resolving bottlenecks
  • Deep expertise in PyTorch, TensorRT, TransformerEngine, Nsight, ONNX Runtime
  • Model compilation, quantization (INT8/FP16), and advanced serving architectures
  • Working knowledge of GPU hardware (NVIDIA)
  • Strong understanding of transformer architectures and modern ML optimization techniques
  • Competitive SF salary and meaningful early equity
  • Visa sponsorship and relocation support
  • Generous health, dental, and vision coverage
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding Engineer, ML Inference
Founding Engineer, ML Inference

Reactor • San Francisco (CA)

On-site
USD 180,000 - 280,000
Competitive salary
Early equity
Health, dental, and vision coverage
+1
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
High-Performance ML Inference Engineer for Diffusion Models
High-Performance ML Inference Engineer for Diffusion Models

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Visa sponsorship
Relocation support
Generous health, dental, and vision coverage
Senior ML Performance Engineer
Senior ML Performance Engineer

well-funded deeptech startup • California (MO)

On-site
USD 200,000 - 250,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
Senior ML Performance Engineer - GPU & Inference
Senior ML Performance Engineer - GPU & Inference

Modal Labs • New York (NY)

On-site
Principal ML Infrastructure Engineer (Relocation Available)
Principal ML Infrastructure Engineer (Relocation Available)

Franklin Fitch • Dallas (TX)

On-site
USD 100,000 - 140,000
Staff Software Engineer, ML Performance & Systems
Staff Software Engineer, ML Performance & Systems

fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Competitive salary and equity
Visa sponsorship
Health, dental, and vision insurance
+1