Member of Technical Staff, Inference

Reactor

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Competitive SF salary
Early equity
Visa sponsorship
Relocation support
Health, dental, and vision coverage

Job summary

Reactor in San Francisco is seeking an ML Inference Engineer to maximize performance of generative media models and push ultra-low-latency, high-throughput inference. You will craft an in-house runtime, implement optimizations with PyTorch tools, and collaborate with partner teams to integrate external models.

The role requires deep expertise in PyTorch, TensorRT, CUDA, and model optimization techniques, with a focus on delivering cutting-edge inference capabilities at scale.

Qualifications

  • Bachelor's degree in Computer Science, Electrical Engineering, or equivalent practical experience.
  • Strong foundation in systems programming and bottleneck resolution.
  • Expertise in PyTorch, TensorRT, TransformerEngine, Nsight, ONNX Runtime.

Responsibilities

  • Drive model performance for diffusion models and high-throughput inference.
  • Design and implement a high-performance in-house inference runtime.
  • Apply optimizations using torch.compile, CUDA kernels, and specialized inference frameworks.
  • Quantize, prune, and modify architectures to optimize neural networks.
  • Profile and benchmark model performance to identify bottlenecks.
  • Collaborate with model partner teams to integrate models into the platform.

Skills

PyTorch
TensorRT
CUDA
Model optimization
Quantization
Low-latency inference
GPU programming
Transformer architectures

Education

Bachelor's degree in CS/EE

Tools

Nsight
ONNX Runtime
TransformerEngine

Job description

Description

We're looking for an ML Inference Engineer with deep expertise in high-performance ML engineering. This is a highly technical, high-impact role focused on squeezing every drop of performance from generative media models.

You'll work across the inference stack, designing novel frameworks, optimizing inference performance, and shaping Reactor's competitive edge in ultra-low-latency, high-throughput environments.

Department

Engineering

Location

San Francisco

What You'll Do
  • Drive our frontier position on model performance for diffusion models
  • Design and implement a high-performance in-house inference runtime
  • Implement optimizations using torch.compile, custom CUDA kernels, and specialized inference frameworks
  • Optimize neural network models through quantization, pruning, and architectural modifications
  • Profile and benchmark model performance to identify computational bottlenecks
  • Collaborate directly with model partner teams to integrate their models into our platform
Required Skills
  • Bachelor's degree in Computer Science, Electrical Engineering, or a related technical field (or equivalent practical experience)
  • Strong foundation in systems programming, with a track record of identifying and resolving bottlenecks
  • Deep expertise in PyTorch, TensorRT, TransformerEngine, Nsight, ONNX Runtime
  • Model compilation, quantization (INT8/FP16), and advanced serving architectures
  • Working knowledge of GPU hardware (NVIDIA)
  • Strong understanding of transformer architectures and modern ML optimization techniques
Benefits
  • Competitive SF salary and meaningful early equity
  • Visa sponsorship and relocation support
  • Generous health, dental, and vision coverage
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Inference Engineer San Francisco · Engineering · Full Time →
ML Inference Engineer San Francisco · Engineering · Full Time →

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Visa sponsorship
Relocation support
Generous health, dental, and vision coverage
Founding Engineer, ML Inference
Founding Engineer, ML Inference

Reactor • San Francisco (CA)

On-site
USD 180,000 - 280,000
Competitive salary
Early equity
Health, dental, and vision coverage
+1
High-Performance ML Inference Engineer
High-Performance ML Inference Engineer

Reactor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive SF salary
Early equity
Visa sponsorship
+2
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
High-Performance ML Inference Engineer for Diffusion Models
High-Performance ML Inference Engineer for Diffusion Models

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Visa sponsorship
Relocation support
Generous health, dental, and vision coverage
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Health, dental, and vision benefits
401(k) company match
Visa sponsorship on case-by-case basis
Software Engineer, Inference - Performance Optimization
Software Engineer, Inference - Performance Optimization

OpenAI • Los Angeles (CA)

On-site
USD 295,000 - 555,000
Equity
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Software Engineer, Inference Runtime
Software Engineer, Inference Runtime

EngRadar • New York (NY)

On-site
USD 150,000 - 230,000
Equity grants
Medical plan
Vision plan
+5