High-Performance ML Inference Engineer

Reactor

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive SF salary
Early equity
Visa sponsorship
Relocation support
Health, dental, and vision coverage

Job summary

Reactor in San Francisco is seeking an ML Inference Engineer to maximize performance of generative media models and push ultra-low-latency, high-throughput inference. You will craft an in-house runtime, implement optimizations with PyTorch tools, and collaborate with partner teams to integrate external models.

The role requires deep expertise in PyTorch, TensorRT, CUDA, and model optimization techniques, with a focus on delivering cutting-edge inference capabilities at scale.

Qualifications

  • Bachelor's degree in Computer Science, Electrical Engineering, or equivalent practical experience.
  • Strong foundation in systems programming and bottleneck resolution.
  • Expertise in PyTorch, TensorRT, TransformerEngine, Nsight, ONNX Runtime.

Responsibilities

  • Drive model performance for diffusion models and high-throughput inference.
  • Design and implement a high-performance in-house inference runtime.
  • Apply optimizations using torch.compile, CUDA kernels, and specialized inference frameworks.
  • Quantize, prune, and modify architectures to optimize neural networks.
  • Profile and benchmark model performance to identify bottlenecks.
  • Collaborate with model partner teams to integrate models into the platform.

Skills

PyTorch
TensorRT
CUDA
Model optimization
Quantization
Low-latency inference
GPU programming
Transformer architectures

Education

Bachelor's degree in CS/EE

Tools

Nsight
ONNX Runtime
TransformerEngine

Job description

Reactor in San Francisco is seeking an ML Inference Engineer to maximize performance of generative media models and push ultra-low-latency, high-throughput inference. You will craft an in-house runtime, implement optimizations with PyTorch tools, and collaborate with partner teams to integrate external models.

The role requires deep expertise in PyTorch, TensorRT, CUDA, and model optimization techniques, with a focus on delivering cutting-edge inference capabilities at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Inference Engineer San Francisco · Engineering · Full Time →
ML Inference Engineer San Francisco · Engineering · Full Time →

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Visa sponsorship
Relocation support
Generous health, dental, and vision coverage
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Reactor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive SF salary
Early equity
Visa sponsorship
+2
High-Performance ML Inference Engineer for Diffusion Models
High-Performance ML Inference Engineer for Diffusion Models

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Visa sponsorship
Relocation support
Generous health, dental, and vision coverage
Founding Engineer, ML Inference
Founding Engineer, ML Inference

Reactor • San Francisco (CA)

On-site
USD 180,000 - 280,000
Competitive salary
Early equity
Health, dental, and vision coverage
+1
Inference Performance Engineer: Optimize Model Serving
Inference Performance Engineer: Optimize Model Serving

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
+1
Senior ML Performance Engineer — Ultra-Fast Inference + Equity
Senior ML Performance Engineer — Ultra-Fast Inference + Equity

well-funded deeptech startup • California (MO)

On-site
USD 200,000 - 250,000
Founding ML Inference Engineer — Ultra-Low Latency AI
Founding ML Inference Engineer — Ultra-Low Latency AI

Reactor • San Francisco (CA)

On-site
USD 180,000 - 280,000
Competitive salary
Early equity
Health, dental, and vision coverage
+1
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Founding ML Inference Performance Engineer
Founding ML Inference Performance Engineer

uRun • San Francisco (CA)

On-site
USD 120,000 - 160,000
Health, dental, and vision
401(k) participation
Flexible spending accounts
+3
Low-Latency ML Inference Engineer
Low-Latency ML Inference Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000