High-Performance ML Inference Engineer for Diffusion Models

Reactor

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Visa sponsorship
Relocation support
Generous health, dental, and vision coverage

Job summary

Reactor is looking for an experienced ML Inference Engineer with deep expertise in high-performance ML engineering. This role focuses on optimizing the performance of generative media models, contributing to Reactor's competitive edge.

The ideal candidate will drive model performance for diffusion models, collaborate with partner teams, and implement optimizations. A Bachelor's degree in a related field is required, along with strong knowledge of GPU hardware and modern ML optimization techniques. The position offers competitive salary and meaningful equity.

Qualifications

  • Strong foundation in systems programming, with a track record of identifying and resolving bottlenecks.
  • Working knowledge of GPU hardware (NVIDIA).
  • Strong understanding of transformer architectures and modern ML optimization techniques.

Responsibilities

  • Drive frontier position on model performance for diffusion models.
  • Design and implement a in-house inference runtime.
  • Profile and benchmark model performance to identify computational bottlenecks.

Skills

Deep expertise in PyTorch
TensorRT
Advanced serving architectures
Model compilation
Quantization

Education

Bachelor's degree in Computer Science, Electrical Engineering, or related

Tools

Nsight
ONNX Runtime

Job description

Reactor is looking for an experienced ML Inference Engineer with deep expertise in high-performance ML engineering. This role focuses on optimizing the performance of generative media models, contributing to Reactor's competitive edge.

The ideal candidate will drive model performance for diffusion models, collaborate with partner teams, and implement optimizations. A Bachelor's degree in a related field is required, along with strong knowledge of GPU hardware and modern ML optimization techniques. The position offers competitive salary and meaningful equity.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Inference Engineer San Francisco · Engineering · Full Time →
ML Inference Engineer San Francisco · Engineering · Full Time →

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Visa sponsorship
Relocation support
Generous health, dental, and vision coverage
Founding Engineer, ML Inference
Founding Engineer, ML Inference

Reactor • San Francisco (CA)

On-site
USD 180,000 - 280,000
Competitive salary
Early equity
Health, dental, and vision coverage
+1
Forward Deployed Engineer
Forward Deployed Engineer

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive salary
Visa sponsorship
Health, dental, and vision coverage
+1
Forward Deployed Engineer San Francisco · Engineering · Full Time →
Forward Deployed Engineer San Francisco · Engineering · Full Time →

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive salary
Early equity
Visa sponsorship
+2
Inference Runtime Engineer for LLMs & Diffusion
Inference Runtime Engineer for LLMs & Diffusion

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Health, dental, and vision benefits
401(k) company match
Visa sponsorship on case-by-case basis
Member of Technical Staff — Inference-Multimodal & Diffusion
Member of Technical Staff — Inference-Multimodal & Diffusion

RadixArk • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Meaningful equity
Comprehensive benefits
+1
Staff Technical Lead for Inference & ML Performance
Staff Technical Lead for Inference & ML Performance

fal • San Francisco (CA)

On-site
USD 150,000 - 200,000
Member of Technical Staff, Inference & Serving
Member of Technical Staff, Inference & Serving

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff ML Systems Engineer - Diffusion LLM Serving
Staff ML Systems Engineer - Diffusion LLM Serving

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Inference-Multimodal & DiffusionPalo Alto, CA
Member of Technical Staff — Inference-Multimodal & DiffusionPalo Alto, CA

RadixArk • Palo Alto (CA)

On-site
USD 180,000 - 240,000