Inference Runtime Engineer for LLMs & Diffusion

Inferact

San Francisco (CA)

Hybrid

USD 200,000 - 400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental, and vision benefits
401(k) company match
Visa sponsorship on case-by-case basis

Job summary

Inferact is seeking an inference runtime engineer to enhance the performance and capabilities of LLM and diffusion model serving. This role requires expertise in optimizing model execution on various hardware architectures and has significant implications for AI inference.

The ideal candidate must possess a bachelor's degree in computer science or related fields, strong programming skills in Python, and experience with LLM inference systems. Remote work options are available for exceptional candidates.

Qualifications

  • Bachelor's degree or equivalent experience in computer science, engineering, or similar.
  • Deep understanding of transformer architectures and their variants.
  • Strong programming skills in Python with experience in PyTorch internals.
  • Experience with LLM inference systems (vLLM, TensorRT-LLM, SGLang, TGI).
  • Ability to read and implement model architectures from research papers.

Responsibilities

  • Push the boundaries of inference model serving.
  • Optimize how models execute across diverse hardware.
  • Directly impact how the world runs AI inference.

Skills

Programming skills in Python
Deep understanding of transformer architectures
Experience with LLM inference systems
Ability to read model architectures

Education

Bachelor's degree or equivalent experience

Tools

PyTorch
TensorRT-LLM
vLLM

Job description

Inferact is seeking an inference runtime engineer to enhance the performance and capabilities of LLM and diffusion model serving. This role requires expertise in optimizing model execution on various hardware architectures and has significant implications for AI inference.

The ideal candidate must possess a bachelor's degree in computer science or related fields, strong programming skills in Python, and experience with LLM inference systems. Remote work options are available for exceptional candidates.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Inference
Member of Technical Staff, Inference

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Health, dental, and vision benefits
401(k) company match
Visa sponsorship on case-by-case basis
Staff ML Systems Engineer - Diffusion LLM Serving
Staff ML Systems Engineer - Diffusion LLM Serving

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Inference Runtime Engineer — On-Device & Cloud ML
Inference Runtime Engineer — On-Device & Cloud ML

Lm-Studio • New York (NY)

Hybrid
USD 140,000 - 210,000
Competitive salary and equity grants
Excellent medical, vision, dental care
Catered team lunch / expensed dinners
+3
Edge-to-Cloud LLM Inference Engineer
Edge-to-Cloud LLM Inference Engineer

Insider, Inc. • United States

On-site
USD 100,000 - 140,000
LLM Inference Optimization Engineer — Frontier Performance
LLM Inference Optimization Engineer — Frontier Performance

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Remote Inference Optimization Engineer
Remote Inference Optimization Engineer

Modular Mailing Systems, Inc. • Los Altos (CA)

Hybrid
USD 198,000 - 286,000
Premier insurance plans
5% 401k matching
Flexible paid time off
+2
LLM Inference Deployment Engineer
LLM Inference Deployment Engineer

Insider, Inc. • United States

On-site
USD 100,000 - 140,000
Member of Technical Staff, Inference & Serving
Member of Technical Staff, Inference & Serving

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Engineer, Model Efficiency & LLM Inference
Staff Engineer, Model Efficiency & LLM Inference

Visa Hunt • New York (NY)

Hybrid
USD 150,000 - 210,000
Lunch stipend
Health and dental benefits
RRSP matching
+4
LLM Inference Systems Performance Engineer
LLM Inference Systems Performance Engineer

3M HEALTHCARE • Austin (TX)

On-site
USD 150,000 - 210,000
Medical, dental, and vision coverage
Income protection benefits
Paid family leave
+1