Remote Inference Engine Engineer - LLMs & Diffusion

Inferact

United States

Remote

USD 130,000 - 180,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Inferact is seeking an inference runtime engineer to push the boundaries of LLM and diffusion model serving. You will optimize how models execute across diverse hardware and architectures, contributing to the core of vLLM and related projects.

This fully remote role offers salary plus equity, visa sponsorship on a case-by-case basis, and comprehensive benefits. Regular overlap with Pacific Time is expected for critical syncs.

Qualifications

  • Bachelor's degree or equivalent experience in computer science, engineering, or similar.
  • Deep understanding of transformer architectures and their variants.
  • Strong programming skills in Python with experience in PyTorch internals.
  • Experience with LLM inference systems (vLLM, TensorRT-LLM, SGLang, TGI).
  • Ability to read and implement model architectures and inference techniques from research papers.
  • Demonstrate the ability to contribute performant and maintainable code and debug in complex ML codebases.

Skills

Python programming
PyTorch internals
Transformer architectures
LLM inference systems
Reading research papers
Debug ML codebases

Education

Bachelor's degree or equivalent experience in CS/Engineering

Tools

vLLM
TensorRT-LLM
SGLang
TGI

Job description

Inferact is seeking an inference runtime engineer to push the boundaries of LLM and diffusion model serving. You will optimize how models execute across diverse hardware and architectures, contributing to the core of vLLM and related projects.

This fully remote role offers salary plus equity, visa sponsorship on a case-by-case basis, and comprehensive benefits. Regular overlap with Pacific Time is expected for critical syncs.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Runtime Engineer for LLMs & Diffusion
Inference Runtime Engineer for LLMs & Diffusion

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Health, dental, and vision benefits
401(k) company match
Visa sponsorship on case-by-case basis
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Health, dental, and vision benefits
401(k) company match
Visa sponsorship on case-by-case basis
Staff ML Systems Engineer - Diffusion LLM Serving
Staff ML Systems Engineer - Diffusion LLM Serving

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Exceptional Generalist (Remote)
Member of Technical Staff, Exceptional Generalist (Remote)

Inferact • United States

Remote
USD 180,000 - 240,000
Competitive salary and equity
Visa sponsorship
Health coverage where applicable
Staff Engineer, Model Efficiency & LLM Inference
Staff Engineer, Model Efficiency & LLM Inference

Visa Hunt • New York (NY)

Hybrid
USD 150,000 - 210,000
Lunch stipend
Health and dental benefits
RRSP matching
+4
Remote Inference Optimization Engineer
Remote Inference Optimization Engineer

Modular Mailing Systems, Inc. • Los Altos (CA)

Hybrid
USD 198,000 - 286,000
Premier insurance plans
5% 401k matching
Flexible paid time off
+2
Member of Technical Staff, Performance and Scale
Member of Technical Staff, Performance and Scale

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Generous health, dental, and vision benefits
401(k) company match
Equity options
Head of Engineering
Head of Engineering

Inferact • San Francisco (CA)

On-site
USD 260,000 - 380,000
Health, dental, vision benefits
401(k) company match
Staff Research Engineer, LLM Inference & Efficiency Remote
Staff Research Engineer, LLM Inference & Efficiency Remote

Cohere • San Francisco (CA), New York (NY)

Hybrid
USD 180,000 - 240,000
Lunch stipend
Health and dental benefits
Parental leave top‑up
+3
Production-Grade LLM Inference Runtime Engineer
Production-Grade LLM Inference Runtime Engineer

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 360,000