Remote-First Lead ML Inference Systems Engineer

Runpod

Mount Laurel Township (NJ)

Remote

USD 150,000 - 220,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Base salary
Equity
Medical/dental/vision
Flexible PTO
Remote-first culture
Home office stipend

Job summary

Runpod is hiring an ML Systems Engineer, Inference to own end-to-end LLM serving performance, optimizing latency and cost across models and hardware generations. This hands-on role requires identifying bottlenecks, implementing fixes, and delivering reliable production configurations.

You will profile serving stacks, define metrics, and collaborate with product and infra teams to shape Runpod's inference offerings. Remote-first team with competitive compensation and equity.

Qualifications

  • 5+ years of professional system engineering experience.
  • Deep, hands-on experience with vLLM, SGLang (or a comparable serving engine) in production or at serious benchmark scale.
  • Strong software engineering skills in Python. You're comfortable working in large, performance-critical codebases.
  • A solid understanding of what drives LLM inference performance: batching, memory, parallelism, and the trade-offs between latency and throughput.
  • Experience with modern inference optimization techniques such as quantization, speculative decoding, or distributed serving.
  • Rigor in benchmarking and performance analysis, plus comfort with GPU profiling tools.
  • The ability to explain your results clearly in writing and turn them into decisions.
  • CUDA or Triton kernel tuning is a plus.
  • Experience with multi-node GPU systems and high-speed networking.
  • Contributions to inference or ML systems projects.
  • Experience at a company where inference cost and latency were core business metrics.

Responsibilities

  • Define how we measure inference performance, including throughput, time to first token, inter-token latency, and cost per token, and build the tooling that makes those measurements rigorous and repeatable.
  • Profile and diagnose performance problems across the serving stack, from scheduling and memory management down to kernels and interconnect.
  • Improve serving efficiency for large, state-of-the-art models on single-node and multi-node GPU deployments.
  • Turn what you learn into production-ready runtimes, configurations, and defaults that customers benefit from automatically.
  • Work closely with product and infrastructure teams to shape how inference is offered on Runpod.
  • Keep up with the fast-moving inference ecosystem, including the open-source community, and decide what's worth adopting, what's worth building, and what's worth contributing back.
  • Trace bottlenecks in the serving engine/runtime and implement fixes when configuration tuning is not enough.

Skills

Python
Performance benchmarking
GPU profiling
System engineering

Tools

vLLM
SGLang
CUDA
Triton

Job description

Runpod is hiring an ML Systems Engineer, Inference to own end-to-end LLM serving performance, optimizing latency and cost across models and hardware generations. This hands-on role requires identifying bottlenecks, implementing fixes, and delivering reliable production configurations.

You will profile serving stacks, define metrics, and collaborate with product and infra teams to shape Runpod's inference offerings. Remote-first team with competitive compensation and equity.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead ML Inference Systems Engineer (Remote)
Lead ML Inference Systems Engineer (Remote)

Showcify, Inc. • United States

Remote
USD 150,000 - 220,000
Equity/options participation
Medical, dental & vision plans
Flexible PTO
+2
Senior ML Inference Systems Engineer (Remote-First)
Senior ML Inference Systems Engineer (Remote-First)

Runpod • United States

Remote
USD 150,000 - 220,000
Equity
Medical, dental, and vision
Remote-first
+1
Senior ML Systems Engineer, Inference
Senior ML Systems Engineer, Inference

Runpod • Mount Laurel Township (NJ)

Remote
USD 150,000 - 220,000
Base salary
Equity
Medical/dental/vision
+3
ML Systems Engineer, Inference (Fully Remote)
ML Systems Engineer, Inference (Fully Remote)

Runpod • United States

Remote
USD 150,000 - 220,000
Equity
Medical, dental, and vision
Remote-first
+1
Remote-First AI Infra Engineer - Customer Impact
Remote-First AI Infra Engineer - Customer Impact

The POD Network • United States

Remote
USD 100,000 - 160,000
Equity
Flexible PTO
Remote-first culture
+1
Engineering Manager, Cloud AI Platform
Engineering Manager, Cloud AI Platform

RunPod Inc. • United States

On-site
USD 110,000 - 220,000
Equity
Medical benefits
Flexible PTO
+2
Engineering Manager, AI Infra & Cloud Platforms
Engineering Manager, AI Infra & Cloud Platforms

Runpod • United States

On-site
USD 110,000 - 220,000
Stock options
Remote-first culture
Home office stipend
Engineering Manager - Cloud
Engineering Manager - Cloud

Runpod • Mount Laurel Township (NJ)

Remote
USD 110,000 - 220,000
Equity
Medical, dental, vision
Flexible PTO
+2
Remote Senior Storage Engineer - AI Infrastructure
Remote Senior Storage Engineer - AI Infrastructure

Runpod • Mount Laurel Township (NJ)

On-site
USD 180,000 - 260,000
Equity participation
Medical, dental & vision plans
Flexible PTO
+2
Remote AI Cloud Engineer — Onboarding & Customer Tech Support
Remote AI Cloud Engineer — Onboarding & Customer Tech Support

Runpod • Mount Laurel Township (NJ)

Remote
USD 100,000 - 160,000
Stock options
Remote work
Home office stipend
+1