Senior ML Inference Systems Engineer (Remote-First)

Runpod

United States

Remote

USD 150,000 - 220,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity
Medical, dental, and vision
Remote-first
Home office stipend

Job summary

Runpod is hiring a seasoned systems engineer to optimize LLM inference performance in a remote-first environment. You will define measurement pipelines for throughput, latency, and cost per token, and build repeatable tooling for rigorous benchmarking.

You will profile the serving stack from scheduler and memory to GPU kernels, diagnose bottlenecks, and implement optimizations for large models on single and multi-node GPU deployments. Join a fast-moving AI infrastructure team.

Qualifications

  • 5+ years of professional system engineering experience with hands-on work in vLLM, SGLang, or a comparable serving engine in production.
  • Strong Python software engineering skills and understanding of what drives LLM inference performance: batching, memory, parallelism, and latency/throughput trade-offs.
  • Experience with modern inference optimization techniques such as quantization, speculative decoding, or distributed serving, plus rigorous benchmarking and GPU profiling.

Responsibilities

  • Define how we measure inference performance, including throughput, time to first token, inter-token latency, and cost per token, and build tooling for rigorous, repeatable measurements.
  • Profile and diagnose performance problems across the serving stack, from scheduling and memory management down to kernels and interconnect.
  • Improve serving efficiency for large, state-of-the-art models on single-node and multi-node GPU deployments.

Skills

Python
LLM serving
Benchmarking
GPU profiling
Throughput optimization

Tools

vLLM
SGLang
Profiling tools

Job description

Runpod is hiring a seasoned systems engineer to optimize LLM inference performance in a remote-first environment. You will define measurement pipelines for throughput, latency, and cost per token, and build repeatable tooling for rigorous benchmarking.

You will profile the serving stack from scheduler and memory to GPU kernels, diagnose bottlenecks, and implement optimizations for large models on single and multi-node GPU deployments. Join a fast-moving AI infrastructure team.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead ML Inference Systems Engineer (Remote)
Lead ML Inference Systems Engineer (Remote)

Showcify, Inc. • United States

Remote
USD 150,000 - 220,000
Equity/options participation
Medical, dental & vision plans
Flexible PTO
+2
Remote-First Lead ML Inference Systems Engineer
Remote-First Lead ML Inference Systems Engineer

Runpod • Mount Laurel Township (NJ)

Remote
USD 150,000 - 220,000
Base salary
Equity
Medical/dental/vision
+3
ML Systems Engineer, Inference (Fully Remote)
ML Systems Engineer, Inference (Fully Remote)

Runpod • United States

Remote
USD 150,000 - 220,000
Equity
Medical, dental, and vision
Remote-first
+1
Senior ML Systems Engineer, Inference
Senior ML Systems Engineer, Inference

Runpod • Mount Laurel Township (NJ)

Remote
USD 150,000 - 220,000
Base salary
Equity
Medical/dental/vision
+3
Remote-First AI Infra Engineer - Customer Impact
Remote-First AI Infra Engineer - Customer Impact

The POD Network • United States

Remote
USD 100,000 - 160,000
Equity
Flexible PTO
Remote-first culture
+1
Remote Full-Stack Software Engineer for AI Infrastructure
Remote Full-Stack Software Engineer for AI Infrastructure

Grapevine Round1 AI • Northern (KY)

Hybrid
USD 110,000 - 170,000
Senior ML Inference Platform Engineer — Remote
Senior ML Inference Platform Engineer — Remote

Bright Vision Technologies • Beaverton (OR)

Remote
USD 105,000 - 143,000
Senior Data Engineer - AI-Driven Data Pipelines (Remote-First)
Senior Data Engineer - AI-Driven Data Pipelines (Remote-First)

Runpod • Mount Laurel Township (NJ)

On-site
USD 175,000 - 220,000
Equity
Remote-friendly
Flexible PTO
+2
Remote Senior Storage Engineer - AI Infrastructure
Remote Senior Storage Engineer - AI Infrastructure

Runpod • Mount Laurel Township (NJ)

On-site
USD 180,000 - 260,000
Equity participation
Medical, dental & vision plans
Flexible PTO
+2
Senior Data Engineer - Build Scalable Data Platform (Remote)
Senior Data Engineer - Build Scalable Data Platform (Remote)

Runpod • United States

Remote
USD 175,000 - 220,000
Equity
Flexible PTO
Remote-first
+1