ML Systems Engineer, Inference (Fully Remote)

Runpod

United States

Remote

USD 150,000 - 220,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity
Medical, dental, and vision
Remote-first
Home office stipend

Job summary

Runpod is hiring a seasoned systems engineer to optimize LLM inference performance in a remote-first environment. You will define measurement pipelines for throughput, latency, and cost per token, and build repeatable tooling for rigorous benchmarking.

You will profile the serving stack from scheduler and memory to GPU kernels, diagnose bottlenecks, and implement optimizations for large models on single and multi-node GPU deployments. Join a fast-moving AI infrastructure team.

Qualifications

  • 5+ years of professional system engineering experience with hands-on work in vLLM, SGLang, or a comparable serving engine in production.
  • Strong Python software engineering skills and understanding of what drives LLM inference performance: batching, memory, parallelism, and latency/throughput trade-offs.
  • Experience with modern inference optimization techniques such as quantization, speculative decoding, or distributed serving, plus rigorous benchmarking and GPU profiling.

Responsibilities

  • Define how we measure inference performance, including throughput, time to first token, inter-token latency, and cost per token, and build tooling for rigorous, repeatable measurements.
  • Profile and diagnose performance problems across the serving stack, from scheduling and memory management down to kernels and interconnect.
  • Improve serving efficiency for large, state-of-the-art models on single-node and multi-node GPU deployments.

Skills

Python
LLM serving
Benchmarking
GPU profiling
Throughput optimization

Tools

vLLM
SGLang
Profiling tools

Job description

  • Define how we measure inference performance, including throughput, time to first token, inter-token latency, and cost per token, and build tooling for rigorous, repeatable measurements.
  • Profile and diagnose performance problems across the serving stack, from scheduling and memory management down to kernels and interconnect.
  • Improve serving efficiency for large, state-of-the-art models on single-node and multi-node GPU deployments.
Requirements:
  • 5+ years of professional system engineering experience with deep, hands-on experience in vLLM, SGLang, or a comparable serving engine in production.
  • Strong software engineering skills in Python and a solid understanding of what drives LLM inference performance: batching, memory, parallelism, and latency/throughput trade-offs.
  • Experience with modern inference optimization techniques such as quantization, speculative decoding, or distributed serving, plus rigor in benchmarking and GPU profiling.
What You'll Receive:
  • Competitive base pay from $150,000 - $220,000, meaningful equity, and generous medical, dental, and vision plans.
  • Flexible PTO, a $1,200 home office and equipment stipend, and the opportunity to join a passionate team on the cutting edge of AI infrastructure.
  • Most roles are remote-first with inclusive, collaborative teams using Slack for internal communication.

Runpod is the AI Developer Cloud, providing a platform for over one million developers to experiment, train, fine-tune, deploy, and scale AI. We're a small, remote-first team that takes ownership seriously, moves fast, and has processed more than 20 billion inference requests.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Systems Engineer, Inference
Senior ML Systems Engineer, Inference

Runpod • Mount Laurel Township (NJ)

Remote
USD 150,000 - 220,000
Base salary
Equity
Medical/dental/vision
+3
Senior ML Inference Systems Engineer (Remote-First)
Senior ML Inference Systems Engineer (Remote-First)

Runpod • United States

Remote
USD 150,000 - 220,000
Equity
Medical, dental, and vision
Remote-first
+1
Lead ML Inference Systems Engineer (Remote)
Lead ML Inference Systems Engineer (Remote)

Showcify, Inc. • United States

Remote
USD 150,000 - 220,000
Equity/options participation
Medical, dental & vision plans
Flexible PTO
+2
Remote-First Lead ML Inference Systems Engineer
Remote-First Lead ML Inference Systems Engineer

Runpod • Mount Laurel Township (NJ)

Remote
USD 150,000 - 220,000
Base salary
Equity
Medical/dental/vision
+3
Inference Performance Engineer
Inference Performance Engineer

Adaption Labs • San Francisco (CA)

On-site
USD 180,000 - 260,000
Flexible work
Adaption Passport
Lunch stipend
+1
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

On-site
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Software Engineer, Inference Runtime
Software Engineer, Inference Runtime

Lm-Studio • New York (NY)

On-site
USD 140,000 - 210,000
Competitive salary and equity grants
Excellent medical, vision, dental care
Catered team lunch / expensed dinners
+3
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Kindredventures • Palo Alto (CA)

On-site
USD 190,000 - 250,000
Comprehensive health insurance
Dental insurance
Vision insurance
+1