Lead ML Inference Systems Engineer (Remote)

Showcify, Inc.

United States

Remote

USD 150,000 - 220,000

Full time

7 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity/options participation
Medical, dental & vision plans
Flexible PTO
Remote-friendly culture
Home office stipend

Job summary

Runpod is hiring a ML Systems Engineer, Inference to lead the performance end-to-end of LLM serving. You will measure, diagnose, and improve latency and cost across models, hardware generations, and workloads in a remote-first environment.

You will ship fixes that directly affect customer experience, working hands-on on modern inference engines and optimization techniques, with opportunities to influence tooling and deployment strategies across Runpod’s platform.

Qualifications

  • 5+ years of professional system engineering experience.
  • Deep, hands-on experience with vLLM, SGLang (or a comparable serving engine).
  • Strong software engineering skills in Python.
  • Solid understanding of what drives LLM inference performance: batching, memory, parallelism, latency-vs-throughput.
  • Experience with modern inference optimization techniques such as quantization, speculative decoding, or distributed serving.
  • Rigor in benchmarking and performance analysis, plus comfort with GPU profiling tools.
  • Ability to explain results clearly in writing and turn them into decisions.

Responsibilities

  • Define how we measure inference performance (throughput, time to first token, inter-token latency, cost per token) and build tooling for rigorous, repeatable measurements.
  • Profile and diagnose performance problems across the serving stack from scheduling to kernels.
  • Improve serving efficiency for large models on single-node and multi-node GPU deployments.
  • Turn learnings into production-ready runtimes, configurations, and defaults for customers.
  • Work with product and infrastructure teams to shape inference offerings on Runpod.
  • Keep up with the inference ecosystem and decide what to adopt, build, or contribute back.
  • Trace bottlenecks in the serving engine/runtime and implement fixes when simple tuning isn’t enough.

Skills

Python
Benchmarking

Tools

vLLM
SGLang
CUDA
Triton

Job description

Runpod is hiring a ML Systems Engineer, Inference to lead the performance end-to-end of LLM serving. You will measure, diagnose, and improve latency and cost across models, hardware generations, and workloads in a remote-first environment.

You will ship fixes that directly affect customer experience, working hands-on on modern inference engines and optimization techniques, with opportunities to influence tooling and deployment strategies across Runpod’s platform.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote-First Lead ML Inference Systems Engineer
Remote-First Lead ML Inference Systems Engineer

Runpod • Mount Laurel Township (NJ)

Remote
USD 150,000 - 220,000
Base salary
Equity
Medical/dental/vision
+3
Senior ML Inference Systems Engineer (Remote-First)
Senior ML Inference Systems Engineer (Remote-First)

Runpod • United States

Remote
USD 150,000 - 220,000
Equity
Medical, dental, and vision
Remote-first
+1
Senior ML Systems Engineer, Inference
Senior ML Systems Engineer, Inference

Runpod • Mount Laurel Township (NJ)

Remote
USD 150,000 - 220,000
Base salary
Equity
Medical/dental/vision
+3
ML Systems Engineer, Inference (Fully Remote)
ML Systems Engineer, Inference (Fully Remote)

Runpod • United States

Remote
USD 150,000 - 220,000
Equity
Medical, dental, and vision
Remote-first
+1
Engineering Manager, Cloud AI Platform
Engineering Manager, Cloud AI Platform

RunPod Inc. • United States

On-site
USD 110,000 - 220,000
Equity
Medical benefits
Flexible PTO
+2
Remote-First AI Infra Engineer - Customer Impact
Remote-First AI Infra Engineer - Customer Impact

The POD Network • United States

Remote
USD 100,000 - 160,000
Equity
Flexible PTO
Remote-first culture
+1
Engineering Manager, AI Infra & Cloud Platforms
Engineering Manager, AI Infra & Cloud Platforms

Runpod • United States

On-site
USD 110,000 - 220,000
Stock options
Remote-first culture
Home office stipend
ML Inference Systems Engineer – Remote, Scalable, Low-Latency
ML Inference Systems Engineer – Remote, Scalable, Low-Latency

Atlassian Corp. • Seattle (WA), Northern (KY)

Hybrid
USD 178,000 - 233,000
Health and wellbeing resources
Volunteer days
Inference Runtime Engineer — On-Device & Cloud ML
Inference Runtime Engineer — On-Device & Cloud ML

Lm-Studio • New York (NY)

Hybrid
USD 140,000 - 210,000
Competitive salary and equity grants
Excellent medical, vision, dental care
Catered team lunch / expensed dinners
+3
Cloud Platform Engineering Manager for Scalable AI
Cloud Platform Engineering Manager for Scalable AI

Runpod • Mount Laurel Township (NJ)

Remote
USD 110,000 - 220,000
Equity
Medical, dental, vision
Flexible PTO
+2