Distributed LLM Inference Engineer - Scale & Resilience

Baseten

San Francisco (CA)

On-site

USD 180,000 - 360,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Health coverage
Flexible PTO
Parental leave
Fertility stipend
401(k)
Startup exposure

Job summary

Baseten, based in San Francisco, seeks a Software Engineer for the Inference Stack to build and operate large-scale LLM inference systems. You will own production-grade infrastructure, spanning deployment orchestration, routing, and observability, and collaborate with model performance engineers to push optimizations to customers.

You will work across the stack, from customer-facing features to low-level components, improving reliability, scalability, and developer experience while owning

Qualifications

  • Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, or a related field.
  • Strong background in distributed systems, backend infrastructure, or platform engineering.
  • Experience building and operating production systems where reliability, latency, and scale are first-class concerns.
  • Strong sense of developer experience: you think about how systems are used, not just how they work.
  • Motivated and willing to learn new languages, frameworks, and systems as needed.
  • Ability to debug complex systems across multiple layers of the stack.
  • Genuine interest in inference engineering. You don’t need to have hands on experience but are willing to learn.
  • Excellent communication and collaboration skills.

Responsibilities

  • Develop infrastructure and orchestration systems for deploying and managing large-scale distributed LLM inference
  • Work across the stack, from customer-facing features to low-level infrastructure components
  • Build platform capabilities related to routing, autoscaling, scheduling, observability, and runtime management
  • Improve the reliability, scalability, and usability of our inference stack
  • Collaborate closely with Model Performance engineers to make new inference optimizations broadly available to customers
  • Help define best practices around testing, release automation, benchmarking, and operational excellence
  • Debug complex production systems spanning Kubernetes, distributed runtimes, networking, and GPU workloads
  • Make thoughtful engineering tradeoffs balancing performance, reliability, operational simplicity, and developer experience
  • Own projects end-to-end: from architecture and implementation through deployment, monitoring, and iteration based on customer feedback

Skills

Distributed systems
Backend infrastructure
Platform engineering
Kubernetes
Performance optimization

Education

Bachelor's degree
Master's degree
Ph.D.

Tools

Kubernetes
Dynamo
vLLM
TensorRT-LLM
SGLang

Job description

Baseten, based in San Francisco, seeks a Software Engineer for the Inference Stack to build and operate large-scale LLM inference systems. You will own production-grade infrastructure, spanning deployment orchestration, routing, and observability, and collaborate with model performance engineers to push optimizations to customers.

You will work across the stack, from customer-facing features to low-level components, improving reliability, scalability, and developer experience while owning

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Inference Stack (LLM Infra)
Software Engineer, Inference Stack (LLM Infra)

Baseten • United States

Remote
USD 180,000 - 240,000
Software Engineer — Inference Stack for Scalable LLMs
Software Engineer — Inference Stack for Scalable LLMs

BaseTen • New York (NY), San Francisco (CA)

On-site
USD 180,000 - 240,000
Equity
Health insurance
Flexible PTO
+2
LLM Inference Platform Engineer
LLM Inference Platform Engineer

Baseten • New York (NY)

On-site
USD 216,000 - 360,000
Equity
Comprehensive health coverage
Flexible PTO
+4
Senior Backend Engineer, LLM Inference Systems
Senior Backend Engineer, LLM Inference Systems

Inception • San Francisco (CA)

On-site
USD 150,000 - 230,000
Distributed LLM Inference Engineer - Scale HighThroughput AI
Distributed LLM Inference Engineer - Scale HighThroughput AI

Cerebras • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Stock Options
Healthcare plans with 99% premium coverage
401k Retirement Plan
+6
Staff ML Systems Engineer - Diffusion LLM Serving
Staff ML Systems Engineer - Diffusion LLM Serving

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Engineer, Model Efficiency & LLM Inference
Staff Engineer, Model Efficiency & LLM Inference

Visa Hunt • New York (NY)

Hybrid
USD 150,000 - 210,000
Lunch stipend
Health and dental benefits
RRSP matching
+4
ML Inference Infrastructure Engineer
ML Inference Infrastructure Engineer

Baseten • United States

Remote
USD 120,000 - 190,000
Equity
Medical coverage for employee and dep.
Flexible PTO including Winter Break
+3
Senior LLM Inference Engineer - End-to-End, Equity
Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA
Software Engineer - AI Inference Platform
Software Engineer - AI Inference Platform

The Consensus • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation with equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy