Software Engineer, Inference Stack (LLM Infra)

Baseten

United States

Remote

USD 180,000 - 240,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Baseten is seeking a Software Engineer for the Inference Stack to own distributed infrastructure for large-scale LLM inference. You will navigate the stack from developer-facing features to low-level systems, enabling robust, scalable model deployments.

You’ll work on routing, autoscaling, observability, and runtime management while collaborating with Model Performance engineers to push optimizations broadly to customers. This role emphasizes reliability, performance, and developer experience.

Qualifications

  • Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, or a related field.
  • Strong background in distributed systems, backend infrastructure, or platform engineering.
  • Proven experience building production-grade systems.

Responsibilities

  • Develop infrastructure and orchestration systems for deploying and managing large-scale distributed LLM inference.
  • Work across the stack from customer-facing features to low-level infrastructure components.
  • Build platform capabilities related to routing, autoscaling, scheduling, observability, and runtime management.
  • Improve the reliability, scalability, and usability of the inference stack.
  • Collaborate with Model Performance engineers to make inference optimizations broadly available to customers and easy to configure.
  • Help define best practices around testing, release automation, benchmarking, and operational excellence.
  • Debug complex production systems spanning Kubernetes, distributed runtimes, networking, and GPU workloads.
  • Make thoughtful engineering tradeoffs balancing performance, reliability, operational simplicity, and developer experience.
  • Own projects end-to-end: from architecture and implementation through deployment, monitoring, and iteration based on customer feedback

Skills

Distributed systems
Backend infrastructure
Platform engineering
Production experience

Education

Bachelor's/Master's/PhD in CS/Engineering or related
Advanced degree preferred

Tools

Kubernetes
Distributed runtimes

Job description

Baseten is seeking a Software Engineer for the Inference Stack to own distributed infrastructure for large-scale LLM inference. You will navigate the stack from developer-facing features to low-level systems, enabling robust, scalable model deployments.

You’ll work on routing, autoscaling, observability, and runtime management while collaborating with Model Performance engineers to push optimizations broadly to customers. This role emphasizes reliability, performance, and developer experience.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer — Inference Stack for Scalable LLMs
Software Engineer — Inference Stack for Scalable LLMs

BaseTen • New York (NY), San Francisco (CA)

On-site
USD 180,000 - 240,000
Equity
Health insurance
Flexible PTO
+2
LLM Inference Platform Engineer
LLM Inference Platform Engineer

Baseten • New York (NY)

On-site
USD 216,000 - 360,000
Equity
Comprehensive health coverage
Flexible PTO
+4
Distributed LLM Inference Engineer - Scale & Resilience
Distributed LLM Inference Engineer - Scale & Resilience

Baseten • San Francisco (CA)

On-site
USD 180,000 - 360,000
Equity
Health coverage
Flexible PTO
+4
Software Engineer - Baseten Inference Stack
Software Engineer - Baseten Inference Stack

Baseten • San Francisco (CA)

On-site
USD 180,000 - 360,000
Equity
Health coverage
Flexible PTO
+4
Software Engineer - Baseten Inference Stack
Software Engineer - Baseten Inference Stack

BaseTen • New York (NY), San Francisco (CA)

On-site
USD 180,000 - 240,000
Equity
Health insurance
Flexible PTO
+2
Software Engineer - AI Inference Platform
Software Engineer - AI Inference Platform

The Consensus • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation with equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
Senior Backend Engineer, LLM Inference Systems
Senior Backend Engineer, LLM Inference Systems

Inception • San Francisco (CA)

On-site
USD 150,000 - 230,000
ML Inference Infrastructure Engineer
ML Inference Infrastructure Engineer

Baseten • United States

Remote
USD 120,000 - 190,000
Equity
Medical coverage for employee and dep.
Flexible PTO including Winter Break
+3
Software Engineer- BIS (Baseten Inference Stack)
Software Engineer- BIS (Baseten Inference Stack)

The Consensus • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation with equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
Senior LLM Inference Systems Engineer
Senior LLM Inference Systems Engineer

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Bonus and equity plan
Medical, dental, vision insurance
401(k) retirement plan
+1