Software Engineer — Inference Stack for Scalable LLMs

BaseTen

New York, San Francisco (NY, CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity
Health insurance
Flexible PTO
Parental leave
401(k)

Job summary

Baseten in New York is seeking a Software Engineer for the Inference Stack to advance large-scale LLM inference. You’ll work across the stack—from developer-facing features to low-level infrastructure, Kubernetes deployments, and traffic routing.

You’ll own production systems and solve hard integration challenges to keep our platform fast, reliable, and easy to use. You should be excited by building in production, collaborating with Model Performance engineers, and shaping testing, release

Qualifications

  • Experience building and operating production systems with high reliability, latency, and scale.
  • Strong focus on developer experience and how end users interact with systems.
  • Willingness to learn new languages, frameworks, and systems as needed.
  • Ability to debug complex multi-layered systems across stacks.
  • Excellent collaboration and communication skills.

Responsibilities

  • Develop infrastructure and orchestration systems for deploying large-scale distributed LLM inference.
  • Work across the stack—from customer-facing features to low-level infrastructure components.
  • Build platform capabilities for routing, autoscaling, scheduling, observability, and runtime management.
  • Improve reliability, scalability, and usability of the inference stack.
  • Collaborate with Model Performance engineers to deploy inference optimizations for customers.
  • Define testing, release automation, benchmarking, and operational practices.
  • Debug production systems spanning Kubernetes, distributed runtimes, networking, and GPU workloads.
  • Make thoughtful engineering tradeoffs balancing performance, reliability, and developer experience.
  • Own projects end-to-end: architecture, implementation, deployment, monitoring, iteration.

Skills

Distributed systems
Backend infrastructure
Platform engineering
Developer experience
Debugging complex systems
Communication skills

Education

Bachelor’s/Master’s/PhD in CS or related

Job description

Baseten in New York is seeking a Software Engineer for the Inference Stack to advance large-scale LLM inference. You’ll work across the stack—from developer-facing features to low-level infrastructure, Kubernetes deployments, and traffic routing.

You’ll own production systems and solve hard integration challenges to keep our platform fast, reliable, and easy to use. You should be excited by building in production, collaborating with Model Performance engineers, and shaping testing, release

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer, Inference Stack (LLM Infra)
Software Engineer, Inference Stack (LLM Infra)

Baseten • United States

Remote
USD 180,000 - 240,000
LLM Inference Platform Engineer
LLM Inference Platform Engineer

Baseten • New York (NY)

On-site
USD 216,000 - 360,000
Equity
Comprehensive health coverage
Flexible PTO
+4
Distributed LLM Inference Engineer - Scale & Resilience
Distributed LLM Inference Engineer - Scale & Resilience

Baseten • San Francisco (CA)

On-site
USD 180,000 - 360,000
Equity
Health coverage
Flexible PTO
+4
Software Engineer - AI Inference Platform
Software Engineer - AI Inference Platform

The Consensus • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation with equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
LLM Inference Systems Engineer
LLM Inference Systems Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior Backend Engineer, LLM Inference Systems
Senior Backend Engineer, LLM Inference Systems

Inception • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior LLM Inference Systems Engineer
Senior LLM Inference Systems Engineer

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Bonus and equity plan
Medical, dental, vision insurance
401(k) retirement plan
+1
LLM Platform Lead — Scalable Training & Inference
LLM Platform Lead — Scalable Training & Inference

Scale AI, Inc. • New York (NY)

On-site
USD 264,800 - 331,000
Comprehensive health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
Inference Runtime Engineer — On-Device & Cloud ML
Inference Runtime Engineer — On-Device & Cloud ML

Lm-Studio • New York (NY)

Hybrid
USD 140,000 - 210,000
Competitive salary and equity grants
Excellent medical, vision, dental care
Catered team lunch / expensed dinners
+3
Software Engineer, Model Inference & LLM Deployment
Software Engineer, Model Inference & LLM Deployment

Google LLC • Mountain View (CA), Northern (KY)

Hybrid
USD 180,000 - 300,000