Software Engineer — Inference Stack for Scalable LLMs

BaseTen

New York, San Francisco (NY, CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Health insurance
Flexible PTO
Parental leave
401(k)

Job summary

Baseten in New York is seeking a Software Engineer for the Inference Stack to advance large-scale LLM inference. You’ll work across the stack—from developer-facing features to low-level infrastructure, Kubernetes deployments, and traffic routing.

You’ll own production systems and solve hard integration challenges to keep our platform fast, reliable, and easy to use. You should be excited by building in production, collaborating with Model Performance engineers, and shaping testing, release

Qualifications

  • Experience building and operating production systems with high reliability, latency, and scale.
  • Strong focus on developer experience and how end users interact with systems.
  • Willingness to learn new languages, frameworks, and systems as needed.
  • Ability to debug complex multi-layered systems across stacks.
  • Excellent collaboration and communication skills.

Responsibilities

  • Develop infrastructure and orchestration systems for deploying large-scale distributed LLM inference.
  • Work across the stack—from customer-facing features to low-level infrastructure components.
  • Build platform capabilities for routing, autoscaling, scheduling, observability, and runtime management.
  • Improve reliability, scalability, and usability of the inference stack.
  • Collaborate with Model Performance engineers to deploy inference optimizations for customers.
  • Define testing, release automation, benchmarking, and operational practices.
  • Debug production systems spanning Kubernetes, distributed runtimes, networking, and GPU workloads.
  • Make thoughtful engineering tradeoffs balancing performance, reliability, and developer experience.
  • Own projects end-to-end: architecture, implementation, deployment, monitoring, iteration.

Skills

Distributed systems
Backend infrastructure
Platform engineering
Developer experience
Debugging complex systems
Communication skills

Education

Bachelor’s/Master’s/PhD in CS or related

Job description

Baseten in New York is seeking a Software Engineer for the Inference Stack to advance large-scale LLM inference. You’ll work across the stack—from developer-facing features to low-level infrastructure, Kubernetes deployments, and traffic routing.

You’ll own production systems and solve hard integration challenges to keep our platform fast, reliable, and easy to use. You should be excited by building in production, collaborating with Model Performance engineers, and shaping testing, release

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Inference Stack (LLM Infra)
Software Engineer, Inference Stack (LLM Infra)

Baseten • United States

Remote
USD 180,000 - 240,000
LLM Inference Platform Engineer
LLM Inference Platform Engineer

Baseten • New York (NY)

On-site
USD 216,000 - 360,000
Equity
Comprehensive health coverage
Flexible PTO
+4
Software Engineer - AI Inference Platform
Software Engineer - AI Inference Platform

The Consensus • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation with equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
Senior Backend Engineer, LLM Inference Systems
Senior Backend Engineer, LLM Inference Systems

Inception • San Francisco (CA)

On-site
USD 150,000 - 230,000
LLM Platform Lead — Scalable Training & Inference
LLM Platform Lead — Scalable Training & Inference

Scale AI, Inc. • New York (NY)

On-site
USD 264,000 - 331,000
Comprehensive health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
Staff Engineer, Model Efficiency & LLM Inference
Staff Engineer, Model Efficiency & LLM Inference

Visa Hunt • New York (NY)

Hybrid
USD 150,000 - 210,000
Lunch stipend
Health and dental benefits
RRSP matching
+4
Inference Runtime Engineer — On-Device & Cloud ML
Inference Runtime Engineer — On-Device & Cloud ML

Lm-Studio • New York (NY)

Hybrid
USD 140,000 - 210,000
Competitive salary and equity grants
Excellent medical, vision, dental care
Catered team lunch / expensed dinners
+3
Staff ML Systems Engineer - Diffusion LLM Serving
Staff ML Systems Engineer - Diffusion LLM Serving

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Engineer, Inference Runtime — Performance & Scale
Staff Engineer, Inference Runtime — Performance & Scale

Menlo Ventures • New York (NY)

Hybrid
USD 405,000 - 485,000
Remote Senior Inference Engineer - Production LLM Pipelines
Remote Senior Inference Engineer - Production LLM Pipelines

Loft Labs, Inc. dba vCluster Labs • United States

On-site
USD 140,000 - 190,000