LLM Inference Platform Engineer

Baseten

New York (NY)

On-site

USD 216,000 - 360,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity
Comprehensive health coverage
Flexible PTO
Parental leave
Fertility stipend
401(k)
Learning opportunities

Job summary

Baseten in the United States is hiring a Software Engineer for the Inference Stack to build distributed runtime for large‑scale LLM inference. You’ll own production systems spanning Kubernetes, routing, autoscaling, observability, and deployment tooling, with a focus on reliability and developer experience.

You’ll collaborate across teams to push performance improvements, scale, and user‑facing ease‑of‑use features, joining a fast‑moving platform engineering group at the forefront of AI

Qualifications

  • Strong background in distributed systems, backend infrastructure, or platform engineering.
  • Experience building and operating production systems where reliability, latency, and scale are first‑class concerns.
  • Excellent communication and collaboration skills.

Responsibilities

  • Develop infrastructure and orchestration systems for deploying and managing large‑scale distributed LLM inference
  • Work across the stack, from customer‑facing features to low‑level infrastructure components
  • Build platform capabilities related to routing, autoscaling, scheduling, observability, and runtime management
  • Improve the reliability, scalability, and usability of our inference stack
  • Collaborate with Model Performance engineers to make inference optimizations broadly available
  • Define best practices around testing, release automation, benchmarking, and operational excellence
  • Debug complex production systems spanning Kubernetes, distributed runtimes, networking, and GPU workloads
  • Make thoughtful engineering tradeoffs balancing performance, reliability, operational simplicity, and developer experience
  • Own projects end‑to‑end: architecture, implementation, deployment, monitoring, iteration based on customer feedback

Skills

Distributed systems
Backend infrastructure
Platform engineering
Developer experience
Debugging complex systems

Education

Bachelor's, Master's, or Ph.D. in CS/Engineering

Tools

Kubernetes

Job description

Baseten in the United States is hiring a Software Engineer for the Inference Stack to build distributed runtime for large‑scale LLM inference. You’ll own production systems spanning Kubernetes, routing, autoscaling, observability, and deployment tooling, with a focus on reliability and developer experience.

You’ll collaborate across teams to push performance improvements, scale, and user‑facing ease‑of‑use features, joining a fast‑moving platform engineering group at the forefront of AI

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer, Inference Stack (LLM Infra)
Software Engineer, Inference Stack (LLM Infra)

Baseten • United States

Remote
USD 180,000 - 240,000
Software Engineer — Inference Stack for Scalable LLMs
Software Engineer — Inference Stack for Scalable LLMs

BaseTen • New York (NY), San Francisco (CA)

On-site
USD 180,000 - 240,000
Equity
Health insurance
Flexible PTO
+2
Distributed LLM Inference Engineer - Scale & Resilience
Distributed LLM Inference Engineer - Scale & Resilience

Baseten • San Francisco (CA)

On-site
USD 180,000 - 360,000
Equity
Health coverage
Flexible PTO
+4
LLM Inference Systems Engineer
LLM Inference Systems Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
LLM Inference Performance Engineer
LLM Inference Performance Engineer

Emploive • New York (NY)

On-site
USD 180,000 - 360,000
Equity
Insurance for dependents
Winter Break
Head of LLM Inference Platform & Architecture
Head of LLM Inference Platform & Architecture

United States Digital Space LLC • San Francisco (CA)

On-site
USD 240,000 - 360,000
ML Inference Infrastructure Engineer
ML Inference Infrastructure Engineer

Baseten • United States

Remote
USD 120,000 - 190,000
Equity
Medical coverage for employee and dep.
Flexible PTO including Winter Break
+3
Staff Engineer, LLM Inference & Infra
Staff Engineer, LLM Inference & Infra

Prime Intellect • United States

Hybrid
USD 120,000 - 150,000
Competitive compensation
Flexible work arrangement
Full visa sponsorship
+2
Principal LLM Inference Runtime Architect
Principal LLM Inference Runtime Architect

Hewlett Packard Enterprise • Spring (TX)

Hybrid
USD 152,000 - 349,000
LLM Platform Lead — Scalable Training & Inference
LLM Platform Lead — Scalable Training & Inference

Scale AI, Inc. • New York (NY)

On-site
USD 264,800 - 331,000
Comprehensive health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2