LLM Inference Platform Engineer

Baseten

New York (NY)

On-site

USD 216,000 - 360,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Comprehensive health coverage
Flexible PTO
Parental leave
Fertility stipend
401(k)
Learning opportunities

Job summary

Baseten in the United States is hiring a Software Engineer for the Inference Stack to build distributed runtime for large‑scale LLM inference. You’ll own production systems spanning Kubernetes, routing, autoscaling, observability, and deployment tooling, with a focus on reliability and developer experience.

You’ll collaborate across teams to push performance improvements, scale, and user‑facing ease‑of‑use features, joining a fast‑moving platform engineering group at the forefront of AI

Qualifications

  • Strong background in distributed systems, backend infrastructure, or platform engineering.
  • Experience building and operating production systems where reliability, latency, and scale are first‑class concerns.
  • Excellent communication and collaboration skills.

Responsibilities

  • Develop infrastructure and orchestration systems for deploying and managing large‑scale distributed LLM inference
  • Work across the stack, from customer‑facing features to low‑level infrastructure components
  • Build platform capabilities related to routing, autoscaling, scheduling, observability, and runtime management
  • Improve the reliability, scalability, and usability of our inference stack
  • Collaborate with Model Performance engineers to make inference optimizations broadly available
  • Define best practices around testing, release automation, benchmarking, and operational excellence
  • Debug complex production systems spanning Kubernetes, distributed runtimes, networking, and GPU workloads
  • Make thoughtful engineering tradeoffs balancing performance, reliability, operational simplicity, and developer experience
  • Own projects end‑to‑end: architecture, implementation, deployment, monitoring, iteration based on customer feedback

Skills

Distributed systems
Backend infrastructure
Platform engineering
Developer experience
Debugging complex systems

Education

Bachelor's, Master's, or Ph.D. in CS/Engineering

Tools

Kubernetes

Job description

Baseten in the United States is hiring a Software Engineer for the Inference Stack to build distributed runtime for large‑scale LLM inference. You’ll own production systems spanning Kubernetes, routing, autoscaling, observability, and deployment tooling, with a focus on reliability and developer experience.

You’ll collaborate across teams to push performance improvements, scale, and user‑facing ease‑of‑use features, joining a fast‑moving platform engineering group at the forefront of AI

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Inference Stack (LLM Infra)
Software Engineer, Inference Stack (LLM Infra)

Baseten • United States

Remote
USD 180,000 - 240,000
Software Engineer — Inference Stack for Scalable LLMs
Software Engineer — Inference Stack for Scalable LLMs

BaseTen • New York (NY), San Francisco (CA)

On-site
USD 180,000 - 240,000
Equity
Health insurance
Flexible PTO
+2
Edge-to-Cloud LLM Inference Engineer
Edge-to-Cloud LLM Inference Engineer

Insider, Inc. • United States

On-site
USD 100,000 - 140,000
Staff Engineer, LLM Inference & Infra
Staff Engineer, LLM Inference & Infra

Prime Intellect • United States

Hybrid
USD 120,000 - 150,000
Competitive compensation
Flexible work arrangement
Full visa sponsorship
+2
LLM Platform Lead — Scalable Training & Inference
LLM Platform Lead — Scalable Training & Inference

Scale AI, Inc. • New York (NY)

On-site
USD 264,000 - 331,000
Comprehensive health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
ML Platform Engineer: Scale AI & Inference
ML Platform Engineer: Scale AI & Inference

Apply • San Francisco (CA)

Hybrid
USD 245,000 - 345,000
Flexible Time Off
Health Insurance
Work From Home Allowance
+2
Staff Research Engineer - LLM Inference & Serving
Staff Research Engineer - LLM Inference & Serving

Modal Labs • New York (NY)

On-site
USD 180,000 - 240,000
Senior LLM Inference Engineer - End-to-End, Equity
Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA
Engineering Manager, Forward Deployed AI & LLM Inference
Engineering Manager, Forward Deployed AI & LLM Inference

BaseTen • New York (NY), San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive equity
Medical, dental, vision coverage
Flexible PTO
+3
Senior Backend Engineer, LLM Inference Systems
Senior Backend Engineer, LLM Inference Systems

Inception • San Francisco (CA)

On-site
USD 150,000 - 230,000