Distributed Systems Engineer for LLM Inference Platform

Baseten

San Francisco (CA)

On-site

USD 180,000 - 360,000

Full time

27 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Competitive equity
Medical insurance
Flexible PTO
Parental leave
Fertility stipend
401(k)
Learning opportunities

Job summary

Baseten, a leader in distributed AI infrastructure, seeks engineers to build the runtime powering large-scale LLM inference and hosted endpoints. You’ll impact model deployment, tooling, and performance across the platform.

You’ll own end-to-end projects from architecture to deployment, balancing performance with reliability and developer experience, while collaborating across teams on challenging integration tasks.

Qualifications

  • Bachelor’s, Master’s, or Ph.D. in Computer Science, Engineering, or related field, or equivalent practical experience.
  • 3+ years building and operating distributed systems, backend infrastructure, or large-scale APIs.

Responsibilities

  • Build infrastructure and orchestration systems for large-scale distributed LLM inference, including routing, autoscaling, scheduling, and runtime management.
  • Design, build, and operate Model APIs with advanced inference capabilities: structured outputs, tool/function calling, multimodal serving.
  • Implement platform fundamentals such as API versioning, validation, usage metering, quotas, and authentication.

Skills

Distributed systems
Backend services
Low-latency design
Dev experience focus
Collaboration

Education

Bachelor's/Master/PhD in CS/Engineering
Equivalent practical experience

Tools

Kubernetes
APIs & gateways
CI/CD

Job description

Baseten, a leader in distributed AI infrastructure, seeks engineers to build the runtime powering large-scale LLM inference and hosted endpoints. You’ll impact model deployment, tooling, and performance across the platform.

You’ll own end-to-end projects from architecture to deployment, balancing performance with reliability and developer experience, while collaborating across teams on challenging integration tasks.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Distributed Systems Engineer — AI Inference Platform
Distributed Systems Engineer — AI Inference Platform

Emploive • New York (NY)

On-site
USD 180,000 - 360,000
Competitive compensation with equity
100% healthcare coverage for employee
Flexible PTO including Winter Break
+4
Distributed Systems Engineer for AI Inference Platform
Distributed Systems Engineer for AI Inference Platform

Baseten • United States

Remote
USD 180,000 - 240,000
Senior Distributed Systems Engineer — AI Inference Platform
Senior Distributed Systems Engineer — AI Inference Platform

Baseten • San Francisco (CA)

On-site
USD 180,000 - 240,000
Equity
Comprehensive medical for US
Flexible PTO
+4
Software Engineer- Inference Platform
Software Engineer- Inference Platform

Baseten • United States

Remote
USD 180,000 - 240,000
LLM Inference Performance Engineer
LLM Inference Performance Engineer

Emploive • New York (NY)

On-site
USD 180,000 - 360,000
Equity
Insurance for dependents
Winter Break
ML Platform Engineer: Scale AI & Inference
ML Platform Engineer: Scale AI & Inference

Apply • San Francisco (CA)

Hybrid
USD 245,000 - 345,000
Flexible Time Off
Health Insurance
Work From Home Allowance
+2
Distributed LLM Inference Engineer
Distributed LLM Inference Engineer

Intel Corporation • Prescott (AZ)

On-site
USD 52,000 - 82,000
Staff Engineer, LLM Inference & Infra
Staff Engineer, LLM Inference & Infra

Prime Intellect • United States

Hybrid
USD 120,000 - 150,000
Competitive compensation
Flexible work arrangement
Full visa sponsorship
+2
Production-Grade LLM Inference Runtime Engineer
Production-Grade LLM Inference Runtime Engineer

AI Chopping Block, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 360,000
Senior LLM Inference Systems Engineer
Senior LLM Inference Systems Engineer

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Medical insurance
Bonus & equity plan
401(k) retirement plan
+1