Senior Distributed Systems Engineer — AI Inference Platform

Baseten

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
Comprehensive medical for US
Flexible PTO
Parental leave
Fertility stipend
401(k)
Learning opportunities

Job summary

Baseten powers mission-critical AI inference at scale. We seek distributed systems engineers to own the runtimes and APIs that run large-scale LLM inference and hosted endpoints.

You’ll work across the stack from developer tooling to orchestration on Kubernetes, driving speed, reliability, and cost-efficiency for customers. You’ll collaborate with the team on model serving, inference performance, and API design while shipping robust, production-grade infrastructure and tools for a rapidly

Qualifications

  • 3+ years building and operating distributed systems, backend infrastructure, or large-scale APIs.
  • “Own low-latency, reliable backend services including rate limiting, auth, quotas, metering, migrations.”
  • Experience with performance profiling, tracing, capacity planning, and SLO management.
  • Strong written communication and collaboration across teams.
  • Interest in inference engineering; ML or model serving experience is a plus.

Responsibilities

  • Build infrastructure and orchestration systems for large-scale distributed LLM inference.
  • Design, build, and operate Model APIs with JSON outputs and tool calling.
  • Implement platform fundamentals: API versioning, validation, metering, quotas, authentication.
  • Instrument observability and run benchmarks for speed, reliability, and quality.
  • Debug and harden production systems spanning Kubernetes, runtimes, networking, and GPUs.
  • Collaborate with Inference Performance engineers to broaden optimization reach.
  • Own projects end to end from architecture to deployment and iteration with customers.

Skills

Distributed systems
Backend infrastructure
Low latency
Developer experience
Code ownership
Observability
Collaboration
Performance tuning

Education

Bachelor's/Master's/Ph.D. in CS or related

Tools

Kubernetes
API gateways
Observability tooling
CI/CD systems

Job description

Baseten powers mission-critical AI inference at scale. We seek distributed systems engineers to own the runtimes and APIs that run large-scale LLM inference and hosted endpoints.

You’ll work across the stack from developer tooling to orchestration on Kubernetes, driving speed, reliability, and cost-efficiency for customers. You’ll collaborate with the team on model serving, inference performance, and API design while shipping robust, production-grade infrastructure and tools for a rapidly

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Distributed Systems Engineer — AI Inference Platform
Distributed Systems Engineer — AI Inference Platform

Emploive • New York (NY)

On-site
USD 180,000 - 360,000
Competitive compensation with equity
100% healthcare coverage for employee
Flexible PTO including Winter Break
+4
Distributed Systems Engineer for AI Inference Platform
Distributed Systems Engineer for AI Inference Platform

Baseten • United States

Remote
USD 180,000 - 240,000
Distributed Systems Engineer for LLM Inference Platform
Distributed Systems Engineer for LLM Inference Platform

Baseten • San Francisco (CA)

On-site
USD 180,000 - 360,000
Competitive equity
Medical insurance
Flexible PTO
+4
Software Engineer- Inference Platform
Software Engineer- Inference Platform

Baseten • United States

Remote
USD 180,000 - 240,000
AI Inference Platform Engineer — Production ML Tools
AI Inference Platform Engineer — Production ML Tools

Baseten • United States

Remote
USD 140,000 - 210,000
Equity
Medical insurance
Dental & Vision
+4
Staff Inference Engineer — Production-Scale AI Platform
Staff Inference Engineer — Production-Scale AI Platform

Designworks Talent LLC • Bellevue (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
401(k) plan with company match
Paid holidays
Senior ML Systems Engineer - Scalable Inference
Senior ML Systems Engineer - Scalable Inference

Atlassian • Seattle (WA)

Hybrid
USD 206,000 - 269,000
Health and wellbeing resources
Paid volunteer days
Senior ML Inference Platform Engineer
Senior ML Inference Platform Engineer

Atlassian • Austin (TX)

Hybrid
USD 206,000 - 269,000
Health and wellbeing resources
Paid volunteer days
Senior Systems Engineer, AI Inference Platform
Senior Systems Engineer, AI Inference Platform

OpenAI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior Inference Engineer, AI Infrastructure & Production
Senior Inference Engineer, AI Infrastructure & Production

Hamilton Barnes • United States

On-site
USD 225,000 - 275,000
Full Benefits