Distributed Systems Engineer — AI Inference Platform

Emploive

New York (NY)

On-site

USD 180,000 - 360,000

Full time

23 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive compensation with equity
100% healthcare coverage for employee
Flexible PTO including Winter Break
Paid parental leave
Fertility and family-building stipend
401(k) (U.S. only)
Exposure to ML startups for learning

Job summary

Baseten is hiring distributed systems engineers to build the runtime powering large-scale LLM inference. You’ll deploy and operate APIs and runtimes, focusing on performance, reliability, and cost efficiency across Kubernetes-based deployments.

You’ll own end-to-end projects, drive architectural decisions, and collaborate with performance engineers to deliver scalable solutions for modern AI workloads while enhancing developer experience and platform usability.

Qualifications

  • Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 3+ years building and operating distributed systems, backend infrastructure, or large-scale APIs where reliability, latency, and scale are first-class concerns.
  • A proven track record of owning low-latency, reliable backend services, including rate limiting, auth, quotas, metering, and migrations.
  • Infrastructure instincts with a feel for performance: profiling, tracing, capacity planning, and SLO management.
  • Comfort debugging performance and reliability issues across multiple layers of the stack, from application behavior down to runtime and infrastructure internals.
  • A strong sense of developer experience. You think about how systems are used, not just how they work.
  • Eagerness to learn new languages, frameworks, and systems, and a real interest in inference engineering. Prior ML or LLM experience isn't required, though experience with model serving or inference systems is a plus.
  • Excellent written communication and collaboration skills, including writing clear design docs and working across functions.

Responsibilities

  • Build the infrastructure and orchestration systems that deploy and run large-scale distributed LLM inference, including routing, autoscaling, scheduling, and runtime management.
  • Design, build, and operate Model APIs, with a focus on advanced inference capabilities: structured outputs (JSON mode, grammar-constrained generation), tool/function calling, and multimodal serving.
  • Implement platform fundamentals such as API versioning, validation, usage metering, quotas, and authentication.
  • Instrument deep observability (metrics, traces, logs) and build repeatable benchmarks for speed, reliability, and quality. Help set best practices for testing, release automation, and operational excellence.
  • Debug and harden complex production systems spanning Kubernetes, distributed runtimes, networking, and GPU workloads to improve reliability and scalability.
  • Partner with Inference Performance engineers and other teams to make new optimizations broadly available to customers and easy to configure.
  • Own projects end to end, from architecture through deployment, monitoring, and iteration on customer feedback. Along the way, make thoughtful tradeoffs between performance, reliability, operational simplicity, and developer experience.

Skills

Distributed systems
Backend infrastructure
APIs
Developer experience
Observability

Education

Bachelor's/Master's/Ph.D. in CS or Engineering

Tools

Kubernetes
API design

Job description

Baseten is hiring distributed systems engineers to build the runtime powering large-scale LLM inference. You’ll deploy and operate APIs and runtimes, focusing on performance, reliability, and cost efficiency across Kubernetes-based deployments.

You’ll own end-to-end projects, drive architectural decisions, and collaborate with performance engineers to deliver scalable solutions for modern AI workloads while enhancing developer experience and platform usability.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Distributed Systems Engineer for AI Inference Platform
Distributed Systems Engineer for AI Inference Platform

Baseten • United States

Remote
USD 180,000 - 240,000
LLM Inference Platform Engineer
LLM Inference Platform Engineer

Baseten • New York (NY)

On-site
USD 216,000 - 360,000
Equity
Comprehensive health coverage
Flexible PTO
+4
Software Engineer- Inference Platform
Software Engineer- Inference Platform

Baseten • United States

Remote
USD 180,000 - 240,000
AI Inference Platform Engineer — Production ML Tools
AI Inference Platform Engineer — Production ML Tools

Baseten • United States

Remote
USD 140,000 - 210,000
Equity
Medical insurance
Dental & Vision
+4
Customer-Facing AI Platform Engineer
Customer-Facing AI Platform Engineer

Baseten • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Equity
Medical, dental, vision coverage for +
Flexible PTO including Winter Break
+4
Distributed LLM Inference Engineer - Scale & Resilience
Distributed LLM Inference Engineer - Scale & Resilience

Baseten • San Francisco (CA)

On-site
USD 180,000 - 360,000
Equity
Health coverage
Flexible PTO
+4
AI Inference Platform Engineer
AI Inference Platform Engineer

BaseTen • New York (NY), San Francisco (CA)

On-site
USD 140,000 - 210,000
Equity
Medical, dental and vision insurance (
Flexible PTO including Winter Break
+4
Software Engineer, Inference Stack (LLM Infra)
Software Engineer, Inference Stack (LLM Infra)

Baseten • United States

Remote
USD 180,000 - 240,000
Software Engineer- Inference Platform Baseten · New York City, NY Full-time · Hybrid $180,000–360,000 12 minutes ago
Software Engineer- Inference Platform Baseten · New York City, NY Full-time · Hybrid $180,000–360,000 12 minutes ago

Emploive • New York (NY)

On-site
USD 180,000 - 360,000
Competitive compensation with equity
100% healthcare coverage for employee
Flexible PTO including Winter Break
+4
Scale AI Infra Engineer | Kubernetes & CI/CD
Scale AI Infra Engineer | Kubernetes & CI/CD

BaseTen • New York (NY), San Francisco (CA)

Hybrid
USD 160,000 - 210,000
Meaningful equity
100% medical, dental, vision coverage
Flexible PTO including Winter Break
+4