Senior Backend Engineer, LLM Inference Systems

Inception

San Francisco (CA)

On-site

USD 150,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Inception in San Francisco is seeking experienced backend engineers to own the systems that serve our diffusion LLMs in production. You will build and operate infrastructure that handles billions of inference requests, optimizing for latency, throughput, cost, and reliability.

This role sits at the intersection of ML systems and backend infrastructure, with responsibilities spanning scalable services, model serving, load balancing, canary deployments, and observability tooling to ensure SLA

Qualifications

  • BS/MS/PhD in Computer Science or a related field (or equivalent experience).
  • 5+ years of experience building production backend systems.
  • Strong proficiency in Python, including async programming and concurrent systems.
  • Solid understanding of distributed systems, networking, and load balancing at scale.
  • Familiarity with Kubernetes, CI/CD pipelines, and cloud infra (AWS and/or Azure).

Responsibilities

  • Design, build, and operate scalable backend services and model serving infrastructure for our diffusion LLMs.
  • Implement and manage load balancing, autoscaling, and traffic routing for model endpoints.
  • Build systems for model versioning, canary deployments, and zero-downtime rollouts.
  • Develop monitoring, alerting, and observability tooling to ensure SLA compliance and rapid incident response.
  • Benchmark and evaluate serving frameworks and hardware configurations to inform infrastructure decisions.

Skills

Python (async programming)
Distributed systems
Load balancing at scale
Networking fundamentals
Backend software design

Education

BS/MS/PhD in Computer Science or related field

Tools

Kubernetes
CI/CD pipelines
AWS / Azure cloud infra
Terraform
Prometheus
Grafana
vLLM / Triton / TensorRT-LLM

Job description

Inception in San Francisco is seeking experienced backend engineers to own the systems that serve our diffusion LLMs in production. You will build and operate infrastructure that handles billions of inference requests, optimizing for latency, throughput, cost, and reliability.

This role sits at the intersection of ML systems and backend infrastructure, with responsibilities spanning scalable services, model serving, load balancing, canary deployments, and observability tooling to ensure SLA

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff ML Systems Engineer - Diffusion LLM Serving
Staff ML Systems Engineer - Diffusion LLM Serving

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Backend, LLM Applications
Member of Technical Staff, Backend, LLM Applications

Inception • San Francisco (CA)

On-site
USD 150,000 - 230,000
Head of ML Systems & Inference
Head of ML Systems & Inference

Doist • San Francisco (CA)

On-site
USD 260,000 - 380,000
Health, dental, vision benefits
401(k) company match
Senior Backend Engineer — AI/LLM Platform & Scale (Hybrid SF)
Senior Backend Engineer — AI/LLM Platform & Scale (Hybrid SF)

SupportFinity™ • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
Competitive compensation
Equity compensation
Daily catered lunch
Staff Research Engineer - LLM Inference & Serving
Staff Research Engineer - LLM Inference & Serving

Modal Labs • New York (NY)

On-site
USD 180,000 - 240,000
Software Engineer, Inference Stack (LLM Infra)
Software Engineer, Inference Stack (LLM Infra)

Baseten • United States

Remote
USD 180,000 - 240,000
LLM Inference Platform Engineer
LLM Inference Platform Engineer

Baseten • New York (NY)

On-site
USD 216,000 - 360,000
Equity
Comprehensive health coverage
Flexible PTO
+4
Software Engineer — Inference Stack for Scalable LLMs
Software Engineer — Inference Stack for Scalable LLMs

BaseTen • New York (NY), San Francisco (CA)

On-site
USD 180,000 - 240,000
Equity
Health insurance
Flexible PTO
+2
Staff Engineer - LLM Inference & Serving at Scale
Staff Engineer - LLM Inference & Serving at Scale

Prime Intellect • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Staff Engineer - Customer-Facing AI Inference Infra
Staff Engineer - Customer-Facing AI Inference Infra

Simplify • San Francisco (CA)

On-site
USD 200,000 - 300,000
Housing stipend
Uber/Waymo rides