Member of Technical Staff, Backend, LLM Applications

Inception

San Francisco (CA)

On-site

USD 150,000 - 230,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Inception in San Francisco is seeking experienced backend engineers to own the systems that serve our diffusion LLMs in production. You will build and operate infrastructure that handles billions of inference requests, optimizing for latency, throughput, cost, and reliability.

This role sits at the intersection of ML systems and backend infrastructure, with responsibilities spanning scalable services, model serving, load balancing, canary deployments, and observability tooling to ensure SLA

Qualifications

  • BS/MS/PhD in Computer Science or a related field (or equivalent experience).
  • 5+ years of experience building production backend systems.
  • Strong proficiency in Python, including async programming and concurrent systems.
  • Solid understanding of distributed systems, networking, and load balancing at scale.
  • Familiarity with Kubernetes, CI/CD pipelines, and cloud infra (AWS and/or Azure).

Responsibilities

  • Design, build, and operate scalable backend services and model serving infrastructure for our diffusion LLMs.
  • Implement and manage load balancing, autoscaling, and traffic routing for model endpoints.
  • Build systems for model versioning, canary deployments, and zero-downtime rollouts.
  • Develop monitoring, alerting, and observability tooling to ensure SLA compliance and rapid incident response.
  • Benchmark and evaluate serving frameworks and hardware configurations to inform infrastructure decisions.

Skills

Python (async programming)
Distributed systems
Load balancing at scale
Networking fundamentals
Backend software design

Education

BS/MS/PhD in Computer Science or related field

Tools

Kubernetes
CI/CD pipelines
AWS / Azure cloud infra
Terraform
Prometheus
Grafana
vLLM / Triton / TensorRT-LLM

Job description

The Role

We seek experienced backend engineers to own the systems that serve our diffusion LLMs in production. You\'ll build and operate infrastructure that handles billions of inference requests — optimizing for latency, throughput, cost, and reliability. This role sits at the intersection of ML systems and backend infrastructure.


Key Responsibilities


  • Design, build, and operate scalable backend services and model serving infrastructure for our diffusion LLMs.

  • Implement and manage load balancing, autoscaling, and traffic routing for model endpoints.

  • Build systems for model versioning, canary deployments, and zero-downtime rollouts.

  • Develop monitoring, alerting, and observability tooling to ensure SLA compliance and rapid incident response.

  • Benchmark and evaluate serving frameworks and hardware configurations to inform infrastructure decisions.


Qualifications


  • BS/MS/PhD in Computer Science or a related field (or equivalent experience).

  • 5+ years of experience building production backend systems.

  • Strong proficiency in Python, including async programming and concurrent systems.

  • Solid understanding of distributed systems, networking, and load balancing at scale.

  • Familiarity with Kubernetes, CI/CD pipelines, and cloud infra (AWS and/or Azure).


Preferred Skills


  • Experience serving LLMs or other large generative models in production at scale.

  • Experience with cloud infrastructure (AWS, Azure), including GPU instance management and cost optimization.

  • Experience with infrastructure as code tools (Terraform) and deployment automation.

  • Experience with monitoring and observability tools (Prometheus, Grafana).

  • Familiarity with model serving frameworks (vLLM, Triton Inference Server, TensorRT-LLM).

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Inference & Serving
Member of Technical Staff, Inference & Serving

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Full Stack, LLM Applications
Member of Technical Staff, Full Stack, LLM Applications

Inception • San Francisco (CA)

On-site
USD 170,000 - 210,000
Senior Backend Engineer, LLM Inference Systems
Senior Backend Engineer, LLM Inference Systems

Inception • San Francisco (CA)

On-site
USD 150,000 - 230,000
Staff Software Engineer- Foundation Model Inference
Staff Software Engineer- Foundation Model Inference

DevHub • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual performance bonus
Equity
AI Developer
AI Developer

Salvo Software LLC • Northern (KY)

Hybrid
USD 120,000 - 190,000
Member of Technical Staff, Training Infra
Member of Technical Staff, Training Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff ML Systems Engineer - Diffusion LLM Serving
Staff ML Systems Engineer - Diffusion LLM Serving

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior/Principal Local LLM & Generative AI Platform Engineer
Senior/Principal Local LLM & Generative AI Platform Engineer

Parallel Wireless • Northern (KY)

Hybrid
USD 150,000 - 210,000
Software Engineer, ML Infrastructure
Software Engineer, ML Infrastructure

Realmlabs • Sunnyvale (CA)

On-site
USD 210,000 - 350,000
Market aligned compensation
Founding engineer equity
Medical, Dental, Vision, and Life insurance
+2
ML Engineer (LLM Systems)
ML Engineer (LLM Systems)

Cynnovative • Arlington (VA)

On-site
USD 120,000 - 150,000