Inference Systems Engineer: Scalable AI Infrastructure

Candidate

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Flexible PTO
Medical, dental, vision benefits
Life insurance and disability benefits
401(k) / Retirement plan
Parental leave
Fertility and family building benefits
Snacks and coffee

Job summary

Sierra is the leading platform for customer-facing AI agents, based in San Francisco. We are seeking a Software Engineer on the Inference team to define architectures for self-hosted and provider-based inference, focusing on latency, reliability, and cost at scale.

You will work across distributed systems to build, optimize, and operate core inference infrastructure. This role emphasizes ownership of complex systems from design to production, and collaboration with frontier labs, providers, and

Qualifications

  • Deep systems thinking and distributed systems fundamentals.
  • Experience designing, building, and operating large-scale production systems.
  • Strong judgment around latency, reliability, capacity, and cost.
  • Experience taking ownership of complex infrastructure from architecture through production operation.
  • Excitement about applying systems expertise to AI infrastructure and evolving tech.

Responsibilities

  • Partner with frontier labs and providers to supply capacity and infrastructure.
  • Shape Sierra’s inference architecture across models, infrastructure, and providers.
  • Build for low latency and high reliability in routing, capacity management, and quotas.
  • Build and operate self-hosted inference on GPU infrastructure and containers.
  • Collaborate with Applied Research on latency and cost optimizations.
  • Work across a hybrid inference stack with Sierra-managed and external platforms.
  • Push the serving stack forward with inference providers to tune engines.
  • Support the broader model lifecycle with infrastructure and cross-team collaboration.

Skills

Deep systems thinking
Distributed systems
Production systems
Latency and reliability tradeoffs
Ownership of infrastructure
AI infrastructure

Tools

vLLM
SGLang

Job description

Sierra is the leading platform for customer-facing AI agents, based in San Francisco. We are seeking a Software Engineer on the Inference team to define architectures for self-hosted and provider-based inference, focusing on latency, reliability, and cost at scale.

You will work across distributed systems to build, optimize, and operate core inference infrastructure. This role emphasizes ownership of complex systems from design to production, and collaboration with frontier labs, providers, and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Inference Systems Engineer (Equity)
Senior AI Inference Systems Engineer (Equity)

Emploive • San Francisco (CA)

On-site
USD 230,000 - 390,000
Flexible PTO
Medical, dental, vision benefits
Retirement plan
+3
AI Inference Systems Engineer
AI Inference Systems Engineer

Engg • San Francisco (CA)

On-site
USD 150,000 - 190,000
Unlimited PTO
Medical, dental, and vision benefits
Life insurance and disabili
Staff Engineer, Scalable AI Inference Infrastructure
Staff Engineer, Scalable AI Inference Infrastructure

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Generous health, dental, and vision benefits
401(k) company match
Equity options
Senior Systems Engineer, AI Inference Platform
Senior Systems Engineer, AI Inference Platform

Slope • San Francisco (CA)

On-site
USD 180,000 - 260,000
Inference Cloud Architect: Build Scalable AI Infrastructure
Inference Cloud Architect: Build Scalable AI Infrastructure

techire.® • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Inference Infrastructure Engineer
AI Inference Infrastructure Engineer

Scaled Cognition • New York (NY)

On-site
USD 120,000 - 170,000
AI Agent Engineer — Production Systems & ADLC Lead
AI Agent Engineer — Production Systems & ADLC Lead

Sierra • New York (NY), Northern (KY)

Hybrid
USD 130,000 - 180,000
Flexible (unlimited) paid time off
Medical, dental, and vision benefits
Retirement plan
+6
Software Engineer, Inference
Software Engineer, Inference

Engg • San Francisco (CA)

On-site
USD 150,000 - 190,000
Unlimited PTO
Medical, dental, and vision benefits
Life insurance and disabili
Senior Inference Systems Engineer for High-Performance AI
Senior Inference Systems Engineer for High-Performance AI

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
AI Systems Engineer: Production-Grade Agentic Workflows
AI Systems Engineer: Production-Grade Agentic Workflows

Sierracorp • San Francisco (CA)

On-site
USD 100,000 - 150,000