AI Inference Systems Engineer

Engg

San Francisco (CA)

On-site

USD 150,000 - 190,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Unlimited PTO
Medical, dental, and vision benefits
Life insurance and disabili

Job summary

Sierra’s AI agents rely on foundation models; as a Software Engineer on Inference, you’ll help define the inference architecture across self-hosted models and third‑party providers. You’ll work on systems for serving and routing inference, managing capacity and quota, and optimizing latency, reliability, and cost.

This is a systems‑first role at the intersection of distributed infrastructure and AI. You don’t need to be an ML researcher—but you should love solving complex problems and applying

Qualifications

  • Deep systems thinking and strong distributed systems fundamentals.
  • Experience designing, building, and operating large‑scale production systems.
  • Strong judgment around tradeoffs involving latency, reliability, capacity, and cost.
  • Experience taking ownership of complex infrastructure from architecture through production operation.
  • Excitement about applying systems expertise to AI infrastructure and learning quickly as the underlying technology evolves.

Responsibilities

  • Partner with frontier labs and providers. At our scale, we rely on frontier labs, and inference providers to supply capacity, training and inference infrastructure.
  • Shape Sierra’s inference architecture. Design how inference traffic flows across models, infrastructure, and providers, including new serving and proxy layers as Sierra scales.
  • Build for low latency and high reliability. Develop systems for routing, failover, capacity management, and quota that keep inference performant and available across large‑scale production workloads.
  • Build and operate self-hosted inference. Run models on GPU infrastructure, from building containers and operating inference engines to managing the underlying compute capacity.
  • Optimize inference performance. Work with the Applied Research team on techniques such as speculative decoding and serving‑engine optimizations that improve latency, throughput, and cost.
  • Build across a hybrid inference stack. Work with both Sierra‑managed infrastructure and leading inference platforms, making architectural decisions about where and how workloads should run.
  • Push the serving stack forward. Work closely with inference providers to tune engines and infrastructure for Sierra’s workloads.
  • Support the broader model lifecycle. Contribute to infrastructure that enables post‑training while partnering closely with our Models and Agent Runtime teams.

Skills

Distributed systems
Self-hosted inference
Latency optimization
GPU infrastructure

Job description

Sierra’s AI agents rely on foundation models; as a Software Engineer on Inference, you’ll help define the inference architecture across self-hosted models and third‑party providers. You’ll work on systems for serving and routing inference, managing capacity and quota, and optimizing latency, reliability, and cost.

This is a systems‑first role at the intersection of distributed infrastructure and AI. You don’t need to be an ML researcher—but you should love solving complex problems and applying

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Inference Systems Engineer: Scalable AI Infrastructure
Inference Systems Engineer: Scalable AI Infrastructure

Candidate • San Francisco (CA)

On-site
USD 180,000 - 240,000
Flexible PTO
Medical, dental, vision benefits
Life insurance and disability benefits
+4
Senior AI Inference Systems Engineer (Equity)
Senior AI Inference Systems Engineer (Equity)

Emploive • San Francisco (CA)

On-site
USD 230,000 - 390,000
Flexible PTO
Medical, dental, vision benefits
Retirement plan
+3
AI Inference Infrastructure Engineer
AI Inference Infrastructure Engineer

Scaled Cognition • New York (NY)

On-site
USD 120,000 - 170,000
AI Inference Engineer
AI Inference Engineer

Evergrid AI, Inc. • New York (NY)

On-site
Confidential
Inference Systems Engineer — High-Performance AI Serving
Inference Systems Engineer — High-Performance AI Serving

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Software Engineer, Inference
Software Engineer, Inference

Engg • San Francisco (CA)

On-site
USD 150,000 - 190,000
Unlimited PTO
Medical, dental, and vision benefits
Life insurance and disabili
Senior AI Engineer: Real-Time Inference & Agent Systems
Senior AI Engineer: Real-Time Inference & Agent Systems

Arcana Analytics • United States

On-site
USD 120,000 - 160,000
Senior AI Inference Infrastructure Engineer
Senior AI Inference Infrastructure Engineer

Modular • United States

Hybrid
USD 167,000 - 273,000
Competitive salary
Premier insurance plans
Flexible paid time off
+1
AI Agent Engineer — Production Systems & ADLC Lead
AI Agent Engineer — Production Systems & ADLC Lead

Sierra • New York (NY), Northern (KY)

Hybrid
USD 130,000 - 180,000
Flexible (unlimited) paid time off
Medical, dental, and vision benefits
Retirement plan
+6
Production AI Agent Engineer
Production AI Agent Engineer

Sierra • New York (NY), Northern (KY)

Hybrid
USD 150,000 - 230,000
Flexible paid time off
Medical, dental, and vision benefits
Retirement plan