Staff Software Engineer- Foundation Model Inference

United States Digital Space LLC

San Francisco (CA)

On-site

USD 190,000 - 265,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

United States Digital Space LLC is seeking seasoned backend/infrastructure engineers to build and optimize LLM inference platforms powering enterprise-grade AI workloads. You will tackle reliability, latency, and scalability challenges across distributed systems and cloud-native infrastructure.

Join a high-visibility Foundation Model Inference team building enterprise-ready APIs with governance and scalable deployment for model serving, training, and vector search.

Qualifications

  • 8+ years of experience in backend or infrastructure engineering.
  • Experience with distributed systems, scalable APIs, or cloud-native infrastructure.
  • Experience with real-time serving, ML infrastructure, or GPU orchestration.
  • Familiarity with service-oriented architecture, deployment pipelines, and system observability.

Responsibilities

  • Build LLM infrastructure powering large-scale inference workloads for customers through partner models (OpenAI, Anthropic, Gemini) and self-hosted models (Qwen, GPT-OSS, Llama).
  • Improve reliability, latency, and efficiency of distributed AI workloads.
  • Collaborate with platform, infra, and ML teams to deliver seamless end-to-end experiences.
  • Shape how developers and data scientists build and interact with AI on the company.

Skills

Distributed systems
Cloud-native infrastructure
Real-time serving
Observability

Tools

SageMaker
Vertex AI
Azure ML
MLflow
PyTorch
Ray
vLLM
SGLang

Job description

P-1930

At the company, we are passionate about enabling data and AI teams to solve the world's toughest problems — from making the next mode of transportation a reality to accelerating the development of medical breakthroughs. We do this by building and running the world's best data and AI infrastructure platform so our customers can use deep data insights to improve their business. Founded by engineers — and customer-obsessed — we leap at every opportunity to solve technical challenges, from designing next-gen UI/UX for interfacing with data to scaling our services and infrastructure across millions of virtual machines. And we're only getting started.

As part of the AI team, you'll build the platforms and products that power everything from data apps, AI agents, model training, model serving, and Vector Search. You'll be joining a high-agency, high-visibility team operating at the frontier of AI infrastructure — with deep ties to research, product, and real-world enterprise use cases. the company Mosaic AI is one of our fastest-growing businesses, helping thousands of our customers democratize AI within their organizations. We're building the products and infrastructure that power the next generation of AI.

The Foundation Model Inference team is the backbone of the company’ generative AI capabilities. We build the infrastructure that enables our customers to serve, scale, and optimize frontier models with enterprise-grade reliability and performance. Our Foundation Model APIs provide a unified platform that gives customers access to LLMs with the governance, flexibility, and scalability required for enterprise production workloads.

We are looking for high-agency engineers who are excited to work on powering model inference at enterprise scale.

The impact you will have:
  • Build LLM infrastructure powering large-scale inference workloads for customers through partner models (OpenAI, Anthropic, Gemini) and self-hosted models (Qwen, GPT-OSS, Llama)
  • Improve reliability, latency, and efficiency of distributed AI workloads
  • Collaborate with platform, infra, and ML teams to deliver seamless end-to-end experiences
  • Shape how developers and data scientists build and interact with AI on the company
What we look for:
  • 8+ years of experience in backend or infrastructure engineering
  • Experience with distributed systems, scalable APIs, or cloud-native infrastructure
  • Experience with real-time serving, ML infrastructure, or GPU orchestration
  • Familiarity with service-oriented architecture, deployment pipelines, and system observability
Bonus points for:
  • Exposure to platforms like SageMaker, Vertex AI, or Azure ML
  • Contributions to OSS projects like MLflow, PyTorch, Ray, vLLM, SGLang
  • Built developer platforms or internal tools supporting AI workflows

Pay Range Transparency

the company is committed to fair and equitable compensation practices. The pay range(s) for this role is listed below and represents the expected salary range for non-commissionable roles or on-target earnings for commissionable roles. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to job-related skills, depth of experience, relevant certifications and training, and specific work location. Based on the factors above, the company anticipates utilizing the full width of the range. The total compensation package for this position may also include eligibility for annual performance bonus, equity, and the benefits listed above. For more information regarding which range your location is in visit our page here.

Local Pay Range
$190,000 — $265,000 USD
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer, Foundation Model API
Staff Software Engineer, Foundation Model API

United States Digital Space LLC • San Francisco (CA)

On-site
USD 190,000 - 265,000
Engineering Manager, Foundation Model Inference (FMAPI)
Engineering Manager, Foundation Model Inference (FMAPI)

United States Digital Space LLC • Mountain View (CA)

On-site
USD 190,000 - 262,000
Staff Software Engineer- Foundation Model Inference San Francisco, California
Staff Software Engineer- Foundation Model Inference San Francisco, California

Databricks Inc. • San Francisco (CA)

On-site
USD 190,000 - 265,000
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
Sr. Software Engineer- Backend
Sr. Software Engineer- Backend

United States Digital Space LLC • New York (NY)

On-site
USD 165,000 - 220,000
Engineering Manager, Foundation Model Inference (FMAPI)
Engineering Manager, Foundation Model Inference (FMAPI)

Databricks • Mountain View (CA)

On-site
USD 190,000 - 261,000
Staff Software Engineer- Foundation Model Inference
Staff Software Engineer- Foundation Model Inference

Databricks • San Francisco (CA)

On-site
USD 190,000 - 265,000
Software Engineer, Model Inference
Software Engineer, Model Inference

OpenAI • San Francisco (CA)

On-site
USD 325,000 - 490,000
AI Infrastructure Engineer, Model Serving Platform
AI Infrastructure Engineer, Model Serving Platform

Scale AI, Inc. • New York (NY)

On-site
USD 180,000 - 225,000
Comprehensive health coverage
Equity compensation
Learning and development stipend
+2
Software Engineer, LLM Infrastructure
Software Engineer, LLM Infrastructure

Fireworks AI • New York (NY)

On-site
USD 175,000 - 220,000
Equity
Competitive salary
Comprehensive benefits