Senior ML Infrastructure Engineer

SambaNova

San Jose (CA)

On-site

USD 180,000 - 240,000

Full time

11 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Medical Insurance
Health Savings Account (HSA)
Dental Insurance
Vision Insurance
Short Term Disability
Long Term Disability
Basic Life Insurance
Voluntary Life Insurance
AD&D Insurance
Flexible Spending Account (FSA)
Headspace Subscription
Gympass+ Membership
One Medical Membership
Employee Assistance Program
Equity

Job summary

SambaNova seeks an engineer to design and operate production-grade inference infrastructure on RDU architecture, optimizing performance and cost while sustaining reliability.

You will own the public API surface and build accuracy verification infrastructure to ensure high-throughput services, with a strong emphasis on scalable distributed systems and modern LLM serving techniques.

Qualifications

  • Requires a B.S. in Computer Science or a related field and 5+ years of experience building large-scale distributed systems.
  • Must be proficient in Python and familiar with modern LLM inference techniques and serving stacks like vLLM.

Responsibilities

  • Design and operate production-grade inference infrastructure on RDU architecture to optimize performance and cost.
  • Own the public API surface and build accuracy verification infrastructure to ensure reliable, high-throughput services.

Skills

Distributed Systems
ML Serving
Python
Systems Design
LLM Inference
vLLM
Algorithms
Data Structures
Concurrency
API Design
Speculative Decoding
Constrained Decoding
Prompt Caching
Long-context Inference
TensorRT-LLM
SGLang

Education

B.S. in Computer Science or related field

Job description

Design and operate production-grade inference infrastructure on RDU architecture to optimize performance and cost. Own the public API surface and build accuracy verification infrastructure to ensure reliable, high-throughput services.

Requirements: Requires a B.S. in Computer Science or a related field and 5+ years of experience building large-scale distributed systems. Must be proficient in Python and familiar with modern LLM inference techniques and serving stacks like vLLM.

Key Skills: Distributed Systems, ML Serving, Python, Systems Design, LLM Inference, vLLM, Algorithms, Data Structures, Concurrency, API Design, Speculative Decoding, Constrained Decoding, Prompt Caching, Long-context Inference, TensorRT-LLM, SGLang

Benefits: Medical Insurance, Health Savings Account (HSA), Dental Insurance, Vision Insurance, Short Term Disability, Long Term Disability, Basic Life Insurance, Voluntary Life Insurance, AD&D Insurance, Flexible Spending Account (FSA), Headspace Subscription, Gympass+ Membership, One Medical Membership, Employee Assistance Program, Equity

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Infrastructure Engineer
Senior ML Infrastructure Engineer

SambaNova • Austin (TX)

On-site
USD 170,000 - 240,000
Medical Insurance
Dependent Insurance
Health Savings Account (HSA)
+13
Staff Software Engineer, AI Inference
Staff Software Engineer, AI Inference

ChatGPT Jobs • New York (NY)

On-site
USD 180,000 - 240,000
Health Insurance
Equity
Member of Technical Staff, Inference & Serving
Member of Technical Staff, Inference & Serving

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Software Engineer- Foundation Model Inference
Staff Software Engineer- Foundation Model Inference

DevHub • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual performance bonus
Equity
Member of Technical Staff, MLSys
Member of Technical Staff, MLSys

Bake AI • San Mateo (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Software Engineer – Inference (SWE1) [D.26.0171]
Software Engineer – Inference (SWE1) [D.26.0171]

Dover Networks LLC • Maryland

On-site
USD 204,000 - 222,000
Software Engineer, Inference Runtime
Software Engineer, Inference Runtime

EngRadar • New York (NY)

On-site
USD 150,000 - 230,000
Equity grants
Medical plan
Vision plan
+5
Member of Technical Staff, Backend, LLM Applications
Member of Technical Staff, Backend, LLM Applications

Inception • San Francisco (CA)

On-site
USD 150,000 - 230,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
Infrastructure Engineer, LLM Inference Optimization
Infrastructure Engineer, LLM Inference Optimization

GMI Cloud • Mountain View (CA)

On-site
USD 170,000 - 230,000