Lead, End-to-End Inference Performance

SambaNovaSystems

San Jose (CA)

On-site

USD 245,000 - 325,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

SambaNova Systems is seeking an Architect for the Inference Systems Performance team to own end-to-end performance of large-scale LLM inference. You will study how requests move from tokenization to decode and how a deployment is sized to meet customer SLOs.

This role spans workload capture, benchmarking, and performance modeling to guide capacity planning and system design. You'll define strategies, build benchmarking capabilities, and develop tooling to profile a distributed inference

Qualifications

  • 12+ years of experience in performance engineering
  • Deep expertise in end-to-end performance analysis of distributed systems
  • Proven command of workload generation and simulation
  • Demonstrated ability to enter unfamiliar domains and apply core performance methods
  • Ability to lead cross-functional efforts and mentor senior engineers
  • Experience representing an organization credibly to customers and partners
  • Track record of scoping and delivering high-complexity work with impact

Responsibilities

  • Define and drive the technical strategy for inference-systems performance including workload capture, benchmarking, modeling, and simulation
  • Build workload-capture and agentic-benchmarking capability to faithfully reflect production traffic
  • Own the performance-modeling and simulation practice to predict how configurations affect output and planning
  • Drive end-to-end profiling tooling to produce actionable profiles of a distributed inference pipeline
  • Serve as the senior technical voice across model-optimization, systems, hardware, and product
  • Represent SambaNova's performance story to customers and partners
  • Mentor and multiply capability by raising engineers' skills and patterns
  • Drive resolution of novel, ambiguous challenges spanning organizational boundaries

Job description

SambaNova Systems is seeking an Architect for the Inference Systems Performance team to own end-to-end performance of large-scale LLM inference. You will study how requests move from tokenization to decode and how a deployment is sized to meet customer SLOs.

This role spans workload capture, benchmarking, and performance modeling to guide capacity planning and system design. You'll define strategies, build benchmarking capabilities, and develop tooling to profile a distributed inference

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Inference Systems Performance Architect
Senior Inference Systems Performance Architect

SambaNova • San Jose (CA)

On-site
USD 245,000 - 325,000
Health insurance
Gympass+
One Medical
AI Systems Performance Engineer - Scalable LLM Inference
AI Systems Performance Engineer - Scalable LLM Inference

SambaNova Systems • San Jose (CA)

On-site
USD 135,000 - 165,000
Equity
Health insurance
Well-being benefits
Senior ML Infra Engineer: High-Throughput AI Inference
Senior ML Infra Engineer: High-Throughput AI Inference

SambaNovaSystems • United States

On-site
USD 200,000 - 275,000
Health insurance
Health Savings Account (HSA)
Headspace subscription
+2
Senior LLM Inference Architect — Performance & Benchmarking
Senior LLM Inference Architect — Performance & Benchmarking

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity and benefits
Inference Systems Performance Architect
Inference Systems Performance Architect

SambaNovaSystems • San Jose (CA)

On-site
USD 245,000 - 325,000
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

SambaNova • San Jose (CA)

On-site
USD 180,000 - 240,000
LLM Inference Systems Performance Engineer
LLM Inference Systems Performance Engineer

3M HEALTHCARE • Austin (TX)

On-site
USD 150,000 - 210,000
Medical, dental, and vision coverage
Income protection benefits
Paid family leave
+1
Inference Systems Performance Architect
Inference Systems Performance Architect

SambaNova • San Jose (CA)

On-site
USD 245,000 - 325,000
Health insurance
Gympass+
One Medical
Senior AI Inference Platform Product Lead
Senior AI Inference Platform Product Lead

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 208,000 - 328,000
Equity
Comprehensive benefits
Inclusive culture
Senior LLM Inference Engineer - End-to-End, Equity
Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA