LLM Inference Systems Performance Architect

Doist

San Jose (CA)

On-site

USD 245,000 - 325,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health Insurance
Dental Insurance
Vision Insurance
Disability Insurance
Life Insurance
AD&D Insurance
HSA with employer contribution
Headspace
Gympass+
One Medical
Employee Assistance Program

Job summary

SambaNova Systems is seeking an Architect on the Inference Systems Performance team to own end-to-end performance for large-scale LLM inference, from tokenization to deployment, across hardware and software components. You will drive reproducible workload capture, benchmarking, modeling, and simulation to inform capacity planning against customer SLOs.

Lead the performance-modeling effort, mentor senior engineers, and collaborate with model optimization, systems, and product teams to balance

Qualifications

  • 12+ years of experience in performance engineering, with a demonstrated record of technical leadership on large-scale, complex systems.
  • Deep expertise in end-to-end performance analysis of distributed systems with many moving parts and the ability to localize bottlenecks that others cannot.
  • Proven command of realistic workload generation and simulation and of performance modeling, including calibrating models against real, variable workloads.
  • Demonstrated ability to enter an unfamiliar domain and apply core performance methods with transferable discipline expertise.
  • Ability to lead cross-functional efforts, mentor senior engineers, and influence organizational direction.
  • Experience representing an organization's credibly to customers and partners
  • Track record of independently scoping and delivering high-complexity, high-ambiguity work with significant impact on products or roadmap

Responsibilities

  • Define and drive the technical strategy for inference-systems performance including workload capture, benchmarking, modeling, and simulation, while developing an architecture that enables many potential futures
  • Build the workload-capture and agentic-benchmarking capability - capture representative production traffic and enforce the discipline of interrogating results
  • Own the performance-modeling and simulation practice - models that predict how a configuration change moves the output, informing capacity planning against customer SLOs and next-generation system and hardware planning
  • Attack the end-to-end profiling gap - drive tooling that produces accurate, actionable profiles of a distributed inference pipeline so bottlenecks can be localized across host, accelerator, and fabric
  • Serve as the senior technical voice across model-optimization, systems, hardware, and product, tying together multiple engineering activities and teams, and weighing trade-offs of reliability, scalability, operational cost, and ease of adoption
  • Act as a resource for the entire organization including representing SambaNova's performance story to customers and partners
  • Mentor and multiply by raising the capability of principal and senior engineers, building the systems, tools, and patterns that make everyone more productive
  • Drive the resolution of the most ambiguous, novel challenges that span organizational boundaries or have no established answer in the field yet

Skills

Performance engineering
Distributed systems
Workload generation
Performance modeling
Leadership
Cross-functional collaboration
Customer storytelling

Tools

LLM inference serving frameworks

Job description

SambaNova Systems is seeking an Architect on the Inference Systems Performance team to own end-to-end performance for large-scale LLM inference, from tokenization to deployment, across hardware and software components. You will drive reproducible workload capture, benchmarking, modeling, and simulation to inform capacity planning against customer SLOs.

Lead the performance-modeling effort, mentor senior engineers, and collaborate with model optimization, systems, and product teams to balance

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Inference Systems Performance Architect
Senior Inference Systems Performance Architect

SambaNova • San Jose (CA)

On-site
USD 245,000 - 325,000
Health insurance
Gympass+
One Medical
Lead, End-to-End Inference Performance
Lead, End-to-End Inference Performance

SambaNovaSystems • San Jose (CA)

On-site
USD 245,000 - 325,000
AI Systems Performance Engineer - Scalable LLM Inference
AI Systems Performance Engineer - Scalable LLM Inference

SambaNova Systems • San Jose (CA)

On-site
USD 135,000 - 165,000
Equity
Health insurance
Well-being benefits
LLM Inference Systems Performance Engineer
LLM Inference Systems Performance Engineer

3M HEALTHCARE • Austin (TX)

On-site
USD 150,000 - 210,000
Medical, dental, and vision coverage
Income protection benefits
Paid family leave
+1
LLM Inference Systems Architect
LLM Inference Systems Architect

Netpreme • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Relocation assistance
Visa sponsorship
Lunch stipend
+1
Senior LLM Inference Architect — Performance & Benchmarking
Senior LLM Inference Architect — Performance & Benchmarking

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity and benefits
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Inference Systems Performance Architect
Inference Systems Performance Architect

SambaNovaSystems • San Jose (CA)

On-site
USD 245,000 - 325,000
LLM Systems Engineer: Inference Hardware & Research
LLM Systems Engineer: Inference Hardware & Research

Netpreme • Cambridge (MA)

On-site
USD 190,000 - 230,000
Relocation assistance
Visa sponsorship
Daily lunch stipend
+2
Senior LLM Inference Engineer — Performance & GPU Optimization
Senior LLM Inference Engineer — Performance & GPU Optimization

Confidential • United States

On-site
USD 180,000 - 240,000