AI Systems Performance Engineer - Scalable LLM Inference

SambaNova Systems

San Jose (CA)

On-site

USD 135,000 - 165,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Health insurance
Well-being benefits

Job summary

SambaNova Systems is seeking an AI Systems Performance Engineer to bring up and optimize foundation models on its reconfigurable dataflow platform. You’ll work with cutting-edge models like LLMs and multimodal architectures, profiling across model, compiler, runtime, and hardware layers to improve throughput, latency, and memory efficiency.

You will collaborate with ML, compiler, runtime, and hardware teams, exploring new techniques in architecture, quantization, scheduling, and memory

Qualifications

  • Bachelor's or Master's degree in computer science, electrical engineering, computer engineering, or a related technical field (completed or expected before start date).
  • Strong programming skills in Python, C++, or a similar language.
  • Foundations in algorithms, data structures, computer architecture, operating systems, or parallel computing.
  • Familiarity with deep learning and at least one major ML framework (PyTorch, TensorFlow, JAX).
  • Strong analytical and problem-solving skills with interest in system performance.
  • Ability and enthusiasm to learn across ML, software systems, and hardware.

Responsibilities

  • Bring up foundation models on SambaNova platform through software stack.
  • Analyze and profile model execution to identify bottlenecks across model, compiler, runtime, and hardware layers.
  • Optimize AI workloads for throughput, latency, memory efficiency, and scalability.
  • Collaborate with ML, compiler, runtime, and hardware engineers on high-performance AI applications.
  • Explore and integrate techniques in model architecture, quantization, scheduling, caching, and memory optimization.
  • Develop tools, benchmarks, and performance analysis methodologies for large-scale AI inference.
  • Investigate new model architectures and translate research into efficient production implementations.
  • Contribute ideas for dataflow, scheduling, and system optimizations for single-node and distributed inference.

Skills

Python
C++
Algorithms & data structures
DL frameworks

Education

Bachelor's or Master's degree in CS/EE/CE or related field

Tools

CUDA
Triton
OpenCL
vLLM
DeepSpeed
Megatron
TensorRT

Job description

SambaNova Systems is seeking an AI Systems Performance Engineer to bring up and optimize foundation models on its reconfigurable dataflow platform. You’ll work with cutting-edge models like LLMs and multimodal architectures, profiling across model, compiler, runtime, and hardware layers to improve throughput, latency, and memory efficiency.

You will collaborate with ML, compiler, runtime, and hardware teams, exploring new techniques in architecture, quantization, scheduling, and memory

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Systems Performance Engineer
Senior AI Systems Performance Engineer

Doist • San Jose (CA)

On-site
USD 140,000 - 190,000
Health insurance
HSA contributions
Dental
+5
LLM Inference Systems Performance Architect
LLM Inference Systems Performance Architect

Doist • San Jose (CA)

On-site
USD 245,000 - 325,000
Health Insurance
Dental Insurance
Vision Insurance
+8
Senior AI Systems Performance Engineer San Jose, California, United States
Senior AI Systems Performance Engineer San Jose, California, United States

SambaNova • Palo Alto (CA)

On-site
USD 120,000 - 150,000
95% premium coverage for employee medical insurance
Health Savings Account with employer contribution
Flexible Spending Account options
Senior Inference Systems Performance Architect
Senior Inference Systems Performance Architect

SambaNova • San Jose (CA)

On-site
USD 245,000 - 325,000
Health insurance
Gympass+
One Medical
Senior ML Infra Engineer: High-Throughput AI Inference
Senior ML Infra Engineer: High-Throughput AI Inference

SambaNovaSystems • United States

On-site
USD 200,000 - 275,000
Health insurance
Health Savings Account (HSA)
Headspace subscription
+2
Principal AI Systems Performance Engineer
Principal AI Systems Performance Engineer

SambaNova • San Jose (CA)

On-site
USD 180,000 - 240,000
Senior Principal ML Engineer: LLMs & Hardware Co-Design
Senior Principal ML Engineer: LLMs & Hardware Co-Design

Doist • San Jose (CA)

On-site
USD 220,000 - 300,000
Health insurance
HSA
Dental
+9
Lead, End-to-End Inference Performance
Lead, End-to-End Inference Performance

SambaNovaSystems • San Jose (CA)

On-site
USD 245,000 - 325,000
AI Systems Performance Engineer - New Graduate
AI Systems Performance Engineer - New Graduate

SambaNova Systems • San Jose (CA)

On-site
USD 135,000 - 165,000
Equity
Health insurance
Well-being benefits
Senior Principal ML Engineer — LLMs & Hardware Co-Design
Senior Principal ML Engineer — LLMs & Hardware Co-Design

SambaNovaSystems • San Jose (CA)

On-site
USD 220,000 - 300,000
Equity
Health insurance
HSA
+1