Senior AI Inference Performance Engineer

SambaNova

San Jose (CA)

On-site

USD 180,000 - 240,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

SambaNova is seeking an ML performance engineer to optimize and scale state-of-the-art foundation models on its reconfigurable dataflow platform. You will work hands-on with models like DeepSeek, GPT OSS, and others to push throughput, latency, and efficiency while coordinating with compiler, runtime, and hardware teams.

Ideal candidates have strong ML performance tuning background, proficiency in Python or C++, and experience with major ML frameworks.

Qualifications

  • Bachelor's or higher degree in computer science, electrical engineering, or a related field.
  • 3+ years of experience in deep learning model development and performance optimization.
  • Experience with compiler, runtime, or kernel-level optimization.
  • Software–hardware co-design or systems performance tuning.
  • Proficiency in Python or C++, with strong foundations in algorithms and numerical computing.
  • Experience with PyTorch, TensorFlow, or JAX.
  • Ability to analyze and optimize performance in real-world ML pipelines.

Responsibilities

  • Bring up and optimize foundation models on the SambaNova platform through the software stack.
  • Profile and enhance model performance across compiler, runtime, and hardware layers for high throughput and low latency.
  • Collaborate with ML, compiler, runtime, and hardware teams to deliver high-performance AI applications.
  • Integrate advances in model architecture, quantization, scheduling, and memory optimization.
  • Develop robust, scalable end-to-end inference solutions aligned with customer needs.
  • Identify bottlenecks and propose dataflow or scheduling optimizations for single-node and distributed systems.

Skills

Deep learning optimization
Python
C++
Algorithms/data structures
Perf analysis

Education

Bachelor's or higher in CS/EE

Tools

PyTorch
TensorFlow
JAX
CUDA
Triton
DeepSpeed
Megatron
vLLM
TensorRT

Job description

SambaNova is seeking an ML performance engineer to optimize and scale state-of-the-art foundation models on its reconfigurable dataflow platform. You will work hands-on with models like DeepSeek, GPT OSS, and others to push throughput, latency, and efficiency while coordinating with compiler, runtime, and hardware teams.

Ideal candidates have strong ML performance tuning background, proficiency in Python or C++, and experience with major ML frameworks.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Systems Performance Engineer
Senior AI Systems Performance Engineer

Doist • San Jose (CA)

On-site
USD 140,000 - 190,000
Health insurance
HSA contributions
Dental
+5
AI Systems Performance Engineer - Scalable LLM Inference
AI Systems Performance Engineer - Scalable LLM Inference

SambaNova Systems • San Jose (CA)

On-site
USD 135,000 - 165,000
Equity
Health insurance
Well-being benefits
Senior AI Systems Performance Engineer San Jose, California, United States
Senior AI Systems Performance Engineer San Jose, California, United States

SambaNova • Palo Alto (CA)

On-site
USD 120,000 - 150,000
95% premium coverage for employee medical insurance
Health Savings Account with employer contribution
Flexible Spending Account options
Senior ML Infra Engineer: High-Throughput AI Inference
Senior ML Infra Engineer: High-Throughput AI Inference

SambaNovaSystems • United States

On-site
USD 200,000 - 275,000
Health insurance
Health Savings Account (HSA)
Headspace subscription
+2
Principal AI Systems Performance Engineer
Principal AI Systems Performance Engineer

SambaNova • San Jose (CA)

On-site
USD 180,000 - 240,000
Principal AI Systems Performance Engineer
Principal AI Systems Performance Engineer

Socket.dev • San Jose (CA)

On-site
USD 140,000 - 190,000
Health insurance
HSA contributions
Dental
+5
Senior AI Inference Performance Engineer — Scale GPUs
Senior AI Inference Performance Engineer — Scale GPUs

NVIDIA • California (MO)

On-site
USD 124,000 - 196,000
Equity eligibility
Senior AI Inference Cloud Platform Engineer
Senior AI Inference Cloud Platform Engineer

SambaNova • Austin (TX)

On-site
USD 144,000 - 189,000
Health insurance
HSA with employer contribution
Dental, Vision
+2
Senior ML Performance Engineer: Scale & Throughput
Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
Senior AI Systems Performance Engineer: Drive SOTA Inference
Senior AI Systems Performance Engineer: Drive SOTA Inference

SambaNova • Palo Alto (CA)

On-site
USD 120,000 - 150,000
95% premium coverage for employee medical insurance
Health Savings Account with employer contribution
Flexible Spending Account options