Principal AI Systems Performance Engineer

Socket.dev

San Jose (CA)

On-site

USD 140,000 - 190,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
HSA contributions
Dental
Vision
Disability insurance
Life insurance
FSA
Headspace

Job summary

SambaNova Systems in San Jose is seeking a talented ML performance engineer to optimize and scale state-of-the-art foundation models on our reconfigurable dataflow platform. You will work hands-on with leading models to push throughput, latency, and efficiency, bridging gaps between compiler, runtime, and hardware teams to deliver world-record AI inference performance.

The role requires a strong foundation in deep learning and systems optimization, 3+ years of experience, and proficiency in

Qualifications

  • Bachelor's or higher in computer science, electrical engineering, or a related field (e.g., applied mathematics, physics, or statistics).
  • 3+ years of experience in Deep learning model development and performance optimization, compiler/runtime/kernel-level optimization, or software–hardware co-design.
  • Proficiency in Python or C++, with strong foundations in algorithms, data structures, and numerical computing.
  • Experience with at least one major ML framework — PyTorch, TensorFlow, or JAX.
  • Demonstrated ability to analyze and optimize performance in real-world ML pipelines.

Responsibilities

  • Bring up and optimize cutting-edge foundation models (e.g., DeepSeek, Llama, Qwen, and others) on the SambaNova platform through the SambaNova software stack.
  • Profile and enhance model performance across compiler, runtime, and hardware layers to achieve SOTA throughput and latency.
  • Collaborate with machine learning, compiler, runtime, and hardware teams to deliver co-designed, high-performance AI applications.
  • Integrate the latest advances in model architecture, quantization, scheduling, and memory optimization from both academia and industry.
  • Develop robust, scalable, and efficient end-to-end inference solutions aligned with customer needs.
  • Identify performance bottlenecks and propose dataflow or scheduling optimizations for both single-node and distributed systems.

Skills

Python
C++
PyTorch
TensorFlow
JAX

Education

Bachelor's degree in CS/EE or related

Tools

CUDA
Triton
DeepSpeed
Megatron
vLLM
TensorRT
CuDNN

Job description

The era of pervasive AI has arrived. In this era, organizations will use generative AI to unlock hidden value in their data, accelerate processes, reduce costs, drive efficiency and innovation to fundamentally transform their businesses and operations at scale. SambaNova Suite™ is the first full-stack, generative AI platform, from chip to model, optimized for enterprise and government organizations. Powered by the intelligent SN40L chip, the SambaNova Suite is a fully integrated platform, delivered on-premises or in the cloud, combined with state-of-the-art open-source models that can be easily and securely fine-tuned using customer data for greater accuracy. Once adapted with customer data, customers retain model ownership in perpetuity, so they can turn generative AI into one of their most valuable assets.

About the role

We are seeking a talented and driven ML performance engineer to optimize and scale state-of-the-art foundation models on SambaNova's reconfigurable dataflow platform. You'll work hands-on with some of the most advanced models in the world — such as DeepSeek R1, GPT OSS, and other frontier architectures — to push the limits of throughput, latency, and efficiency. In this role, you'll bridge the gap between deep learning and systems performance, collaborating across compiler, runtime, and hardware layers to deliver world-record performance for large-scale AI inference.

Responsibilities
  • Bring up and optimize cutting-edge foundation models (e.g., DeepSeek, Llama, Qwen, and others) on the SambaNova platform through the SambaNova software stack.
  • Profile and enhance model performance across compiler, runtime, and hardware layers to achieve SOTA throughput and latency.
  • Collaborate with machine learning, compiler, runtime, and hardware teams to deliver co-designed, high-performance AI applications.
  • Integrate the latest advances in model architecture, quantization, scheduling, and memory optimization from both academia and industry.
  • Develop robust, scalable, and efficient end-to-end inference solutions aligned with customer needs.
  • Identify performance bottlenecks and propose dataflow or scheduling optimizations for both single-node and distributed systems.
Basic Qualifications
  • Bachelor's or higher degree in computer science, electrical engineering, or a related field (e.g., applied mathematics, physics, or statistics).
  • 3+ years of experience in one or more of the following areas: Deep learning model development and performance optimization, Compiler, runtime, or kernel-level optimization, Software–hardware co-design or systems performance tuning.
  • Proficiency in Python or C++, with strong foundations in algorithms, data structures, and numerical computing.
  • Experience with at least one major ML framework — PyTorch, TensorFlow, or JAX.
  • Demonstrated ability to analyze and optimize performance in real-world ML pipelines.
Preferred Qualifications
  • Hands-on experience with LLM or multimodal model training and inference.
  • Background in large-scale distributed training, continuous batching, and high-throughput inference systems.
  • Familiarity with quantization, graph optimization, kernel fusion, and model partitioning.
  • Experience with frameworks such as DeepSpeed, Megatron, vLLM, or TensorRT.
  • Strong GPU programming skills (CUDA, Triton, or OpenCL); experience with cuDNN, cuBLAS, or similar libraries is a plus.
  • Knowledge of memory hierarchy optimization, caching, and scheduling for large-scale model execution.
  • Publication record or open-source contributions in ML systems or performance optimization is a plus.
EEO Policy

SambaNova Systems is an Equal Opportunity/Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard basis of age (40 and over), color, disability, gender identity, genetic information, marital status, military or veteran status, national origin/ancestry, race, religion, creed, sex (including pregnancy, childbirth, breastfeeding), sexual orientation, and any other applicable status protected by federal, state, or local laws.

Benefits Summary for US-Based, Full-Time Employment Positions

SambaNova offers a competitive total rewards package, including the base salary, plus equity and benefits.

  • 95% premium coverage for employee medical insurance
  • 77% premium coverage for dependents and offer a Health Savings Account (HSA) with employer contribution
  • Dental
  • Vision
  • Short/Long term Disability
  • Basic Life
  • Voluntary Life
  • AD&D insurance plans
  • Flexible Spending Account (FSA) options like Health Care, Limited Purpose, and Dependent Care
  • full subscription to Headspace, Gympass+ membership with access to physical gyms, One Medical membership, counseling services with an Employee Assistance Program, and much more.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal AI Systems Performance Engineer
Principal AI Systems Performance Engineer

SambaNova • San Jose (CA)

On-site
USD 180,000 - 240,000
Senior AI Systems Performance Engineer San Jose, California, United States
Senior AI Systems Performance Engineer San Jose, California, United States

SambaNova • Palo Alto (CA)

On-site
USD 120,000 - 150,000
95% premium coverage for employee medical insurance
Health Savings Account with employer contribution
Flexible Spending Account options
AI Systems Performance Engineer - New Graduate
AI Systems Performance Engineer - New Graduate

SambaNova Systems • San Jose (CA)

On-site
USD 135,000 - 165,000
Equity
Health insurance
Well-being benefits
Inference Systems Performance Architect
Inference Systems Performance Architect

Socket.dev • San Jose (CA)

On-site
USD 245,000 - 325,000
Health Insurance
Dental Insurance
Vision Insurance
+8
Forward Deployment Engineer SambaNova Remote - US
Forward Deployment Engineer SambaNova Remote - US

Neura Market • Northern (KY)

Hybrid
USD 138,000 - 170,000
Equity
Health insurance
Well-being benefits
Forward Deployment Engineer
Forward Deployment Engineer

External SambaNova Systems • United States

On-site
USD 138,000 - 170,000
Equity
Health insurance
Health Savings Account (HSA)
Inference Systems Performance Architect
Inference Systems Performance Architect

SambaNovaSystems • San Jose (CA)

On-site
USD 245,000 - 325,000
Senior Software Engineer - ML Infrastructure
Senior Software Engineer - ML Infrastructure

SambaNovaSystems • United States

On-site
USD 200,000 - 275,000
Health insurance
Health Savings Account (HSA)
Headspace subscription
+2
Inference Systems Performance Architect
Inference Systems Performance Architect

SambaNova • San Jose (CA)

On-site
USD 245,000 - 325,000
Health insurance
Gympass+
One Medical
Software Architect
Software Architect

SambaNovaSystems • San Jose (CA)

On-site
USD 245,000 - 325,000