Member of Technical Staff, ML Inference Engineering

Sanas

Palo Alto (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Sanas is building real-time speech and language models deployed on-premise inside sovereign data centers, delivering low-latency, high-throughput AI for multi-node GPU workloads.

As a Senior Engineer, you will shape core infrastructure and architecture decisions, lead performance optimizations, and own the inference engine to scale research and production workloads. This role is in the SF Bay Area with hands-on impact in a fast-growing startup.

Qualifications

  • 5+ years of high-performance software development experience.
  • Strong familiarity with NVIDIA GPU architecture and CUDA.
  • Experience with the LLM serving stack from kernels to autoscaling.
  • Background in LLM/Speech-to-Text/Text-to-Speech inference is preferred.
  • Record of shipping research or systems usable by others.

Responsibilities

  • Optimize system and GPU performance for high-throughput AI workloads.
  • Analyze latency, throughput, memory, and compute efficiency.
  • Profile and fix GPU- and kernel-level bottlenecks.
  • Own and evolve the inference engine for reliability at scale.

Skills

High-performance coding
CUDA familiarity
LLM serving stack
Speech/LLM inference
Research/Systems shipping

Tools

Kubernetes
InfiniBand
RoCE
Bare-metal provisioning
Distributed storage

Job description

Member of Technical Staff, ML Inference Engineering

Sanas is pioneering the future of human communication. Founded by a team of Stanford researchers and entrepreneurs with deep industry experience, Sanas has developed the world's first real-time speech AI platform capable of accent translation, noise cancellation, speech enhancement, cross-language communication, and more.

Sanas makes conversations clearer, more inclusive, and more effective, removing barriers that prevent people from being understood, regardless of accent, background noise, or native language.

Sanas is currently one of the fastest growing startups in Silicon Valley, growing from $16M to $50M ARR in 2025. The company's core business is profitable and is on track to end 2026 with >$120M ARR. Our team combines deep expertise in model innovation and systems engineering with a design-minded product engineering culture to build and ship cutting-edge AI models and experiences — entirely in-house.

Sanas is a 130 person team, established in 2020. In this short span, we've successfully secured over $100 million in funding. Our innovation has been supported by the industry's leading investors, including Insight Partners, Google Ventures, Quadrille Capital, General Catalyst, Quiet Capital, and other influential investors. Our reputation is further solidified by collaborations with numerous Fortune 100 companies. With Sanas, you're not just adopting a product; you're investing in the future of communication.

If you’re looking to have a significant role in roadmapping and driving technical directions, if you’re looking to deploy challenging and big ideas without much overhead or slowness, if you're looking to leave your mark on an ambitious, generational mission to change how the worlds thinks about speech + AI, then Sanas is a well-suited place for you.

Sanas is bringing real-time speech and language models on-premise — deployed at scale directly inside sovereign data centers, not served from behind a hosted cloud endpoint. It's one of the most demanding environments in the industry: strict latency budgets, massive concurrency, and infrastructure that needs to be private and reliable.

We're looking for a deeply hands-on, senior engineer to help lead that build. This is someone who shapes core infrastructure and architecture decisions rather than just executing against a specification, and who naturally raises the level of the engineers working alongside them.

Performance Optimization
  • Optimize system and GPU performance for high-throughput AI workloads across multi-node training and inference
  • Analyze and improve latency, throughput, memory usage, and compute efficiency
  • Profile system performance to detect and resolve GPU- and kernel-level bottlenecks
  • Implement low-level optimizations using CUDA, Triton, and other performance tooling
  • Improve support for mixed precision, quantization, and model graph optimization
  • Build and maintain performance benchmarking and monitoring infrastructure
  • Scale inference and training systems across multi-GPU, multi-node environments
Inference Systems & Reliability
  • Own and evolve our inference engine, enabling reliability and performance at scale
  • Develop and optimize runtime inference services for large-scale AI applications
  • Implement robust, fault-tolerant systems for data ingestion and processing
Requirements
Must-have:
  • 5+ years of experience writing high-quality, high-performance code
  • Familiarity with NVIDIA GPU architecture and CUDA
  • Fluency in the LLM serving stack, from kernels and quantization up to schedulers and autoscaling
  • A research-leaning or systems background in LLM, Speech-to-Text, Text-to-Speech, or Speech-to-Speech inference, with work you can point to
  • A record of shipping research or systems that other people build on, whether in a lab or in industry
Nice-to-have:
  • Experience serving low-precision (FP4/FP8) models, multiple LoRA adapters within one model instance (Multi-LoRA), or models distributed across several GPU nodes
  • Experience maintaining or contributing to open-source ML projects
  • Experience managing machine learning workloads on Kubernetes clusters
  • Experience with InfiniBand or RoCE networking
  • Experience with bare-metal provisioning and lifecycle management
  • Experience operating large-scale AI training or inference clusters
  • Experience with hardware health monitoring and predictive failure detection
  • Experience with distributed storage systems
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, LLM Post-Training, Applied
Member of Technical Staff, LLM Post-Training, Applied

Sanas • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Staff+ Production Engineer
Staff+ Production Engineer

Sanas.AI Inc. • Palo Alto (CA), Northern (KY)

Hybrid
USD 140,000 - 170,000
Staff+ Production Engineer
Staff+ Production Engineer

Tensec • Palo Alto (CA)

On-site
USD 130,000 - 160,000
Senior ML Inference Engineer — Real-Time Systems
Senior ML Inference Engineer — Real-Time Systems

Sanas • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Research Evaluations
Member of Technical Staff, Research Evaluations

Sanas • Palo Alto (CA)

On-site
USD 140,000 - 190,000
Member of Technical Staff, MLSys
Member of Technical Staff, MLSys

Bake AI • San Mateo (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Staff Applied AI Inference Engineer
Staff Applied AI Inference Engineer

Crusoe Energy Systems • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health benefits
401(k) match
Paid time off
+1
Senior Director of Accounting
Senior Director of Accounting

Quiet Capital • Palo Alto (CA)

On-site
USD 120,000 - 180,000
Senior Technical Program Manager (Engineering) - AI Tooling & Systems
Senior Technical Program Manager (Engineering) - AI Tooling & Systems

Madrona Venture Labs • United States

On-site
USD 150,000 - 230,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits