Senior ML Inference Engineer — Real-Time Systems

Sanas

Palo Alto (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Sanas is building real-time speech and language models deployed on-premise inside sovereign data centers, delivering low-latency, high-throughput AI for multi-node GPU workloads.

As a Senior Engineer, you will shape core infrastructure and architecture decisions, lead performance optimizations, and own the inference engine to scale research and production workloads. This role is in the SF Bay Area with hands-on impact in a fast-growing startup.

Qualifications

  • 5+ years of high-performance software development experience.
  • Strong familiarity with NVIDIA GPU architecture and CUDA.
  • Experience with the LLM serving stack from kernels to autoscaling.
  • Background in LLM/Speech-to-Text/Text-to-Speech inference is preferred.
  • Record of shipping research or systems usable by others.

Responsibilities

  • Optimize system and GPU performance for high-throughput AI workloads.
  • Analyze latency, throughput, memory, and compute efficiency.
  • Profile and fix GPU- and kernel-level bottlenecks.
  • Own and evolve the inference engine for reliability at scale.

Skills

High-performance coding
CUDA familiarity
LLM serving stack
Speech/LLM inference
Research/Systems shipping

Tools

Kubernetes
InfiniBand
RoCE
Bare-metal provisioning
Distributed storage

Job description

Sanas is building real-time speech and language models deployed on-premise inside sovereign data centers, delivering low-latency, high-throughput AI for multi-node GPU workloads.

As a Senior Engineer, you will shape core infrastructure and architecture decisions, lead performance optimizations, and own the inference engine to scale research and production workloads. This role is in the SF Bay Area with hands-on impact in a fast-growing startup.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, ML Inference Engineering
Member of Technical Staff, ML Inference Engineering

Sanas • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Senior ML Engineer, Voice AI — Real-Time Inference Lead
Senior ML Engineer, Voice AI — Real-Time Inference Lead

Together AI • San Francisco (CA)

On-site
USD 200,000 - 260,000
Senior ML Engineer, Voice AI — Real-Time Inference at Scale
Senior ML Engineer, Voice AI — Real-Time Inference at Scale

Together AI • San Francisco (CA)

On-site
USD 200,000 - 260,000
Competitive salary
Startup equity
Health insurance
+1
Staff ML Engineer, Voice AI - Real-Time Inference Lead
Staff ML Engineer, Voice AI - Real-Time Inference Lead

Together AI • San Francisco (CA)

On-site
USD 220,000 - 280,000
Health insurance
Startup equity
Competitive benefits
Research Scientist (Model Evaluation)
Research Scientist (Model Evaluation)

Jobless • Palo Alto (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Senior ML Inference Engineer — Production Systems
Senior ML Inference Engineer — Production Systems

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Real-Time Multimodal Inference Architect
Senior Real-Time Multimodal Inference Architect

Amazon • Sunnyvale (CA)

On-site
USD 192,000 - 260,000
Health insurance
401(k) matching
RSU/Stock options
Member of Technical Staff, LLM Post-Training, Applied
Member of Technical Staff, LLM Post-Training, Applied

Sanas • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Senior ML Infra Engineer — Real-Time, Low-Latency
Senior ML Infra Engineer — Real-Time, Low-Latency

LMArena • San Francisco (CA)

On-site
USD 170,000 - 260,000
Competitive compensation
Comprehensive health benefits
Opportunity to work on cutting-edge AI
Senior Speech ML Engineer — Scalable ASR & Real‑Time Insights
Senior Speech ML Engineer — Scalable ASR & Real‑Time Insights

Level AI • Mountain View (CA)

On-site
USD 120,000 - 170,000