ML Model Serving Engineer — High-Throughput Inference

Sesame

Bellevue (WA)

On-site

USD 175,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

401(k) max employer match: 3.5% of compensation
100% employer‑paid health, vision, and dental benefits
Unlimited PTO and sick time
Flexible spending account with employer matching
Guardian Employee Assistance Program (EAP)
Competitive stock options

Job summary

Sesame, located in Bellevue, Washington, is seeking a talented engineer to join our team focused on revolutionizing the way computers interact with humans. The role involves optimizing machine learning models for a new consumer product category, working with state-of-the-art LLM and vision models. Candidates should have significant systems programming and performance engineering experience.

You will enjoy benefits like unlimited PTO, 100% employer-paid health coverage, and competitive stock options. The compensation range for this role is $175K - $280K.

Qualifications

  • Expert in some differentiable array computing framework, preferably PyTorch.
  • Experience in optimizing machine learning models for high throughput with low latency.
  • Extensive experience with complex server systems and performance engineering.

Responsibilities

  • Turbocharge our serving layer with various LLM, speech, and vision models.
  • Partner with ML engineers to build an efficient and reliable serving layer.
  • Modify LLM serving frameworks for high-performance model serving.

Skills

Expert in some differentiable array computing framework
Optimizing machine learning models
Significant systems programming experience
Bottleneck analysis
Staying up to date on model serving optimization

Tools

PyTorch
GCP
AWS
Azure
Kubernetes
Ray

Job description

Sesame, located in Bellevue, Washington, is seeking a talented engineer to join our team focused on revolutionizing the way computers interact with humans. The role involves optimizing machine learning models for a new consumer product category, working with state-of-the-art LLM and vision models. Candidates should have significant systems programming and performance engineering experience.

You will enjoy benefits like unlimited PTO, 100% employer-paid health coverage, and competitive stock options. The compensation range for this role is $175K - $280K.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Model Serving Engineer - High-Performance Inference
ML Model Serving Engineer - High-Performance Inference

Sesame • New York (NY)

On-site
USD 175,000 - 280,000
401(k) max employer match: 3.5%
100% employer‑paid health, vision, and dental benefits
Unlimited PTO and sick time
+3
ML Model Serving Engineer
ML Model Serving Engineer

Sesame • Bellevue (WA)

On-site
USD 175,000 - 280,000
401(k) max employer match: 3.5% of compensation
100% employer‑paid health, vision, and dental benefits
Unlimited PTO and sick time
+3
AI Infrastructure Engineer, Model Serving Platform
AI Infrastructure Engineer, Model Serving Platform

Scale AI, Inc. • New York (NY)

On-site
USD 180,000 - 225,000
Comprehensive health coverage
Equity compensation
Learning and development stipend
+2
Machine Learning Engineering Manager - LLM Serving (Remote - US)
Machine Learning Engineering Manager - LLM Serving (Remote - US)

Jobgether • United States

Remote
USD 176,000 - 252,000
Competitive salary range
Comprehensive health insurance
Paid parental leave
+4
ML Model Serving Infra Engineer - On-Prem & Edge
ML Model Serving Infra Engineer - On-Prem & Edge

Palantir Technologies • New York (NY)

Hybrid
USD 145,000 - 200,000
Medical, dental, and vision insurance
401k plan
Paid time off
+1
Staff Engineer, Model Efficiency & LLM Inference
Staff Engineer, Model Efficiency & LLM Inference

Visa Hunt • New York (NY)

Hybrid
USD 150,000 - 210,000
Lunch stipend
Health and dental benefits
RRSP matching
+4
Staff Engineer - LLM Inference & Serving at Scale
Staff Engineer - LLM Inference & Serving at Scale

Prime Intellect • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Research Engineer - LLM/VLM Inference Optimization (Seed Infra) Seattle Regular
Research Engineer - LLM/VLM Inference Optimization (Seed Infra) Seattle Regular

ByteDance • Seattle (WA)

On-site
USD 232,000 - 428,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+1
Remote Model Serving Engineer for High-Scale ML
Remote Model Serving Engineer for High-Scale ML

Bright Vision Technologies • Canton Charter Township (MI)

On-site
USD 74,000 - 98,000
ML Engineer: Distributed Inference & Model Serving
ML Engineer: Distributed Inference & Model Serving

ByteDance • Seattle (WA)

On-site
USD 177,000 - 417,000