ML Model Serving Engineer

Sesame

Bellevue (WA)

On-site

USD 175,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

401(k) max employer match: 3.5% of compensation
100% employer‑paid health, vision, and dental benefits
Unlimited PTO and sick time
Flexible spending account with employer matching
Guardian Employee Assistance Program (EAP)
Competitive stock options

Job summary

Sesame, located in Bellevue, Washington, is seeking a talented engineer to join our team focused on revolutionizing the way computers interact with humans. The role involves optimizing machine learning models for a new consumer product category, working with state-of-the-art LLM and vision models. Candidates should have significant systems programming and performance engineering experience.

You will enjoy benefits like unlimited PTO, 100% employer-paid health coverage, and competitive stock options. The compensation range for this role is $175K - $280K.

Qualifications

  • Expert in some differentiable array computing framework, preferably PyTorch.
  • Experience in optimizing machine learning models for high throughput with low latency.
  • Extensive experience with complex server systems and performance engineering.

Responsibilities

  • Turbocharge our serving layer with various LLM, speech, and vision models.
  • Partner with ML engineers to build an efficient and reliable serving layer.
  • Modify LLM serving frameworks for high-performance model serving.

Skills

Expert in some differentiable array computing framework
Optimizing machine learning models
Significant systems programming experience
Bottleneck analysis
Staying up to date on model serving optimization

Tools

PyTorch
GCP
AWS
Azure
Kubernetes
Ray

Job description

About Sesame

Sesame believes in a future where computers are lifelike – with the ability to see, hear, and collaborate with us in ways that feel natural and human. With this vision, we're designing a new kind of computer, focused on making voice agents part of our daily lives. Our team brings together founders from Oculus and Ubiquity6, alongside proven leaders from Meta, Google, and Apple, with deep expertise spanning hardware and software. Join us in shaping a future where computers truly come alive.

Responsibilities
  • Turbocharge our serving layer, consisting of a variety of LLM, speech, and vision models.
  • Partner with ML infrastructure and training engineers to build a fast, cost‑effective, accurate, and reliable serving layer to power a new consumer product category.
  • Modify and extend LLM serving frameworks like VLLM and SGLang to take advantage of the latest techniques in high‑performance model serving.
  • Work with the training team to identify opportunities to produce faster models without sacrificing quality.
  • Use techniques like in‑flight batching, caching, and custom kernels to speed up inference.
  • Find ways to reduce model initialization times without sacrificing quality.
Required Qualifications
  • Expert in some differentiable array computing framework, preferably PyTorch.
  • Expert in optimizing machine learning models for serving reliably at high throughput, with low latency.
  • Significant systems programming experience; ex. Experience working on high‑performance server systems—you’d be just as comfortable with the internals of VLLM as you would with a complex PyTorch codebase.
  • Significant performance engineering experience; ex. Bottleneck analysis in high‑scale server systems or profiling low‑level systems code.
  • Always up to date on the latest techniques for model serving optimization.
Preferred Qualifications
  • Familiarity with high-performance LLM serving; ex. experience with VLLM, SGlang deployment, and internals.
  • Experience with a public cloud platform such as GCP, AWS, or Azure.
  • Experience deploying and scaling inference workloads in the cloud using Kubernetes, Ray, etc.
  • You like to ship and have a track record of leading complex multi-month projects without assistance.
  • You’re excited to learn new things and work in a multitude of roles.
Benefits
  • 401(k) max employer match: 3.5% of compensation
  • 100% employer‑paid health, vision, and dental benefits for you and your dependents
  • Unlimited PTO and sick time
  • Flexible spending account with employer matching up to $1,650/year (medical FSA)
  • Guardian Employee Assistance Program (EAP)
  • Opportunity to share in the company's success with competitive stock options

Benefits do not apply to contingent/contract workers.

Compensation Range: $175K - $280K

Sesame is committed to a workplace where everyone feels valued, respected, and empowered. We welcome all qualified applicants, embracing diversity in race, gender, identity, orientation, ability, and more. We provide reasonable accommodations for applicants with disabilities.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Model Serving Engineer - High-Performance Inference
ML Model Serving Engineer - High-Performance Inference

Sesame • New York (NY)

On-site
USD 175,000 - 280,000
401(k) max employer match: 3.5%
100% employer‑paid health, vision, and dental benefits
Unlimited PTO and sick time
+3
Staff Software Engineer, Backend Infrastructure
Staff Software Engineer, Backend Infrastructure

Sesame • San Francisco (CA)

On-site
USD 110,000 - 170,000
401 (k) employer match
100% employer-paid health benefits
Unlimited PTO and sick time
SWE - Backend Infrastructure Engineer
SWE - Backend Infrastructure Engineer

Sesame • New York (NY)

On-site
USD 175,000 - 280,000
401(k) max employer match: 3.5%
100% employer-paid health, vision, and dental benefits
Unlimited PTO and sick time
+1
Research Engineer
Research Engineer

Sesame • Bellevue (WA)

On-site
USD 190,000 - 320,000
401(k) max match
Health, vision & dental benefits
Unlimited PTO"
+3
ML Model Serving Engineer — High-Throughput Inference
ML Model Serving Engineer — High-Throughput Inference

Sesame • Bellevue (WA)

On-site
USD 175,000 - 280,000
401(k) max employer match: 3.5% of compensation
100% employer‑paid health, vision, and dental benefits
Unlimited PTO and sick time
+3
Research Engineer
Research Engineer

Sesame • San Francisco (CA)

On-site
USD 120,000 - 160,000
401(k) max employer match
100% employer-paid health benefits
Unlimited PTO and sick time
+3
Software Engineer - Backend
Software Engineer - Backend

Sesame • San Francisco (CA)

On-site
USD 175,000 - 280,000
401(k) matching
Health, vision, dental benefits
Unlimited PTO
+3
SWE - Developer Infrastructure Engineer
SWE - Developer Infrastructure Engineer

Sesame • New York (NY)

On-site
USD 175,000 - 280,000
401k matching
100% employer-paid health, vision, and dental benefits
Unlimited PTO and sick time
+1
Data Engineer, Machine Learning
Data Engineer, Machine Learning

Sesame • San Francisco (CA)

On-site
USD 120,000 - 160,000
401(k) max employer match: 3.5% of compensation
100% employer-paid health, vision, and dental benefits
Unlimited PTO and sick time
+3
Research Scientist
Research Scientist

Sesame • Bellevue (WA)

On-site
USD 190,000 - 320,000
401k match
Health, vision, and dental insurance
Unlimited PTO
+3