ML Model Serving Engineer - High-Performance Inference

Sesame

New York (NY)

On-site

USD 175,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

401(k) max employer match: 3.5%
100% employer‑paid health, vision, and dental benefits
Unlimited PTO and sick time
Flexible spending account with employer matching
Guardian Employee Assistance Program
Competitive stock options

Job summary

Sesame in New York is seeking an expert in optimizing machine learning models to turbocharge their serving layer, integrating LLM, speech, and vision models. The ideal candidate has significant experience in systems programming and performance engineering, aiming to improve high-throughput, low-latency serving.

Join a team dedicated to pioneering advancements in voice agents and enjoy competitive benefits, including health coverage and stock options, with a salary range of $175K to $280K.

Qualifications

  • Expert in some differentiable array computing framework, preferably PyTorch.
  • Expert in optimizing machine learning models for serving reliably at high throughput, with low latency.
  • Significant systems programming experience on high‑performance server systems.
  • Significant performance engineering experience with bottleneck analysis.
  • Up to date on the latest techniques for model serving optimization.

Responsibilities

  • Turbocharge the serving layer with LLM, speech, and vision models.
  • Collaborate with ML engineers to build a fast and reliable serving layer.
  • Modify LLM serving frameworks like VLLM for performance.
  • Identify opportunities with the training team to produce faster models.
  • Reduce model initialization times without sacrificing quality.

Skills

Expert in PyTorch
Optimizing machine learning models
Systems programming experience
Performance engineering experience
Model serving optimization techniques

Tools

Kubernetes
Ray
GCP
AWS
Azure

Job description

Sesame in New York is seeking an expert in optimizing machine learning models to turbocharge their serving layer, integrating LLM, speech, and vision models. The ideal candidate has significant experience in systems programming and performance engineering, aiming to improve high-throughput, low-latency serving.

Join a team dedicated to pioneering advancements in voice agents and enjoy competitive benefits, including health coverage and stock options, with a salary range of $175K to $280K.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Model Serving Engineer — High-Throughput Inference
ML Model Serving Engineer — High-Throughput Inference

Sesame • Bellevue (WA)

On-site
USD 175,000 - 280,000
401(k) max employer match: 3.5% of compensation
100% employer‑paid health, vision, and dental benefits
Unlimited PTO and sick time
+3
ML Model Serving Engineer
ML Model Serving Engineer

Sesame • Bellevue (WA)

On-site
USD 175,000 - 280,000
401(k) max employer match: 3.5% of compensation
100% employer‑paid health, vision, and dental benefits
Unlimited PTO and sick time
+3
AI Infrastructure Engineer, Model Serving Platform
AI Infrastructure Engineer, Model Serving Platform

Scale AI, Inc. • New York (NY)

On-site
USD 180,000 - 225,000
Comprehensive health coverage
Equity compensation
Learning and development stipend
+2
Machine Learning Engineering Manager - LLM Serving (Remote - US)
Machine Learning Engineering Manager - LLM Serving (Remote - US)

Jobgether • United States

Remote
USD 176,000 - 252,000
Competitive salary range
Comprehensive health insurance
Paid parental leave
+4
Inference Performance Engineer: Optimize Model Serving
Inference Performance Engineer: Optimize Model Serving

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
+1
ML Model Serving Infra Engineer - On-Prem & Edge
ML Model Serving Infra Engineer - On-Prem & Edge

Palantir Technologies • New York (NY)

Hybrid
USD 145,000 - 200,000
Medical, dental, and vision insurance
401k plan
Paid time off
+1
Member of Technical Staff, Inference & Serving
Member of Technical Staff, Inference & Serving

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Remote Model Serving Engineer for High-Scale ML
Remote Model Serving Engineer for High-Scale ML

Bright Vision Technologies • Canton Charter Township (MI)

On-site
USD 74,000 - 98,000
Senior ML Serving Engineer for LLMs & Inference
Senior ML Serving Engineer for LLMs & Inference

Alldus • San Jose (CA)

On-site
USD 180,000 - 220,000
Staff ML Engineer, Voice AI - Real-Time Inference Lead
Staff ML Engineer, Voice AI - Real-Time Inference Lead

Together AI • San Francisco (CA)

On-site
USD 220,000 - 280,000
Health insurance
Startup equity
Competitive benefits