Senior Model Serving Engineer - Low-Latency AI Platform

Menlo Ventures

San Francisco (CA)

On-site

USD 192,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Comprehensive benefits
Eligibility for annual performance bonus
Equity opportunities

Job summary

A leading data and AI company in San Francisco is seeking a Staff Engineer to design and implement systems for their AI/ML Model Serving platform. You will collaborate with product, infrastructure, and research teams to ensure high-performance system delivery. The ideal candidate has over 10 years of experience in distributed systems and model serving. This role offers a competitive salary and a commitment to diversity and inclusion.

Qualifications

  • 10+ years of experience building and operating large-scale distributed systems.
  • Deep expertise in model serving and related infrastructure.
  • Proven ability to deliver technically complex, high-impact initiatives.

Responsibilities

  • Design and implement systems and APIs for Databricks Model Serving.
  • Define the technical roadmap for serving workloads.
  • Lead initiatives to improve performance and cost-effectiveness.

Skills

Building and operating large-scale distributed systems
Model serving and inference systems
Algorithms and system design
Strong communication skills
Mentoring and technical guidance

Job description

A leading data and AI company in San Francisco is seeking a Staff Engineer to design and implement systems for their AI/ML Model Serving platform. You will collaborate with product, infrastructure, and research teams to ensure high-performance system delivery. The ideal candidate has over 10 years of experience in distributed systems and model serving. This role offers a competitive salary and a commitment to diversity and inclusion.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Model Serving Engineer — Low-Latency Inference
Senior AI Model Serving Engineer — Low-Latency Inference

Menlo Ventures • San Francisco (CA)

On-site
USD 166,000 - 225,000
Annual performance bonus
Equity options
Comprehensive benefits package
Lead AI Model-Serving Platform Engineer
Lead AI Model-Serving Platform Engineer

Sciforium • San Francisco (CA)

On-site
USD 180,000 - 240,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
Senior Engineer, Model Serving & Inference
Senior Engineer, Model Serving & Inference

Databricks • San Francisco (CA)

On-site
USD 166,000 - 225,000
Staff ML Infrastructure & Performance Engineer
Staff ML Infrastructure & Performance Engineer

Embedding VC • San Mateo (CA)

On-site
USD 120,000 - 150,000
Staff Engineer, AI-Driven Customer Engagement Platform
Staff Engineer, AI-Driven Customer Engagement Platform

Databricks Inc. • San Francisco (CA)

On-site
USD 190,000 - 260,000
Comprehensive benefits
Inclusive culture
Opportunities for advancement
Senior Model Inference Engineer for Production-Scale AI
Senior Model Inference Engineer for Production-Scale AI

OpenAI • San Francisco (CA)

On-site
USD 325,000 - 490,000
Senior ML Engineer - Low-Latency Inference & Systems
Senior ML Engineer - Low-Latency Inference & Systems

Inworld • Germany (OH)

Hybrid
USD 120,000 - 160,000
Senior ML Infra Engineer — Real-Time, Low-Latency
Senior ML Infra Engineer — Real-Time, Low-Latency

LMArena • San Francisco (CA)

On-site
USD 170,000 - 260,000
Competitive compensation
Comprehensive health benefits
Opportunity to work on cutting-edge AI
Staff Engineer - Foundation Model Serving & Inference
Staff Engineer - Foundation Model Serving & Inference

Menlo Ventures • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Model API Engineer – Low-Latency AI Serving
Senior Model API Engineer – Low-Latency AI Serving

Baseten • United States

Remote
USD 150,000 - 230,000