Staff Engineer - Foundation Model Serving & Inference

Menlo Ventures

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

A leading tech enterprise in San Francisco is seeking a Staff Engineer to shape their foundation model API product. You will design and build systems that ensure efficient performance on high-throughput, low-latency GPU workloads. The ideal candidate will have experience with operational sensitive systems but does not need prior AI experience. Collaboration with various teams is essential to improve product offerings and architectural decisions.

Qualifications

  • No prior ML or AI experience is necessary.
  • Strong engineering skills required.
  • Ability to work in a collaborative environment.

Responsibilities

  • Design and build systems for high-throughput, low-latency inference.
  • Influence architectural direction for AI model serving.
  • Collaborate across various teams to enhance product experience.

Skills

Experience with high scale operational sensitive systems
Interest in building LLM APIs
Experience in customer facing APIs

Job description

A leading tech enterprise in San Francisco is seeking a Staff Engineer to shape their foundation model API product. You will design and build systems that ensure efficient performance on high-throughput, low-latency GPU workloads. The ideal candidate will have experience with operational sensitive systems but does not need prior AI experience. Collaboration with various teams is essential to improve product offerings and architectural decisions.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Engineer: Foundation Model API & GPU Inference
Staff Engineer: Foundation Model API & GPU Inference

Databricks Inc. • San Francisco (CA)

On-site
USD 192,000 - 260,000
Comprehensive benefits
Diversity and inclusion initiatives
Staff Backend Engineer, Foundation Model Serving
Staff Backend Engineer, Foundation Model Serving

Menlo Ventures • San Francisco (CA)

On-site
USD 192,000 - 260,000
Comprehensive benefits
Annual performance bonus
Equity options
Staff Software Engineer, Foundational Model Serving
Staff Software Engineer, Foundational Model Serving

Cacheflow • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Model Serving Engineer - Low-Latency AI Platform
Senior Model Serving Engineer - Low-Latency AI Platform

Menlo Ventures • San Francisco (CA)

On-site
USD 192,000 - 260,000
Comprehensive benefits
Eligibility for annual performance bonus
Equity opportunities
Senior AI Model Serving Engineer — Low-Latency Inference
Senior AI Model Serving Engineer — Low-Latency Inference

Menlo Ventures • San Francisco (CA)

On-site
USD 166,000 - 225,000
Annual performance bonus
Equity options
Comprehensive benefits package
Staff Inference Engineer — Production-Scale AI Platform
Staff Inference Engineer — Production-Scale AI Platform

Designworks Talent LLC • Bellevue (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
401(k) plan with company match
Paid holidays
Foundation Model Architect & AI Infrastructure Leader
Foundation Model Architect & AI Infrastructure Leader

Getvinci • Palo Alto (CA)

Hybrid
USD 130,000 - 200,000
High ownership opportunities
Work with Tier-1 clients
Innovative projects in AI
Senior Foundation Model Infra Architect
Senior Foundation Model Infra Architect

Apple Inc. • Seattle (WA)

On-site
USD 185,000 - 325,000
Employee stock programs
Medical and dental coverage
Retirement benefits
+2
Senior ML Foundation Model Compute Infra Engineer
Senior ML Foundation Model Compute Infra Engineer

Apple Inc. • Cupertino (CA)

On-site
USD 184,700 - 324,800
Stock programs
Relocation
Tuition reimbursement
+1
Foundation Model AI Infra Architect - Production-Scale
Foundation Model AI Infra Architect - Production-Scale

Vinci4D.ai • Palo Alto (CA)

On-site
USD 180,000 - 220,000