Machine Learning Inference Engineer

Oscar Technology

San Francisco (CA)

Hybrid

USD 225,000 - 275,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
401k matching
Medical coverage
Competitive base salary

Job summary

Oscar Associates Limited (US) in San Francisco offers a hybrid Senior ML Inference Engineer role. You will enhance AI-native infrastructure powering generative and multimodal models, with ownership across inference systems and production performance.

The ideal candidate has 3+ years of experience, strong Python and PyTorch skills, and familiarity with GPU tooling like Triton, TensorRT, and vLLM. Equity and full benefits accompany a competitive base salary.

Qualifications

  • 3+ years of professional ML experience required.
  • Strong understanding of GPU infrastructure for AI-native systems.
  • Hands-on with Python and PyTorch; building model-serving microservices.

Responsibilities

  • Improve efficiency of AI-native infrastructure for generative and multimodal models.
  • Own inference systems and production model performance.
  • Collaborate across teams to optimize deployment.
  • Experience with diffusion and multimodal models is a plus.

Skills

GPU infra
Python
PyTorch
Triton
TensorRT
vLLM
Diffusion
Model-serving
Multimodal models

Tools

Triton
TensorRT

Job description

Title: ML Inference Engineer

Location: San Francisco, CA

Salary: $250k base + equity

An AI Unicorn startup is hiring a Senior Machine Learning Inference Engineer for a full-time role.

You will be responsible for improving efficiency for AI-native infrastructure powered by generative and multimodal models. The ideal candidate has over 3 years of professional experience and a strong understanding of GPU infrastructure, Python, and PyTorch. This is a highly autonomous role with significant ownership across inference systems and model performance in production.

This role is hybrid in San Francisco Bay Area and offers full benefits and equity.

Experience:

  • Building AI applications at scale from the ground up

  • Strong understanding of GPU infrastructure including Triton, TensorRT, or vLLM frameworks

Hands-on experience with Python and PyTorch

  • Building model-serving Microservices

  • Diffusion and Multimodal model experience is a plus

Benefits:

  • Competitive base salary

  • Equity

  • $401k matching

  • Medical coverage

Oscar Associates Limited (US) is acting as an Employment Agency in relation to this vacancy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Inference Engineer
Machine Learning Inference Engineer

Oscar • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Equity
401k matching
Medical coverage
Senior ML Inference Engineer — AI Infrastructure & Equity
Senior ML Inference Engineer — AI Infrastructure & Equity

Oscar Technology • San Francisco (CA)

Hybrid
USD 225,000 - 275,000
Equity
401k matching
Medical coverage
+1
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Kindredventures • Palo Alto (CA)

On-site
USD 190,000 - 250,000
Comprehensive health insurance
Dental insurance
Vision insurance
+1
Senior ML Inference Engineer: Scale AI in Production
Senior ML Inference Engineer: Scale AI in Production

Oscar • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Equity
401k matching
Medical coverage
Founding Engineer, ML Inference
Founding Engineer, ML Inference

Reactor • San Francisco (CA)

On-site
USD 180,000 - 280,000
Competitive salary
Early equity
Health, dental, and vision coverage
+1
ML Inference Engineer San Francisco · Engineering · Full Time →
ML Inference Engineer San Francisco · Engineering · Full Time →

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Visa sponsorship
Relocation support
Generous health, dental, and vision coverage
Machine Learning Engineer
Machine Learning Engineer

IC Resources • San Francisco (CA)

On-site
USD 200,000 - 290,000
401(k)
Unlimited PTO
Modern engineering workspace
+1
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits