Senior ML Inference Systems Engineer

Gimlet Labs

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

A tech startup focused on AI workloads is seeking a Member of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and improving execution performance across various components. Ideal candidates should have strong software engineering skills and experience with ML inference systems, particularly in Python and C++. This position is an opportunity to contribute to cutting-edge AI technology in a dynamic environment.

Qualifications

  • Experience with building or operating ML inference or model serving systems.
  • Ability to reason about performance, memory usage, and system behavior under load.

Responsibilities

  • Design and optimize end-to-end inference pipelines.
  • Build inference runtimes balancing latency and throughput.
  • Profile and debug inference performance issues.

Skills

Strong software engineering fundamentals
Experience building ML inference systems
Performance and memory reasoning

Tools

TensorRT-LLM
vLLM
Python
C++

Job description

A tech startup focused on AI workloads is seeking a Member of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and improving execution performance across various components. Ideal candidates should have strong software engineering skills and experience with ML inference systems, particularly in Python and C++. This position is an opportunity to contribute to cutting-edge AI technology in a dynamic environment.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Inference Systems Engineer
ML Inference Systems Engineer

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior ML Systems Engineer - Model Inference & Efficiency
Senior ML Systems Engineer - Model Inference & Efficiency

Cohere • New York (NY)

Hybrid
USD 100,000 - 150,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+4
Staff Engineer - ML Inference & Model Efficiency
Staff Engineer - ML Inference & Model Efficiency

Cohere • San Francisco (CA)

Remote
USD 180,000 - 240,000
Inclusive work culture
Weekly lunch stipend
Full health and dental benefits
+4
Senior ML Inference Systems Engineer
Senior ML Inference Systems Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Inference & RL Systems Engineer (Scalable ML Infra)
Senior Inference & RL Systems Engineer (Scalable ML Infra)

Magic AI, Inc • San Francisco (CA)

On-site
USD 300,000 - 550,000
Equity compensation
401(k) matching
Health, dental and vision insurance
+4
Senior Inference Performance Engineer — GPU & CUDA
Senior Inference Performance Engineer — GPU & CUDA

Inference • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Staff ML Engineer: Build Ultra-Fast AI at Scale (Relocation)
Staff ML Engineer: Build Ultra-Fast AI at Scale (Relocation)

Inworld AI • Mountain View (CA)

On-site
USD 270,000 - 500,000
Relocation assistance
Equity options
Comprehensive benefits package
Senior ML Systems Engineer - Distributed AI Inference
Senior ML Systems Engineer - Distributed AI Inference

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 120,000 - 160,000
Senior Inference Systems Engineer — Large-Scale GPUs
Senior Inference Systems Engineer — Large-Scale GPUs

RadixArk • Palo Alto (CA)

On-site
USD 190,000 - 260,000
Competitive compensation
Meaningful equity
Comprehensive benefits
+1