Senior ML Inference Systems Engineer

Gimlet Labs

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A tech startup focused on AI workloads is seeking a Member of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and improving execution performance across various components. Ideal candidates should have strong software engineering skills and experience with ML inference systems, particularly in Python and C++. This position is an opportunity to contribute to cutting-edge AI technology in a dynamic environment.

Qualifications

  • Experience with building or operating ML inference or model serving systems.
  • Ability to reason about performance, memory usage, and system behavior under load.

Responsibilities

  • Design and optimize end-to-end inference pipelines.
  • Build inference runtimes balancing latency and throughput.
  • Profile and debug inference performance issues.

Skills

Strong software engineering fundamentals
Experience building ML inference systems
Performance and memory reasoning

Tools

TensorRT-LLM
vLLM
Python
C++

Job description

A tech startup focused on AI workloads is seeking a Member of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and improving execution performance across various components. Ideal candidates should have strong software engineering skills and experience with ML inference systems, particularly in Python and C++. This position is an opportunity to contribute to cutting-edge AI technology in a dynamic environment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Inference Systems Engineer
ML Inference Systems Engineer

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Senior ML Systems Engineer - Model Inference & Efficiency
Senior ML Systems Engineer - Model Inference & Efficiency

Cohere • New York (NY)

Hybrid
USD 100,000 - 150,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+4
Senior ML Engineer - Low-Latency Inference & Systems
Senior ML Engineer - Low-Latency Inference & Systems

Inworld • Germany (OH)

Hybrid
USD 120,000 - 160,000
Senior ML Engineer - Real-Time Inference & Systems
Senior ML Engineer - Real-Time Inference & Systems

Inworld AI • Germany (OH)

On-site
USD 120,000 - 180,000
Staff Engineer - ML Inference & Model Efficiency
Staff Engineer - ML Inference & Model Efficiency

Cohere • San Francisco (CA)

Remote
USD 180,000 - 240,000
Inclusive work culture
Weekly lunch stipend
Full health and dental benefits
+4
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000
ML Inference & Systems Architect
ML Inference & Systems Architect

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Senior Staff Tech Lead — Inference & ML Performance
Senior Staff Tech Lead — Inference & ML Performance

fal • San Francisco (CA)

On-site
USD 150,000 - 200,000
Senior ML Systems Engineer — Inference & Scale
Senior ML Systems Engineer — Inference & Scale

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000