ML Inference Systems Engineer

Gimlet Labs, Inc.

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Gimlet Labs, Inc. is looking for a Member of Technical Staff focused on ML systems and inference in San Francisco. You will design and build inference systems under real production constraints, ensuring fast, predictable, and scalable performance. Key responsibilities include optimizing inference pipelines and managing cache allocation. Candidates should have strong foundations in software engineering, experience with ML inference systems, and performance tuning capabilities. Familiarity with Python and C++ is preferred.

Qualifications

  • Strong software engineering fundamentals.
  • Experience building or operating ML inference or model serving systems.
  • Comfort reasoning about performance, memory usage, and system behavior under load.

Responsibilities

  • Design and optimize end‑to‑end inference pipelines from request ingestion through execution and response.
  • Build and evolve inference runtimes that balance latency, throughput, and concurrency under real‑world load.
  • Manage KV cache allocation, placement, reuse, and eviction across models and requests.

Skills

Software engineering fundamentals
Building ML inference systems
Performance reasoning

Tools

TensorRT-LLM
vLLM
Python
C++

Job description

Gimlet Labs, Inc. is looking for a Member of Technical Staff focused on ML systems and inference in San Francisco. You will design and build inference systems under real production constraints, ensuring fast, predictable, and scalable performance. Key responsibilities include optimizing inference pipelines and managing cache allocation. Candidates should have strong foundations in software engineering, experience with ML inference systems, and performance tuning capabilities. Familiarity with Python and C++ is preferred.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
ML Inference & Systems Architect
ML Inference & Systems Architect

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Senior ML Inference Systems Engineer
Senior ML Inference Systems Engineer

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Staff Engineer - ML Inference & Model Efficiency
Staff Engineer - ML Inference & Model Efficiency

Cohere • San Francisco (CA)

Remote
USD 180,000 - 240,000
Inclusive work culture
Weekly lunch stipend
Full health and dental benefits
+4
Senior Staff Tech Lead — Inference & ML Performance
Senior Staff Tech Lead — Inference & ML Performance

fal • San Francisco (CA)

On-site
USD 150,000 - 200,000
Head of ML Systems & Inference
Head of ML Systems & Inference

Doist • San Francisco (CA)

On-site
USD 260,000 - 380,000
Health, dental, vision benefits
401(k) company match
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
ML Inference Infrastructure Engineer
ML Inference Infrastructure Engineer

Baseten • San Francisco (CA)

On-site
USD 90,000 - 130,000
100% coverage of medical, dental, and vision insurance
Generous PTO policy including Winter Break
Company-facilitated 401(k)
Senior ML Engineer - Real-Time Inference & Scalable Systems
Senior ML Engineer - Real-Time Inference & Scalable Systems

careers.bitkraft.vc - Jobboard • Germany (OH)

On-site
USD 120,000 - 160,000