Member of Technical Staff - ML Systems & Inference

Gimlet Labs, Inc.

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Gimlet Labs, Inc. is looking for a Member of Technical Staff focused on ML systems and inference in San Francisco. You will design and build inference systems under real production constraints, ensuring fast, predictable, and scalable performance. Key responsibilities include optimizing inference pipelines and managing cache allocation. Candidates should have strong foundations in software engineering, experience with ML inference systems, and performance tuning capabilities. Familiarity with Python and C++ is preferred.

Qualifications

  • Strong software engineering fundamentals.
  • Experience building or operating ML inference or model serving systems.
  • Comfort reasoning about performance, memory usage, and system behavior under load.

Responsibilities

  • Design and optimize end‑to‑end inference pipelines from request ingestion through execution and response.
  • Build and evolve inference runtimes that balance latency, throughput, and concurrency under real‑world load.
  • Manage KV cache allocation, placement, reuse, and eviction across models and requests.

Skills

Software engineering fundamentals
Building ML inference systems
Performance reasoning

Tools

TensorRT-LLM
vLLM
Python
C++

Job description

About Us

Gimlet Labs is building the first heterogeneous neocloud for AI workloads. As AI systems scale, the industry is hitting fundamental limits in power, capacity, and cost with today’s homogeneous, vertically integrated infrastructure. Gimlet addresses this by decoupling AI workloads from the underlying hardware. Our platform intelligently partitions workloads into components and orchestrates each component to hardware that best fits its performance and efficiency needs. This approach enables heterogeneous systems across multi‑vendor and multi‑generation hardware, including the latest emerging accelerators. These systems unlock step‑function improvements in performance and cost efficiency at scale. On top of this foundation, Gimlet is building a production‑grade neocloud for agentic workloads. Customers use Gimlet to deploy and manage their workloads through stable, production‑ready APIs, without having to reason about hardware selection, placement, or low‑level performance optimization. Gimlet works with foundation labs, hyperscalers, and AI native companies to power real production workloads built to scale to gigawatt‑class AI datacenters.

Mission

Gimlet Labs is seeking a Member of Technical Staff focused on ML systems and inference. In this role, you will design and build inference systems that execute full models end‑to‑end under real production constraints. You will work at the intersection of model architecture, runtime behavior, and system performance to ensure inference is fast, predictable, and scalable. This role is ideal for engineers who deeply understand how modern models execute in practice and who care about latency, throughput, and memory behavior across the full inference lifecycle.

Responsibilities
  • Design and optimize end‑to‑end inference pipelines from request ingestion through execution and response.
  • Build and evolve inference runtimes that balance latency, throughput, and concurrency under real‑world load.
  • Reason about batching, queuing, and scheduling tradeoffs, including their impact on tail latency and fairness.
  • Manage KV cache allocation, placement, reuse, and eviction across models and requests.
  • Optimize prefill and decode paths, including attention mechanisms and memory usage.
  • Profile and debug inference performance issues across model, runtime, and system boundaries.
  • Work closely with compilers, kernels, networking, and distributed systems to deliver end‑to‑end performance improvements.
Qualifications
  • Strong software engineering fundamentals.
  • Experience building or operating ML inference or model serving systems.
  • Comfort reasoning about performance, memory usage, and system behavior under load.
Preferred Qualifications
  • Experience with inference runtimes such as TensorRT‑LLM, vLLM, or custom serving systems.
  • Deep understanding of modern model architectures and attention mechanisms.
  • Experience with batching, scheduling, and concurrency control in inference systems.
  • Familiarity with KV cache management and memory placement strategies.
  • Experience profiling and tuning latency‑and throughput‑critical systems.
  • Software development experience in Python and C++.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff - Infrastructure
Member of Technical Staff - Infrastructure

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000
Member of Technical Staff - Kernels & GPU Performance
Member of Technical Staff - Kernels & GPU Performance

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 150,000 - 350,000
Member of Technical Staff - Kernels & GPU Performance
Member of Technical Staff - Kernels & GPU Performance

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Machine Learning Engineer (LLM inference)
Machine Learning Engineer (LLM inference)

GMI Cloud • Mountain View (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - Distributed Systems
Member of Technical Staff - Distributed Systems

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Machine Learning Engineer, LLM Inference Optimization in Sonoma
Machine Learning Engineer, LLM Inference Optimization in Sonoma

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000