Member of Technical Staff - Compiler Engineer

Gimlet Labs

San Francisco (CA)

On-site

USD 190,000 - 260,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Gimlet Labs is building the first multi-silicon neocloud designed for fast, efficient AI inference. You will shape compiler infrastructure, optimize how workloads are represented and executed across diverse architectures, and influence scheduling, memory movement, and kernel orchestration for production-scale inference.

This ML-systems–hardware hybrid role focuses on delivering low latency and high throughput by partitioning workloads across devices and by enabling new accelerator architectures

Qualifications

  • Experience building compiler, runtime, or execution infrastructure.
  • Experience with IR transformations, compiler passes, lowering, or code generation.
  • Strong systems and performance-engineering fundamentals.
  • Ability to reason about execution behavior, memory systems, scheduling, and hardware efficiency.
  • Strong C++ and/or Python skills.
  • A bachelor's degree in a relevant field or equivalent practical experience.

Responsibilities

  • Improve latency, throughput, and efficiency of production inference workloads.
  • Design execution strategies for partitioning workloads across heterogeneous hardware.
  • Develop compiler optimizations spanning IR transformations, scheduling, memory movement, and kernel orchestration.
  • Enable new models, accelerator architectures, and serving techniques to run efficiently in production.

Skills

C++
Python
Compiler infra

Education

Bachelor's degree in a relevant field

Tools

MLIR
LLVM
XLA
TVM
Triton

Job description

About us

Gimlet is building the first multi-silicon neocloud designed for fast, efficient AI inference.


We combine large-scale compute infrastructure with an execution platform that partitions AI workloads and maps each stage to the hardware best suited to run it.


We work with foundation labs, hyperscalers, and AI-native companies, giving our team access to technical problems spanning frontier models, production infrastructure, and emerging hardware.


About the role

As a Member of Technical Staff, you will build the compiler infrastructure that determines how AI workloads are represented, optimized, and executed across hardware with different architectures, performance characteristics, and memory systems.


This is compiler engineering at the boundary of ML systems and distributed execution. The problems do not end when code is generated: compiler decisions interact directly with scheduling, communication, memory movement, kernel execution, and serving performance.


You will work across intermediate representations, graph transformations, lowering, execution planning, and runtime interfaces. Compiler decisions directly shape where computation runs, how intermediate state moves between devices, which kernels execute, and ultimately the latency, throughput, and efficiency of the serving system. You will develop strategies for partitioning computation across devices, bring new models and accelerator architectures onto the platform, and partner with ML systems, kernel, and distributed systems engineers to improve end-to-end execution.


Our work on Corsair and low-latency speculative decoding is one example of the problems this team tackles.


What success looks like

In your first 12-18 months, you will:



  • Improve the latency, throughput, and efficiency of production inference workloads

  • Design execution strategies for partitioning workloads across heterogeneous hardware

  • Develop compiler optimizations spanning IR transformations, scheduling, memory movement, and kernel orchestration

  • Enable new models, accelerator architectures, and serving techniques to run efficiently in production


You may be a good fit if you have


  • Experience building compiler, runtime, or execution infrastructure

  • Experience with IR transformations, compiler passes, lowering, or code generation

  • Strong systems and performance-engineering fundamentals

  • The ability to reason about execution behavior, memory systems, scheduling, and hardware efficiency

  • Strong C++ and/or Python skills

  • A bachelor's degree in a relevant field or equivalent practical experience


Strong candidates may also have


  • Experience with MLIR, LLVM, XLA, TVM, Triton, or similar compiler/runtime infrastructure

  • Experience optimizing ML inference or serving workloads

  • Familiarity with runtime systems, kernel dispatch, launch APIs, or memory allocators

  • Experience working with GPUs, AI accelerators, or heterogeneous hardware systems

  • Experience profiling and debugging performance-critical systems

  • Familiarity with scheduling, partitioning, or kernel-level optimizations


Why join now?

Gimlet is expanding from its core technology into a production neocloud spanning new hardware, customers, and data centers.



  • Solve hard problems.

  • Own meaningful work.

  • Build for production.

  • Help define what's next.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Compilers
Member of Technical Staff - Compilers

The Consensus • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 230,000
Member of Technical Staff, GPU Compiler
Member of Technical Staff, GPU Compiler

SF Tensor • San Francisco (CA)

On-site
USD 285,000 - 315,000
Relocation assistance
Member of Technical Staff, GPU Kernels
Member of Technical Staff, GPU Kernels

SF Tensor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, GPU Compiler
Member of Technical Staff, GPU Compiler

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 285,000 - 315,000
Relocation assistance
Office in San Francisco
Member of Technical Staff: AI Inference Compiler Engineer
Member of Technical Staff: AI Inference Compiler Engineer

Gimlet Labs • San Francisco (CA)

On-site
USD 190,000 - 260,000
Member of Technical Staff, AI-Driven Compilation
Member of Technical Staff, AI-Driven Compilation

SF Tensor • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance
Equity and benefits
Office in San Francisco
Member of Technical Staff - Applied AI Research
Member of Technical Staff - Applied AI Research

Gimlet Labs, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff, GPU Kernels
Member of Technical Staff, GPU Kernels

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 285,000 - 315,000
Relocation assistance
Equity
Member of Technical Staff - Kernels & GPU Performance
Member of Technical Staff - Kernels & GPU Performance

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000