Member of Technical Staff - Compilers

The Consensus

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 230,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Gimlet is hiring an experienced systems/compiler infra engineer to shape the execution stack for diverse AI workloads. You will work across compiler infrastructure, runtime systems, scheduling, memory movement, kernel orchestration, and serving optimization to run workloads efficiently in production.

We expect strong C++ and Python skills, a solid foundation in compiler theory, and a degree in a relevant field.

Qualifications

  • Bachelor's degree in a relevant field.
  • Strong systems and performance engineering fundamentals.
  • Experience building compiler systems or execution/runtime infrastructure.
  • Experience implementing IR transformations, compiler passes, or code generation systems.

Responsibilities

  • Build compiler and runtime infrastructure for low latency and high throughput in AI inference workloads.
  • Design execution strategies to partition and coordinate workloads across heterogeneous hardware.
  • Develop compiler optimizations spanning IR transformations, scheduling, memory movement, and kernel orchestration.
  • Enable new model architectures and serving techniques for production efficiency.
  • Influence the architecture of an execution platform for future AI workloads.

Skills

Systems engineering
Compiler experience
IR transformations
Scheduling/Execution
C++
Python

Education

Bachelor's degree in a relevant field

Job description

About Us

Gimlet is building the first multi-silicon neocloud designed for fast, efficient inference.

As AI workloads become more complex and new hardware architectures emerge, simply deploying more GPUs isn't enough. The challenge is making increasingly diverse compute work together.

Gimlet's platform intelligently partitions and routes workloads across heterogeneous hardware, enabling step-function improvements in performance and efficiency. Customers deploy through production-grade APIs without needing to think about hardware selection, placement, or optimization.

We work with foundation labs, hyperscalers, and AI-native companies to power production workloads at massive scale and help define the infrastructure layer for the future of AI. This gives our team access to systems research problems grounded in frontier models, cutting-edge production workloads, and emerging hardware architectures.

About the role

At Gimlet, we believe every hire changes the company.

As a an early-stage company, talent density matters more than headcount. The engineers we hire today will shape the systems, culture, and standards that define Gimlet for years to come.

The future of AI infrastructure will not be built on a single hardware platform. It will be built on software capable of intelligently orchestrating increasingly heterogeneous compute to unprecedented scale.

Compilers sit at the center of that challenge. The performance gains unlocked at this layer compound across every workload that runs on the platform.

This role is an opportunity to help build the execution stack that transforms modern AI workloads into efficient programs running across diverse hardware architectures.

You will work across compiler infrastructure, runtime systems, scheduling, memory movement, kernel orchestration, and serving optimization to improve how AI workloads are executed in production.

This is not a traditional compiler role.

We are not building a language compiler in isolation.

We are building the systems that determine how AI workloads are partitioned, optimized, scheduled, and executed across the next generation of AI infrastructure.

You'll work on MLIR transformations, execution planning, speculative decoding optimization, heterogeneous scheduling, runtime optimization, and serving infrastructure that powers production AI workloads at scale.

To learn more about the kinds of systems we build, see our work on Corsair and low-latency speculative decoding:

see our work on Corsair and low-latency speculative decoding: https://gimletlabs.ai/blog/low-latency-spec-decode-corsair

What success looks like

In your first 12-18 months, you will help:

  • Build compiler and runtime infrastructure that improves latency, throughput, and efficiency for large-scale AI inference workloads.

  • Design execution strategies that intelligently partition and coordinate workloads across heterogeneous hardware.

  • Develop compiler optimizations spanning IR transformations, scheduling, memory movement, and kernel orchestration.

  • Enable new model architectures and serving techniques to run efficiently in production environments.

  • Influence the architecture of an execution platform that will help define how AI workloads are deployed over the next decade.

You may be a good fit if
  • Strong systems and performance engineering fundamentals

  • Experience building compiler systems, compiler-adjacent infrastructure, or execution/runtime systems

  • Experience implementing IR transformations, compiler passes, lowering logic, or code generation systems

  • Ability to reason about execution behavior, memory systems, scheduling, and hardware efficiency

  • Strong software engineering skills in C++ and/or Python

  • Bachelor's degree in a relevant field, or an equivalent combination of education, training, and professional experience.

Strong candidates may also have
  • Experience with MLIR, LLVM, XLA, TVM, Triton, or similar compiler/runtime infrastructure

  • Experience optimizing ML inference or serving workloads

  • Familiarity with runtime systems, kernel dispatch, launch APIs, or memory allocators

  • Experience working with GPUs, AI accelerators, or heterogeneous hardware systems

  • Experience profiling and debugging performance-critical systems

  • Familiarity with scheduling, partitioning, or kernel-level optimizations

Why join now?

Gimlet is at the very beginning of its journey, and that's what makes this moment special. Most AI infrastructure companies are focused on deploying more compute. We are focused on making increasingly diverse compute work together, and that ambition touches every part of how we build and run this company.

As an early member of the team, you will have significant ownership over your work, partner directly with a small group of highly capable people, and help shape not just what we build, but how we scale the company.

We value people who are excited to work across domains, take ownership of meaningful problems, and help define what Gimlet becomes over the next several years.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding Engineer - ML Infrastructure
Founding Engineer - ML Infrastructure

uRun • San Francisco (CA)

On-site
USD 120,000 - 160,000
Health, dental, and vision
401(k)
Paid time off
+2
Member of Technical Staff - Infrastructure
Member of Technical Staff - Infrastructure

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Founding Engineer - ML Performance
Founding Engineer - ML Performance

uRun • San Francisco (CA)

On-site
USD 120,000 - 160,000
Health, dental, and vision
401(k) participation
Flexible spending accounts
+3
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff, GPU Compiler
Member of Technical Staff, GPU Compiler

SF Tensor • San Francisco (CA)

On-site
USD 285,000 - 315,000
Relocation assistance
Member of Technical Staff, GPU Kernels
Member of Technical Staff, GPU Kernels

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 285,000 - 315,000
Relocation assistance
Equity
Member of Technical Staff, AI-Driven Compilation
Member of Technical Staff, AI-Driven Compilation

SF Tensor • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance
Equity and benefits
Office in San Francisco
Member of Technical Staff, GPU Compiler
Member of Technical Staff, GPU Compiler

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 285,000 - 315,000
Relocation assistance
Office in San Francisco
Staff Compiler Engineer for AI Systems & Heterogeneous HW
Staff Compiler Engineer for AI Systems & Heterogeneous HW

The Consensus • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 230,000
Developer Relations
Developer Relations

gimlet • San Francisco (CA)

On-site
USD 120,000 - 180,000