An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Gimlet Labs is building the first multi-silicon neocloud designed for fast, efficient AI inference. You will shape compiler infrastructure, optimize how workloads are represented and executed across diverse architectures, and influence scheduling, memory movement, and kernel orchestration for production-scale inference.
This ML-systems–hardware hybrid role focuses on delivering low latency and high throughput by partitioning workloads across devices and by enabling new accelerator architectures
Gimlet is building the first multi-silicon neocloud designed for fast, efficient AI inference.
We combine large-scale compute infrastructure with an execution platform that partitions AI workloads and maps each stage to the hardware best suited to run it.
We work with foundation labs, hyperscalers, and AI-native companies, giving our team access to technical problems spanning frontier models, production infrastructure, and emerging hardware.
As a Member of Technical Staff, you will build the compiler infrastructure that determines how AI workloads are represented, optimized, and executed across hardware with different architectures, performance characteristics, and memory systems.
This is compiler engineering at the boundary of ML systems and distributed execution. The problems do not end when code is generated: compiler decisions interact directly with scheduling, communication, memory movement, kernel execution, and serving performance.
You will work across intermediate representations, graph transformations, lowering, execution planning, and runtime interfaces. Compiler decisions directly shape where computation runs, how intermediate state moves between devices, which kernels execute, and ultimately the latency, throughput, and efficiency of the serving system. You will develop strategies for partitioning computation across devices, bring new models and accelerator architectures onto the platform, and partner with ML systems, kernel, and distributed systems engineers to improve end-to-end execution.
Our work on Corsair and low-latency speculative decoding is one example of the problems this team tackles.
In your first 12-18 months, you will:
Gimlet is expanding from its core technology into a production neocloud spanning new hardware, customers, and data centers.