We are partnering with an early-stage AI hardware company developing a new accelerator architecture and the complete software stack required to make it programmable.
They are seeking a Founding AI Compiler Engineer to take technical ownership of the compiler from its earliest architectural decisions through to production deployment on custom silicon.
This is an opportunity for an experienced compiler engineer to become the technical foundation of a new organisation, remain deeply hands-on and ultimately help build the team around them.
What you’ll do
- Own the technical architecture and development of a new AI compiler
- Build the initial compiler pipeline from ML frameworks and graph representations through intermediate representations to executable code
- Determine where LLVM, MLIR or other compiler infrastructure should be adopted, extended or replaced
- Translate representative AI workloads into compiler and hardware requirements
- Collaborate with architecture teams on ISA design, instruction selection, register usage, memory systems and synchronization
- Enable initial models and workloads on pre-silicon simulators before supporting first-silicon bring-up
- Investigate compiler-generated performance at the graph, kernel and instruction levels
- Define engineering standards and a technical roadmap for the compiler stack
- Help recruit and mentor the compiler team as the company grows
- Engage with early customers and application teams to understand programmability and workload requirements
What we’re looking for
- Significant experience building production compilers, preferably for GPUs, AI accelerators, DSPs, CPUs or other parallel architectures
- Deep knowledge of LLVM, MLIR or comparable compiler infrastructure
- Experience working across several stages of a compiler rather than contributing exclusively to an isolated pass
- Strong understanding of intermediate representations, lowering, code generation and compiler optimisation
- Experience developing hardware-aware transformations such as fusion, tiling, vectorisation, scheduling, memory planning or layout optimisation
- Understanding of parallel-compute architectures and the relationship between compiler decisions and hardware performance
- Ability to operate effectively in an early-stage environment with incomplete specifications and evolving hardware