Turn this role into an interview — a resume and cover letter built around what this employer wants.
Evollabs is seeking a senior compiler engineer to design MLIR dialects and lowering pipelines for our AI accelerator. You will lower ML graphs to optimized kernels, focusing on fusion, tiling and vectorization, and profile for latency and throughput on target hardware.
You will collaborate with Silicon, firmware and runtime teams to co‑design features, build a portable compiler stack, and drive compiler innovations across generations of hardware.
We are a tech company specializing in the design and development of cutting-edge, customized server hardware solutions optimized for artificial intelligence and machine learning applications. Our mission is to empower businesses and researchers to accelerate their AI initiatives by providing them with high-performance, scalable, and energy-efficient hardware infrastructure.
As a rapidly growing company at the forefront of AI hardware innovation, we are constantly seeking talented and motivated individuals to join our team. We offer a dynamic and challenging work environment, with opportunities to make a significant impact on the future of AI technology.
Silicon, firmware, runtime, datacenter-software and architecture teams to build a mlir-compiler and software stack that brings next-gen AI/NPU hardware into modern datacenter environments and production AI workloads.
Define and implement MLIR dialects and lowering pipelines targeting our AI-accelerator/NPU platform.
Lower ML graphs (e.g., from ONNX/PyTorch) into optimized kernels focusing on fusion, tiling and vectorization.
Profile, benchmark and optimize compiler output to meet latency, throughput and power targets.
Work with hardware architects and runtime/framework teams to co-design compiler features.
Build a future‑ready, portable compiler stack that supports multiple hardware generations.
5+ years of experience in compiler engineering or high‑performance systems.
Strong programming skills in C++ (modern standards), and scripting or tooling in Python.
Hands‑on experience with MLIR and/or LLVM (designing/extending dialects, lowering, optimization).
Deep understanding of compiler internals: IRs, code scheduling, vectorization, loop transformations.
Strong familiarity with AI/ML frameworks (e.g.,PyTorch, ONNX, TensorFlow).
Excellent understanding of ML operations such as tensor operators, matmuls, quantization, dynamic shapes.
Robust understanding of computer architecture and memory hierarchy.
Bachelor’s (or higher) degree in Computer Science, Computer Engineering or equivalent.
Prior experience targeting NPU/AI accelerator hardware.
Contributions to open‑source compiler projects (MLIR, LLVM,Torch-MLIR, ONNX-MLIR).
Familiarity with runtime systems, scheduling, resource sharing, memory movement engines and, interconnects.
Previous leadership role driving compiler architecture, setting standards, mentoring teams and owning roadmap.
Familiarity with software development tooling: Git, CI/CD, debuggers, profilers.
Advanced degree (M.S./Ph.D.) preferred.
A production compiler IR/dialect is designed and integrated, successfully lowering key ML workloads to the target NPU.
Key operator kernels (matmul, convolution, etc) show measurable performance gains (reduced latency, increased throughput) on the hardware or simulator.
The compiler pipeline handles multi‑die/chiplet topology correctly and efficiently schedules across devices.
Runtime interfaces and resource scheduling between host and accelerator are functional and validated with real workloads.
Telemetry/debug hooks and performance counters from compiled code are available and used for performance analysis by other teams.
The architecture and roadmap for next‑gen hardware are defined, and junior engineers are actively mentored within the compiler team.
Join us in our mission to democratize AI compute — where your firmware expertise becomes the bedrock of tomorrow's AI breakthroughs.