An application made for this job — a tailored resume and cover letter that speak straight to the posting.
microTECH Global Limited in the United Kingdom is seeking an AI Compiler Optimization Engineer to push inference performance on CPU and hybrid CPU/XPU systems. You will implement advanced compiler techniques and profiling to uncover optimization opportunities.
You will profile frameworks such as TensorFlow, PyTorch, ONNX, and llama.cpp, propose kernel and graph-level optimizations, and contribute insights from industry research through technical reports.
We are seeking a skilled AI Compiler Optimization Engineer to optimize AI model inference performance through advanced compiler technologies. You will focus on performance tuning for CPU or hybrid CPU/XPU heterogeneous architectures, profiling AI frameworks to discover new optimization opportunities, and delivering cutting-edge insights from industry research.
Key Responsibilities:
Compiler-Based Performance Optimization:
Implement compiler techniques (e.g., MLIR level optimizations, LLVM backend optimizations) to enhance inference performance on CPU and CPU/XPU hybrid systems
Optimize JIT level compute graphs with operator fusion, memory allocation and etc. for latency/throughput improvements
Preferred: Experience with LLVM/MLIR development
AI Model Profiling & Framework Optimization:
Profile end-to-end inference workflows on frameworks like TensorFlow, PyTorch, ONNX, and llama.cpp to identify hotspots and bottlenecks
Propose and implement optimization strategies (e.g., kernel tuning, graph-level optimizations)
Preferred: Experience optimizing models on multiple AI frameworks
Research & Insight Development:
Track and analyze the latest advancements in AI & compiler research (academic papers, open-source projects)
Produce actionable insight reports summarizing trends, benchmarks, and potential optimizations
Preferred: Strong technical writing skills with prior publications or reports